Choosing a granularity
docray emits three output shapes. This is the most important decision you make as a consumer — it changes payload size by more than an order of magnitude.
The short version: if an LLM reads the output, use
element, then select the token-leanleanoutput format when you do not need the JSON provenance envelope. It carries the text, position, and style of every element at ~7% of the lossless payload. Only move towordwhen you need word-level highlighting, and tocharwhen you need the full archival hierarchy.
The three levels
| Level | Shape | Measured size¹ | Use when |
|---|---|---|---|
element | one text string + bbox per element | −92.9% | LLM/RAG consumption, semantic processing, citations |
word | flat [text, x0, y0, x1, y1] tuples | −89.2% | word-precise highlighting, search-hit boxes |
char (default) | full char → word → line hierarchy | baseline | archival, ML training data, anything lossless |
¹ Measured across a mixed corpus of real documents (bank statements, pitch
decks, a 49-page signed contract) totalling 22 MB of char output.
element
docray extract file.pdf --granularity element
# or: POST /v1/extract?granularity=element
{
"type": "text",
"bbox": [399.6, 90.7, 531.5, 99.4],
"text": "Customer service information",
"font": { "name": "ConnectionsBold_CZEX0AA0", "size": 9.5, "bold": true },
"color": { "fill": [35, 31, 32] }
}
Images reduce to {"type": "image", "bbox": [...]}. Paths keep their bbox
plus optional fill, stroke, and stroke_width; annotations keep their
subtype and uri. The bbox is still precise enough for click-to-source
highlighting.
word
{
"type": "text",
"bbox": [399.6, 90.7, 531.5, 99.4],
"font": { "name": "ConnectionsBold_CZEX0AA0", "size": 9.5, "bold": true },
"words": [
["Customer", 399.6, 90.8, 442.0, 99.4],
["service", 444.8, 90.8, 476.3, 99.4],
["information", 479.1, 90.7, 531.5, 99.4]
]
}
Each word is a positional tuple: [text, x0, y0, x1, y1]. Words appear in
content-stream (extraction) order — docray does not infer reading
order.
char (the default)
Omit the parameter and you get the lossless v1.1 contract, byte-identical across versions: every text run with nested lines, words, and per-character boxes, full font/color detail on every element, image quads and content hashes, path stroke properties. This is the archival shape — see the JSON contract.
Rules shared by the compact levels
- Coordinates round to 1 decimal (0.05 pt max displacement — chosen over integers, which produced degenerate zero-area boxes on thin paths).
- Omitted when default for text:
bold/italicwhen false,fillwhen black, andstrokewhen absent or black. Path paint is omitted only when absent, because black fill/stroke is still authored visual content. - Never omitted: page dimensions, element type, bbox,
scannedflags, and any non-emptywarningsarray. Silent-failure freedom survives every granularity. - Compact responses report
"schema_version": "1.6"and echo the"granularity"you asked for. - Every compact granularity carries each page’s non-visible
hiddenitems verbatim; granularity changes visible text detail, not supplemental context.
Granularity controls which information is retained; output format controls how that information is encoded. See output formats for the measured JSON-versus-lean tradeoff and complete lean specification.
Flow documents
DOCX and DOCM support only element and default to it when the parameter is
omitted. Their schema 1.7 flow blocks have no page bboxes, so word and char
are unavailable. For token-conscious consumers, element plus lean keeps
headings, lists, tables, runs, and hidden context without a JSON envelope. See
Word extraction.
Deliberate non-optimizations
Two further size levers were measured and rejected — documented here so you know they were choices, not oversights:
- One-letter type tags (
"t"vs"text") save only 0.3% more while making the payload harder for models and humans to read. - Document-level font/color tables save ~13% more but break self-containment: a single page or element retrieved into a RAG context would no longer describe itself.