Choosing a granularity
docray emits three output shapes. This is the most important decision you make as a consumer — it changes payload size by more than an order of magnitude.
The short version: if an LLM reads the output, use
element, then select the token-leanleanoutput format when you need positions, ormdwhen you need clean reading-order prose and do not need the JSON provenance envelope. Element JSON carries text, position, and style of every element at ~7% of the lossless payload. Only move towordwhen you need word-level highlighting, and tocharwhen you need the full archival hierarchy.
The three levels
| Level | Shape | Measured size¹ | Use when |
|---|---|---|---|
element | one text string + bbox per element | −92.9% | LLM/RAG consumption, semantic processing, citations |
word | flat [text, x0, y0, x1, y1] tuples | −89.2% | word-precise highlighting, search-hit boxes |
char (default) | full char → word → line hierarchy | baseline | archival, ML training data, anything lossless |
¹ Measured across a mixed corpus of real documents (bank statements, pitch
decks, a 49-page signed contract) totalling 22 MB of char output.
element
docray extract file.pdf --granularity element
# or: POST /v1/extract?granularity=element
{
"type": "text",
"bbox": [399.6, 90.7, 531.5, 99.4],
"text": "Customer service information",
"font": { "name": "ConnectionsBold_CZEX0AA0", "size": 9.5, "bold": true },
"color": { "fill": [35, 31, 32] }
}
Images reduce to {"type": "image", "bbox": [...]}. Paths keep their bbox
plus optional fill, stroke, and stroke_width; annotations keep their
subtype and uri. The bbox is still precise enough for click-to-source
highlighting.
word
{
"type": "text",
"bbox": [399.6, 90.7, 531.5, 99.4],
"font": { "name": "ConnectionsBold_CZEX0AA0", "size": 9.5, "bold": true },
"words": [
["Customer", 399.6, 90.8, 442.0, 99.4],
["service", 444.8, 90.8, 476.3, 99.4],
["information", 479.1, 90.7, 531.5, 99.4]
]
}
Each word is a positional tuple: [text, x0, y0, x1, y1]. Words appear in
content-stream (extraction) order — docray does not infer reading
order.
char (the default)
Omit the parameter and you get the lossless v1.1 contract, byte-identical
across versions: every text run with nested lines, words, and per-character
boxes, full font/color detail on every element, image quads and content
hashes, path stroke properties. This is the archival shape — see
the JSON contract. It never regroups anything, even on
the glyph-fragmented pages described below: a PDF that places one glyph per
text object still reports one char per text object, exactly as extracted.
Glyph-fragmented pages
Some PDF producers (certain CAD/plotter exports, font-subsetted
re-renderers) emit one text object per glyph instead of batching a run into
a single show-text operator. Left alone, that would surface as dozens of
single-character element/word entries per line instead of readable text.
element and word detect this pattern per page and regroup the glyphs into
geometric lines and words before projecting — so a glyph-fragmented page
reads the same as a normally-authored one (“hello world” as one element, not
eleven). Word boundaries that a batched run on the same page had already
resolved are carried through the regrouping rather than re-derived from glyph
positions, which at small font sizes cannot recover them. Detection and regrouping only ever run on the element/word
projection; normal pages (glyphs already batched into runs) are left
untouched, and char always reports the original, ungrouped geometry.
Rules shared by the compact levels
- Coordinates round to 1 decimal (0.05 pt max displacement — chosen over integers, which produced degenerate zero-area boxes on thin paths).
- Omitted when default for text:
bold/italicwhen false,fillwhen black, andstrokewhen absent or black. Path paint is omitted only when absent, because black fill/stroke is still authored visual content. - Never omitted: page dimensions, element type, bbox,
scannedflags, and any non-emptywarningsarray. Silent-failure freedom survives every granularity. element/wordcompact responses report"schema_version": "1.9". Requesting the explicitchargranularity reports"schema_version": "1.6"— distinct from the frozen"1.1"you get when the parameter is omitted entirely. All three echo the"granularity"you asked for.- Every compact granularity carries each page’s non-visible
hiddenitems verbatim; granularity changes visible text detail, not supplemental context.
Granularity controls which information is retained; output format controls how that information is encoded. See output formats for the measured JSON-versus-lean tradeoff and complete lean specification.
Flow documents
DOCX and DOCM support only element and default to it when the parameter is
omitted. Their schema 1.7 flow blocks have no page bboxes, so word and char
are unavailable. For token-conscious consumers, element plus lean keeps
headings, lists, tables, runs, and hidden context without a JSON envelope. See
Word extraction.
Deliberate non-optimizations
Two further size levers were measured and rejected — documented here so you know they were choices, not oversights:
- One-letter type tags (
"t"vs"text") save only 0.3% more while making the payload harder for models and humans to read. - Document-level font/color tables save ~13% more but break self-containment: a single page or element retrieved into a RAG context would no longer describe itself.