Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Choosing a granularity

docray emits three output shapes. This is the most important decision you make as a consumer — it changes payload size by more than an order of magnitude.

The short version: if an LLM reads the output, use element, then select the token-lean lean output format when you need positions, or md when you need clean reading-order prose and do not need the JSON provenance envelope. Element JSON carries text, position, and style of every element at ~7% of the lossless payload. Only move to word when you need word-level highlighting, and to char when you need the full archival hierarchy.

The three levels

LevelShapeMeasured size¹Use when
elementone text string + bbox per element−92.9%LLM/RAG consumption, semantic processing, citations
wordflat [text, x0, y0, x1, y1] tuples−89.2%word-precise highlighting, search-hit boxes
char (default)full char → word → line hierarchybaselinearchival, ML training data, anything lossless

¹ Measured across a mixed corpus of real documents (bank statements, pitch decks, a 49-page signed contract) totalling 22 MB of char output.

element

docray extract file.pdf --granularity element
# or: POST /v1/extract?granularity=element
{
  "type": "text",
  "bbox": [399.6, 90.7, 531.5, 99.4],
  "text": "Customer service information",
  "font": { "name": "ConnectionsBold_CZEX0AA0", "size": 9.5, "bold": true },
  "color": { "fill": [35, 31, 32] }
}

Images reduce to {"type": "image", "bbox": [...]}. Paths keep their bbox plus optional fill, stroke, and stroke_width; annotations keep their subtype and uri. The bbox is still precise enough for click-to-source highlighting.

word

{
  "type": "text",
  "bbox": [399.6, 90.7, 531.5, 99.4],
  "font": { "name": "ConnectionsBold_CZEX0AA0", "size": 9.5, "bold": true },
  "words": [
    ["Customer", 399.6, 90.8, 442.0, 99.4],
    ["service", 444.8, 90.8, 476.3, 99.4],
    ["information", 479.1, 90.7, 531.5, 99.4]
  ]
}

Each word is a positional tuple: [text, x0, y0, x1, y1]. Words appear in content-stream (extraction) order — docray does not infer reading order.

char (the default)

Omit the parameter and you get the lossless v1.1 contract, byte-identical across versions: every text run with nested lines, words, and per-character boxes, full font/color detail on every element, image quads and content hashes, path stroke properties. This is the archival shape — see the JSON contract. It never regroups anything, even on the glyph-fragmented pages described below: a PDF that places one glyph per text object still reports one char per text object, exactly as extracted.

Glyph-fragmented pages

Some PDF producers (certain CAD/plotter exports, font-subsetted re-renderers) emit one text object per glyph instead of batching a run into a single show-text operator. Left alone, that would surface as dozens of single-character element/word entries per line instead of readable text.

element and word detect this pattern per page and regroup the glyphs into geometric lines and words before projecting — so a glyph-fragmented page reads the same as a normally-authored one (“hello world” as one element, not eleven). Word boundaries that a batched run on the same page had already resolved are carried through the regrouping rather than re-derived from glyph positions, which at small font sizes cannot recover them. Detection and regrouping only ever run on the element/word projection; normal pages (glyphs already batched into runs) are left untouched, and char always reports the original, ungrouped geometry.

Rules shared by the compact levels

  • Coordinates round to 1 decimal (0.05 pt max displacement — chosen over integers, which produced degenerate zero-area boxes on thin paths).
  • Omitted when default for text: bold/italic when false, fill when black, and stroke when absent or black. Path paint is omitted only when absent, because black fill/stroke is still authored visual content.
  • Never omitted: page dimensions, element type, bbox, scanned flags, and any non-empty warnings array. Silent-failure freedom survives every granularity.
  • element/word compact responses report "schema_version": "1.9". Requesting the explicit char granularity reports "schema_version": "1.6" — distinct from the frozen "1.1" you get when the parameter is omitted entirely. All three echo the "granularity" you asked for.
  • Every compact granularity carries each page’s non-visible hidden items verbatim; granularity changes visible text detail, not supplemental context.

Granularity controls which information is retained; output format controls how that information is encoded. See output formats for the measured JSON-versus-lean tradeoff and complete lean specification.

Flow documents

DOCX and DOCM support only element and default to it when the parameter is omitted. Their schema 1.7 flow blocks have no page bboxes, so word and char are unavailable. For token-conscious consumers, element plus lean keeps headings, lists, tables, runs, and hidden context without a JSON envelope. See Word extraction.

Deliberate non-optimizations

Two further size levers were measured and rejected — documented here so you know they were choices, not oversights:

  • One-letter type tags ("t" vs "text") save only 0.3% more while making the payload harder for models and humans to read.
  • Document-level font/color tables save ~13% more but break self-containment: a single page or element retrieved into a RAG context would no longer describe itself.