docray
X-ray for documents. docray takes PDF or PPTX input and returns JSON describing physical elements on every page or slide — text, images, vector paths, and annotations — each with a bounding box, font, and color information. PDF provides the full character → word → line hierarchy; PPTX provides native element-granularity slide structure.
docray extract report.pdf --granularity element
{
"type": "text",
"bbox": [72.0, 61.1, 143.4, 74.5],
"text": "Quarterly results",
"font": { "name": "Helvetica-Bold", "size": 18.0, "bold": true }
}
What it’s for
- LLM / RAG pipelines — token-efficient page content with coordinates for
click-to-source citations. Start with the
granularity guide; if you are token-conscious, you want
element. - Document viewers — overlay highlights and selections on the original page using the same coordinates the extraction reports.
- ML and data extraction — deterministic, lossless physical structure at
chargranularity for training data and downstream parsing.
Design principles
Lossless and unopinionated. At full granularity docray records what the PDF physically contains — it does not guess at headings, tables, or reading order. Consumers filter down; the extractor does not interpret.
Deterministic. The same input bytes always produce byte-identical JSON. The test suite enforces this with golden files and double-extraction checks.
No silent failures. Anything skipped, unsupported, or partially parsed is
recorded in a warnings array. Scanned (raster-only) pages are flagged
"scanned": true so you know exactly which pages need OCR.
Hardened against hostile input. Documents are untrusted. The HTTP server never parses them in-process — every document runs in an isolated worker subprocess with a wall-clock timeout, memory limit, and output cap. A malformed file that crashes the parser kills one job, never the service.
The pieces
| Piece | What it is |
|---|---|
docray | CLI: PDF/PPTX in, JSON or lean text on stdout |
docray-server | HTTP API: sync extraction + async job queue |
/playground | Browser workbench: pages beside their bounding boxes, live JSON |
docray-{model,core,pdf,pptx} | Rust crates: schema, extraction traits, and format extractors |