Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The JSON contract

This page documents the lossless char-level contract (schema 1.1, the default when no granularity is requested). The compact shapes are documented in choosing a granularity.

DOCX/DOCM uses the separate schema 1.7 flow contract. Its envelope contains layout: "flow", optional approx_pages, and sections instead of pages:

{
  "granularity": "element", "schema_version": "1.7", "layout": "flow",
  "source": {"format": "docx", "sha256": "…", "size_bytes": 1234},
  "document": {"metadata": {"title": "…", "author": "…"}},
  "warnings": [], "approx_pages": null,
  "sections": [{
    "page_width": 612.0, "page_height": 792.0,
    "margins": {"top": 72.0, "right": 72.0, "bottom": 72.0, "left": 72.0},
    "headers": [], "footers": [], "blocks": []
  }]
}

Flow block types are paragraph, table, image, textbox, and break. Paragraphs contain stable block IDs, semantic roles, resolved runs, optional list labels and approximate-page hints, and authored breaks. Tables use authored column widths and merge-anchor cells, and can carry the approximate page where the table starts. Positioned tables, images, and textboxes carry tagged placement constraints, never resolved bounding boxes. See Word extraction for provenance and limits.

Coordinate system

Everything you need to place a box on a rendered page:

  • Origin top-left, y increases downward — what viewers and CV pipelines expect.
  • Units are PDF points (1/72 inch).
  • Coordinates are reported after page rotation — a 612×792 page with /Rotate 90 reports width: 792, height: 612 and boxes in that rotated, visible space. What you see is what the coordinates mean.
  • All values are rounded to 3 decimals; output is deterministic — byte-identical for identical input on a given platform and PDFium build. (Documents using non-embedded fonts pick up the platform’s substitute font metrics, so coordinates can differ by fractions of a point across operating systems.)
  • Bounding boxes are objects — {"x0", "y0", "x1", "y1"} — never bare arrays at this level.

Envelope

{
  "schema_version": "1.1",
  "source": { "format": "pdf", "sha256": "…", "size_bytes": 123456 },
  "document": { "page_count": 12, "metadata": { "title": "…", "author": "…" } },
  "warnings": [],
  "pages": [
    { "page_number": 1, "width": 612.0, "height": 792.0,
      "rotation": 0, "scanned": false, "elements": [] }
  ]
}

warnings is the no-silent-failure channel: skipped object kinds, per-page parse problems, geometry that couldn’t be read, and pages whose text is likely garbled because many glyphs are unmapped, replacement characters, or control characters. Empty means a fully clean extraction.

scanned is true when a page has no text elements and a single image covering ≥ 85% of the page area — the signal that a page’s text is not machine-readable and needs OCR to recover. It also flags pre-rendered (rasterized-slide) pages, which have the same property.

Opt-in page classification (schema 1.8)

An explicit classify request adds this object to every physical page:

"classification": {
  "kind": "mixed",
  "confidence": 0.82,
  "needs_ocr": false,
  "reasons": ["text_coverage=0.012", "text_elements=3", "image_coverage=0.248"]
}

kind is text, scanned, image, or mixed. The deterministic v1 rules use summed clipped glyph-box area divided by page area, mapped glyph and text element counts, summed and largest image coverage, vector-path density, and the existing page-scoped suspected_garbled_text warning:

  • mapped text is meaningful at 0.001 page coverage or eight glyphs;
  • an image is meaningful at 0.01 page coverage and participates in mixed at 0.20 coverage when meaningful text is also present;
  • little/no mapped text plus an image covering at least 0.85 is scanned;
  • meaningful mapped text without meaningful image coverage is text; other visual/no-text pages are image.

needs_ocr is true for scanned and image, a largest-image ratio of at least 0.60, text coverage below 0.001, or suspected garbled text. A garbled page therefore needs OCR even when its physical text operators make the kind text. confidence is a rounded 0..1 decision-margin score: distance from the thresholds used for the selected kind, not a statistically calibrated probability. reasons records the rounded signals in stable evaluation order.

Classification does not read or mutate the legacy scanned field. Omitting the option serializes the original schema 1.1 model directly, so no classification field or version change can leak into default bytes.

Granularity-shaped pages (schema 1.6 for explicit char, 1.9 for element/word) can also carry a hidden array. The field is omitted when empty and is copied unchanged across granularities:

"hidden": [
  { "kind": "role", "element": "p1-e0", "content": "title" },
  { "kind": "notes", "content": "Presenter script" }
]

element is omitted for page-targeted items. Hidden content is supplemental, non-visible document context and must not be treated as text rendered on the page. The kind namespace is stable:

KindTargetPPTXPDF
roleelementPlaceholder type, defaulting to bodynot emitted
notespageSpeaker-notes body textnot emitted
altelementShape/picture alternative textnot emitted
hidden-slidepagetrue for a slide with show="0"not emitted
source-layerelementmaster or layout for inherited visible shapesnot emitted
fieldblockDOCX field instructionnot emitted
commentblockDOCX comment bodynot emitted
tracked-insertblockDOCX accepted insertionnot emitted
tracked-deleteblockDOCX rejected deletionnot emitted
footnoteblockDOCX note linked to its referencenot emitted

Elements

One element per native PDF page object, in z-order, discriminated by "type". IDs are stable within a response: p{page}-e{index}. Content inside Form XObjects (containers PowerPoint exports wrap everything in) is recursively extracted and flattened into the page’s element stream with correct page-space coordinates.

text

{
  "id": "p1-e4", "type": "text",
  "bbox": {"x0": 294.9, "y0": 48.0, "x1": 300.3, "y1": 61.2},
  "content": "Introduction to parsing",
  "font": { "name": "NimbusRomNo9L-Regu", "size": 10.909, "bold": false, "italic": false },
  "color": { "fill": [0, 0, 0], "stroke": null },
  "lines": [
    { "bbox": {}, "baseline_y": 61.2,
      "words": [
        { "content": "Introduction", "bbox": {},
          "chars": [ { "content": "I", "bbox": {}, "unicode": 73 } ] }
      ] }
  ]
}

One text element per native text run. Lines and words are grouped geometrically and deterministically; whitespace characters separate words and are not emitted as chars. Word order is content-stream order — reading order is not inferred.

Granularity-shaped text elements (schema 1.6 for explicit char, 1.9 for element/word) can additionally carry runs. Each run preserves its own content, resolved font, color, and optional external hyperlink target:

"runs": [
  {
    "content": "linked text",
    "font": { "name": "Aptos", "size": 18.0, "bold": true },
    "color": { "fill": [31, 78, 121] },
    "href": "https://example.com"
  }
]

href is omitted for an unlinked run. Element/word compact output applies the same compact font and color rules to runs as to their parent: false emphasis flags, black fill, and empty color objects are omitted. PDF text has no separate native shape run layer and omits runs, preserving the frozen no-parameter schema 1.1 bytes. PPTX text keeps content, font, and color as the concatenated and dominant summary while using runs for the per-run detail.

table

Granularity-shaped output (schema 1.6 for explicit char, 1.9 for element/word) carries first-class table elements for PPTX:

{
  "id": "p1-e2", "type": "table",
  "bbox": {"x0": 72.0, "y0": 90.0, "x1": 272.0, "y1": 170.0},
  "rows": 2, "cols": 2,
  "cells": [
    {
      "bbox": {"x0": 72.0, "y0": 90.0, "x1": 272.0, "y1": 120.0},
      "row": 0, "col": 0, "row_span": 1, "col_span": 2,
      "content": "Merged heading",
      "runs": []
    }
  ]
}

rows and cols are the source grid dimensions. Only merge-anchor cells are emitted; continuation cells are omitted and the anchor carries the clamped span and merged bounding box. Cell paragraphs are joined with \n, and cell runs use the same shape as text-element runs. PDF emits no table elements.

chart

Granularity-shaped output (schema 1.6 for explicit char, 1.9 for element/word) carries first-class chart elements for PPTX:

{
  "id": "p1-e3", "type": "chart",
  "bbox": {"x0": 72.0, "y0": 72.0, "x1": 432.0, "y1": 288.0},
  "chart_type": "doughnut",
  "title": "Channel mix",
  "series": [
    {
      "name": "Share",
      "points": [
        {"category": "Direct", "value": "41%"},
        {"category": "Reseller", "value": "59%"}
      ]
    }
  ]
}

chart_type is bar, pie, doughnut, line, area, scatter, or other, derived from the chart node in plotArea. Combo charts use the first chart node in document order while retaining every series in document order. The optional chart title and series name are omitted when absent. Points pair categories and finite values by their source index; unmatched values remain as points without category. Values are strings formatted with the series’ OOXML formatCode, so a stored 0.41 displayed as 0% is returned as "41%". PDF emits no chart elements.

image

Bounding box plus a quad (four corner points — meaningful when the image is placed with rotation or skew), pixel dimensions, colorspace, and a content_hash (sha256 of the raw image data) for deduplication. Pixel data itself is never embedded.

path

Bounding box plus paint: fill/stroke colors and stroke width. Path operator lists are not included. Schema 1.9 compact element/word paths retain the same optional fill, stroke, and stroke_width fields while rounding the bbox and stroke width to one decimal. An absent paint field is omitted; compact images remain bbox-only.

annotation

Subtype (link, highlight, widget, …), bounding box, and uri for links.

Stability

SchemaSelected by
1.1PDF JSON with no shaping parameters
1.6Explicit char granularity for paged output
1.7DOCX/DOCM authored-flow output
1.8Opt-in classified paged output, with or without granularity
1.9Explicit element or word granularity for paged output
  • The no-parameter response is frozen at schema 1.1 — new fields are only ever additive. Explicit char granularity carries schema 1.6; explicit element/word granularity carries schema 1.9, both with a granularity discriminator. element/word additionally detect glyph-fragmented pages (one text object per glyph, with no batched runs) and regroup them into geometric lines and words before projecting — char always reports the original per-glyph geometry, unregrouped. Flow responses use schema 1.7. Opt-in classified paged responses use schema 1.8 and may also carry the requested granularity discriminator. PDF emits no hidden items, runs, tables, or charts, so its no-parameter 1.1 bytes remain unchanged.
  • Element IDs, field names, and the coordinate system are load-bearing contract; they do not change within a major schema version. Hidden kind strings are equally stable and are never renamed.