Processors| Processor | What it does | Rate (per 1,000 pages) |
|---|
| Convert — fast / balanced | Document to markdown or HTML | 4 |
| Convert — accurate | Highest-fidelity conversion | 10 |
| Segment — page level | Detect boundaries between combined documents at the page level. Returns the page ranges and titles only — no OCR text | 0.5 |
| Segment — block level | Boundaries can fall mid-page at specific content blocks, driven by an optional custom prompt ("start a new segment at each invoice"). Returns the parsed markdown, so it runs Convert too | 4.5 |
| Extraction — fast | Structured fields, single pass (up to 50 pages per document) | 6 |
| Extraction — balanced | Structured fields, multi-pass with per-field verification | 15 + fees* |
| Extraction — accurate | Structured fields, multi-pass with per-field verification, strongest model | 20 + fees* |
| Custom processor | Your own pipeline | 20 |
| Eval | Score output quality against a rubric | 2 |
| Form fill | Populate a form's fields | 10‡ |
| Create document | Generate a document | 6 |
Add-ons| Add-on | What it does | Rate (per 1,000 pages) |
|---|
| Track changes | Diff revisions in a Word document | 6 |
| Chart understanding (add-on) | Adds chart parsing to a run | +3 |
| Infographic (add-on) | Adds infographic parsing to a run | +4 |
| Word bounding boxes (add-on) | A box and confidence score for every word | +3 |
| Cross-page merging (beta) | Stitch content split across a page break back together: tables (long or wide), paragraphs broken mid-sentence, and lists | Variable** |
| Word bounding boxes | Identify bounding boxes at the word level with confidence scores so you can efficiently route OCR issues for human review and audit mistakes in your pipeline | +3 |
| Table cell bounding boxes (includes word bboxes) | Identify granular bounding boxes for each table cell (instead of the whole table). Also includes word level bounding boxes. | +6 |
| List bounding boxes (includes word bboxes) | Identify granular bounding boxes for each item in a list or list group (instead of the entire list as one block). Also includes word level bounding boxes. | +6 |
| Word, table cell, and list bounding boxes | Get all 3 granular bounding boxes at the word, table cell, and list level. | +9 |
* Around 5% of Balanced and Accurate extractions bill above our fixed
per-page rates. That happens when the extraction agent has to work through
a document — long documents, and schemas that pull many repeated rows, are
the usual cases; the rate covers the rest, and you are never charged both
the rate and the full compute. When it does apply it is typically a dollar
or two, more for a long or unusually dense document (Accurate runs a
stronger model, so it both includes and can bill more). Schemas over 750
fields aren't supported.
** Cross-page merging is in beta. It has no per-page rate — it bills only
a variable compute surcharge reflecting the merge work your document
actually needs, which is typically a few cents per document.
‡ Form filling runs an agent over your form, so its cost follows the work
the form needs rather than its page count. Almost every form bills this
rate and nothing more; an unusually dense one bills what the work actually
cost instead. The rate is a floor, not something a surcharge adds to, so
you are never charged twice for the same fill.
Spreadsheets bill by cells, not pages — 2,500 cells per page capped at
$0.60 per sheet (simple) or 500 cells per page (intelligent), chosen
automatically. Processing in the EU adds 25% to the summed total. A single
discount can also apply — 25% for opting into data retention, or 33% under
the startup program — but the two discounts can't both apply.