Engineering, methodology, and what we ship.
Latest updates, tutorials, and insights from the Datalab team on document intelligence, OCR, and AI-powered document processing.
Form filling v2 - significantly more accurate
Form filling now runs on our document agent: geometry is measured rather than guessed, and every fill is checked against the page it produced. Building the benchmark to prove it also turned up three defects in the version we had been shipping, all of them silent.
Segmentation is now $0.50 per 1,000 pages
We cut the price of document segmentation by 92% — and for the most common use, by 95%.
Tagged, accessible PDFs in minutes
A new processor that turns a PDF — including a scan that is nothing but pixels — into a tagged PDF built for assistive technology, checked against a battery of automated tests and a real screen reader.
Datalab leads another competitor's extraction benchmark
The benchmark shipped with major scoring bugs. After fixing them, we score 93.6% on the benchmark, leading all other contenders. We encourage you to run your own evals, and not just trust vendor benchmarks.
From scanned papers to publisher-ready JATS XML
A new processor that turns raw scientific paper PDFs into DTD-validated JATS 1.2 XML, with every run carrying a conformance verdict you can audit.
ApparelWerks turns handwritten measurements into a system where every inch counts
How a made-to-measure clothing studio used Datalab to build a production system where every inch counts.
Improved localization on our customers' hardest documents
Bounding boxes for bad scans, dense pages, and sideways text.
Improving High Accuracy Mode
High accuracy mode is now more accurate while making fewer changes to the page, powered by new verification models trained on-distribution with Chandra.
Marker 2: faster, CPU-ready, and more accurate
Marker 2 is a rewrite of our open-source PDF-to-markdown converter. It runs on CPU, picks a speed/accuracy mode automatically, and beats comparable pipeline OCR systems on both accuracy and throughput on olmOCR-bench.
When a citation needs to survive a courtroom
How Kin rebuilt the way adjusters cite insurance policies, achieving compliance-grade accuracy with Chandra and custom processors.
We lead an external extraction benchmark (but run your own evals)
A competitor commissioned a long-table extraction benchmark, and scored Datalab at 34% recall. We rebuilt our extraction and now have the top recall (99.1%) and lead on precision too. Still, you should take all vendor benchmarks with a grain of salt and run your own evals.
Learnboost fixes the layer beneath the AI and lifts student retention
How an AI-powered learning platform boosted retention by replacing its document conversion with Marker.
Tell us how you want your documents parsed
Custom Processors turn plain-language instructions and a few examples into a processor that parses your documents exactly the way you need.
How we achieve a near-zero OCR hallucination rate
Hallucinations are a fact of how LLMs work. For mission-critical documents, even one is too many. Here's the defense-in-depth system we use to drive them toward zero.
A box and a confidence score for every word
Making OCR output auditable.
Free to start, pay as you go
Datalab now has a free tier and pay-as-you-go pricing. Start parsing for free with a monthly usage allowance, then pay only for the pages you actually process — no subscription, no minimum, no plan to pick.
Introducing lift: open-weights structured extraction
lift is a 9B open-weights vision model that extracts structured JSON from PDFs and images by passing a schema. It scores 90.2% field accuracy on our 225-document extraction benchmark - the best small self-hostable model we've tested.
Chandra 2.1: Improved Multilingual and Table Accuracy
Announcing Chandra 2.1, a smaller, faster model that improves on multilingual and table accuracy.
Turbo Mode: Structured Extraction At 12s/doc
Turbo mode for the Datalab extraction API: structured JSON from any document at a median of ~12 seconds per document.
How Nevada County Historical Archive made 200K+ pages of Gold Rush history searchable with Datalab
How the Nevada County Historical Archive uses Datalab's Chandra model to transcribe 200K+ pages of handwritten Gold Rush-era records 150X faster, making a century of California history searchable.
Introducing Balanced Mode for Structured Extraction
Balanced is our new default extraction mode: higher accuracy, with reasoning and independent verification baked into every response. It knows when a value isn't in the document instead of making one up.
EU Data Residency now available on Datalab
Datalab now offers data residency controls, starting with a new processing location in the Netherlands. Keep your document content and processing entirely within the EU, using the same API you already use.
Announcing Surya OCR 2: small, accurate, multilingual
Surya OCR 2 is a 650M-parameter open-source OCR model that scores 83.3% on olmOCR-bench, hits 87.2% on a 91-language multilingual eval, and runs on CPU, GPU, and MPS.
Managed Batch Processing is now Live
Introducing Datalab's Managed Batch Processing — give us a bucket, we handle the rest. Process millions of pages without managing GPUs, scaling infrastructure, or orchestrating workloads.
How Radical AI Accelerates Materials Science Discovery with Datalab
Radical AI ($55M Seed+) ingests millions of pages of dense academic literature through Datalab to feed the autonomous-discovery systems shrinking materials R&D from decades to weeks.
Announcing Chandra OCR 2: 90+ Languages, Top Benchmarks
Chandra 2 is a 4B parameter OCR model with state of the art benchmarks, layout blocks with bounding boxes, and structured output for diagrams and charts.
Chandra 1.5: Better Tables, Chemistry Support, and Faster Processing
Announcing Chandra 1.5 with dramatically improved table extraction, chemistry support with SMILES output, diagram rendering, and significant latency improvements.
How Rely trusts Datalab to process massive data rooms in hours, not days
Rely audits commercial real-estate transactions worth millions in misvalued assets. Running Datalab on-prem at 12.5 pages/sec lets them ingest 500,000-page data rooms before deals close — without resident documents ever leaving their network.
How Cofactr (YC W'22) Automates Complex Logistics Workflows to Support Critical Hardware Supply Chains
Cofactr (YC W'22) serves regulated hardware manufacturers — Amazon Robotics among them. Datalab parses inconsistent BOMs, quotes, and spec sheets across 5,000+ suppliers with the fidelity their compliance and traceability work demands.
Extracting Hyperlinks from PDFs
Learn how to extract hyperlinks from PDFs using the Datalab API, preserving both link text and destinations.
Introducing Form Filling: Automatically Fill PDF Forms with AI
Automatically fill PDF forms using AI with our new form filling feature.
Introducing: Forge Evals
A new tool to evaluate how various settings and modes impact document parsing quality across PDFs, DOCX files, and spreadsheets.
Datalab Benchmarks + Evals: View Scores & Compare Output On Your Documents
View scores and compare model output across a sample of diverse pages, or get evals on your own documents.
Launch Week - Day 5: Faster Tracked Changes Outputs
Performance improvements for Tracked Changes output parsing, plus native spreadsheet support in Forge.
Launch Week - Day 4: Spreadsheet Parsing
Native spreadsheet support in the Datalab API for accurate table segmentation and parsing.
Launch Week - Day 3: Introducing Agni: Solving Multi-Page Section Hierarchy in OCR
Introducing Agni, a new model that maintains consistent section hierarchy across multi-page documents.
Launch Week - Day 2: Chandra is Faster (Again)
Announcing Chandra Small, a latency-optimized model that achieves 2-3x faster speeds with minimal performance degradation.
Launch Week - Day 1: Chandra 1.1
Announcing Chandra 1.1 with improved layout, math, tables, and multilingual performance.
Speeding Up Chandra
How we achieved 3x faster Chandra performance using Eagle3 speculative decoding without sacrificing accuracy.
Saturating the olmOCR Benchmark
Our analysis reveals fundamental limitations in the olmOCR benchmark and why traditional edit-distance benchmarks provide diminishing returns for OCR evaluation.
Extract Tracked Changes Metadata from Word Documents into Markdown & HTML
New feature enabling extraction of tracked changes metadata from Word documents for contract review workflows.
Grounded Intelligence: How High-Fidelity OCR Drives Accurate Structured Extraction
How superior OCR technology enhances the accuracy of structured data extraction from documents.
Introducing our newest model: Chandra
Meet Chandra, our latest OCR model that tops independent benchmarks and brings powerful new capabilities to document processing.
View Your API Requests in Our Playground
Visualize your API requests directly within the Playground to audit and debug your document processing.
Launch Week Day 5: Playground Examples
We've added practical examples to the Playground that you can explore directly.
Launch Week Day 4: High Accuracy Mode
Introducing High Accuracy Mode, which combines our proprietary models with frontier LLMs to handle challenging documents.
Launch Week Day 3: Layout Model Updates
Significant upgrades to our OCR and Layout models, achieving state-of-the-art performance on olmOCR-bench.
Launch Week Day 2: Document Segmentation
New Document Segmentation feature that automatically detects page boundaries within multi-document PDF files.
Launch Week Day 1: Datalab's New Playground
Unveiling our updated Playground with parsing, refinement, segmentation, and extraction capabilities.
Cracking Math OCR: How We're Unlocking High-Quality Data for Reasoning Models
How high-quality mathematical text extraction enables training of reasoning-capable language models.
Structured Extraction with Datalab API, and Handling Long Documents
How to use Datalab's Marker API for structured extraction from PDFs using JSON schemas, plus strategies for processing lengthy documents.
Citation Needed - Auditable Structured Extraction
Structured Extraction now includes citation support, enabling users to verify information sources within documents.
Parse PDFs *Just the Way You Want*
Introducing Marker Prompt API and Forge Parse for customizing PDF parsing outputs through prompts.
Introducing Forge Extract: Turn PDFs into Structured JSON
Extract structured data from PDFs using custom schemas with our new Forge Extract feature.
Purchaser.ai: Six-figure cost savings and 50% higher retention with Datalab
Manufacturing procurement runs on messy quotes and thousand-row BOMs. Datalab handles the inconsistency so Purchaser.ai's team ships faster purchasing decisions — and keeps the customers who'd otherwise churn on bad parses.
The Datalab SDK: Transform Document Processing from Hours to Minutes
Announcing the official Datalab SDK for Python with simplified APIs and powerful features.
RevisionDojo (YC F24): Achieving 10x cost reduction and seamless scalability with Datalab
RevisionDojo's edtech platform serves 310K+ students who need study materials parsed on demand. Retiring their self-hosted screenshot pipeline for Datalab cut processing cost 10x and brought average parse time to ~10 seconds.
Gamma: Unlocking a 40% Conversion Boost with reliable and accurate PDF conversion
Gamma's 50M users turn PDFs into working presentations. Switching parsers to Datalab lifted PDF-import conversion 40% — over a million PDFs a week now flow through the engine without dropping layout or table fidelity.
Building better document intelligence for an AI-first world
Why highly accurate document intelligence matters more than ever in an AI-first world.