BLOG

Engineering, methodology, and what we ship.

Latest updates, tutorials, and insights from the Datalab team on document intelligence, OCR, and AI-powered document processing.

POST · 059

Form filling v2 - significantly more accurate

Form filling now runs on our document agent: geometry is measured rather than guessed, and every fill is checked against the page it produced. Building the benchmark to prove it also turned up three defects in the version we had been shipping, all of them silent.

PRODUCT UPDATES AUG 24 · 2026 6 MIN
POST · 058

Segmentation is now $0.50 per 1,000 pages

We cut the price of document segmentation by 92% — and for the most common use, by 95%.

PRODUCT UPDATES AUG 19 · 2026 3 MIN
POST · 057

Tagged, accessible PDFs in minutes

A new processor that turns a PDF — including a scan that is nothing but pixels — into a tagged PDF built for assistive technology, checked against a battery of automated tests and a real screen reader.

PRODUCT UPDATES AUG 19 · 2026 4 MIN
POST · 056

Datalab leads another competitor's extraction benchmark

The benchmark shipped with major scoring bugs. After fixing them, we score 93.6% on the benchmark, leading all other contenders. We encourage you to run your own evals, and not just trust vendor benchmarks.

PRODUCT UPDATES AUG 13 · 2026 5 MIN
POST · 055

From scanned papers to publisher-ready JATS XML

A new processor that turns raw scientific paper PDFs into DTD-validated JATS 1.2 XML, with every run carrying a conformance verdict you can audit.

PRODUCT UPDATES AUG 13 · 2026 4 MIN
POST · 054

ApparelWerks turns handwritten measurements into a system where every inch counts

How a made-to-measure clothing studio used Datalab to build a production system where every inch counts.

CASE STUDIES AUG 10 · 2026 4 MIN
POST · 053

Improved localization on our customers' hardest documents

Bounding boxes for bad scans, dense pages, and sideways text.

PRODUCT UPDATES AUG 4 · 2026 5 MIN
POST · 052

Improving High Accuracy Mode

High accuracy mode is now more accurate while making fewer changes to the page, powered by new verification models trained on-distribution with Chandra.

PRODUCT UPDATES JUL 24 · 2026 3 MIN
POST · 051

Marker 2: faster, CPU-ready, and more accurate

Marker 2 is a rewrite of our open-source PDF-to-markdown converter. It runs on CPU, picks a speed/accuracy mode automatically, and beats comparable pipeline OCR systems on both accuracy and throughput on olmOCR-bench.

PRODUCT UPDATES JUL 20 · 2026 8 MIN
POST · 050

When a citation needs to survive a courtroom

How Kin rebuilt the way adjusters cite insurance policies, achieving compliance-grade accuracy with Chandra and custom processors.

CASE STUDIES JUL 17 · 2026 4 MIN
POST · 049

We lead an external extraction benchmark (but run your own evals)

A competitor commissioned a long-table extraction benchmark, and scored Datalab at 34% recall. We rebuilt our extraction and now have the top recall (99.1%) and lead on precision too. Still, you should take all vendor benchmarks with a grain of salt and run your own evals.

PRODUCT UPDATES JUL 14 · 2026 5 MIN
POST · 048

Learnboost fixes the layer beneath the AI and lifts student retention

How an AI-powered learning platform boosted retention by replacing its document conversion with Marker.

CASE STUDIES JUL 8 · 2026 4 MIN
POST · 047

Tell us how you want your documents parsed

Custom Processors turn plain-language instructions and a few examples into a processor that parses your documents exactly the way you need.

PRODUCT UPDATES JUL 7 · 2026 4 MIN
POST · 046

How we achieve a near-zero OCR hallucination rate

Hallucinations are a fact of how LLMs work. For mission-critical documents, even one is too many. Here's the defense-in-depth system we use to drive them toward zero.

PRODUCT UPDATES JUN 30 · 2026 6 MIN
POST · 045

A box and a confidence score for every word

Making OCR output auditable.

PRODUCT UPDATES JUN 23 · 2026 4 MIN
POST · 044

Free to start, pay as you go

Datalab now has a free tier and pay-as-you-go pricing. Start parsing for free with a monthly usage allowance, then pay only for the pages you actually process — no subscription, no minimum, no plan to pick.

PRODUCT UPDATES JUN 18 · 2026 2 MIN
POST · 043

Introducing lift: open-weights structured extraction

lift is a 9B open-weights vision model that extracts structured JSON from PDFs and images by passing a schema. It scores 90.2% field accuracy on our 225-document extraction benchmark - the best small self-hostable model we've tested.

PRODUCT UPDATES JUN 18 · 2026 5 MIN
POST · 042

Chandra 2.1: Improved Multilingual and Table Accuracy

Announcing Chandra 2.1, a smaller, faster model that improves on multilingual and table accuracy.

PRODUCT UPDATES JUN 17 · 2026 4 MIN
POST · 041

Turbo Mode: Structured Extraction At 12s/doc

Turbo mode for the Datalab extraction API: structured JSON from any document at a median of ~12 seconds per document.

PRODUCT UPDATES JUN 11 · 2026 3 MIN
POST · 040

How Nevada County Historical Archive made 200K+ pages of Gold Rush history searchable with Datalab

How the Nevada County Historical Archive uses Datalab's Chandra model to transcribe 200K+ pages of handwritten Gold Rush-era records 150X faster, making a century of California history searchable.

CASE STUDIES JUN 5 · 2026 4 MIN
POST · 039

Introducing Balanced Mode for Structured Extraction

Balanced is our new default extraction mode: higher accuracy, with reasoning and independent verification baked into every response. It knows when a value isn't in the document instead of making one up.

PRODUCT UPDATES JUN 4 · 2026 6 MIN
POST · 038

EU Data Residency now available on Datalab

Datalab now offers data residency controls, starting with a new processing location in the Netherlands. Keep your document content and processing entirely within the EU, using the same API you already use.

PRODUCT UPDATES MAY 29 · 2026 2 MIN
POST · 037

Announcing Surya OCR 2: small, accurate, multilingual

Surya OCR 2 is a 650M-parameter open-source OCR model that scores 83.3% on olmOCR-bench, hits 87.2% on a 91-language multilingual eval, and runs on CPU, GPU, and MPS.

PRODUCT UPDATES MAY 27 · 2026 6 MIN
POST · 036

Managed Batch Processing is now Live

Introducing Datalab's Managed Batch Processing — give us a bucket, we handle the rest. Process millions of pages without managing GPUs, scaling infrastructure, or orchestrating workloads.

PRODUCT UPDATES APR 29 · 2026 5 MIN
POST · 035

How Radical AI Accelerates Materials Science Discovery with Datalab

Radical AI ($55M Seed+) ingests millions of pages of dense academic literature through Datalab to feed the autonomous-discovery systems shrinking materials R&D from decades to weeks.

CASE STUDIES MAR 26 · 2026 4 MIN
POST · 034

Announcing Chandra OCR 2: 90+ Languages, Top Benchmarks

Chandra 2 is a 4B parameter OCR model with state of the art benchmarks, layout blocks with bounding boxes, and structured output for diagrams and charts.

PRODUCT UPDATES MAR 18 · 2026 8 MIN
POST · 033

Chandra 1.5: Better Tables, Chemistry Support, and Faster Processing

Announcing Chandra 1.5 with dramatically improved table extraction, chemistry support with SMILES output, diagram rendering, and significant latency improvements.

PRODUCT UPDATES JAN 22 · 2026 5 MIN
POST · 032

How Rely trusts Datalab to process massive data rooms in hours, not days

Rely audits commercial real-estate transactions worth millions in misvalued assets. Running Datalab on-prem at 12.5 pages/sec lets them ingest 500,000-page data rooms before deals close — without resident documents ever leaving their network.

CASE STUDIES JAN 16 · 2026 3 MIN
POST · 031

How Cofactr (YC W'22) Automates Complex Logistics Workflows to Support Critical Hardware Supply Chains

Cofactr (YC W'22) serves regulated hardware manufacturers — Amazon Robotics among them. Datalab parses inconsistent BOMs, quotes, and spec sheets across 5,000+ suppliers with the fidelity their compliance and traceability work demands.

CASE STUDIES JAN 6 · 2026 4 MIN
POST · 030

Extracting Hyperlinks from PDFs

Learn how to extract hyperlinks from PDFs using the Datalab API, preserving both link text and destinations.

PRODUCT UPDATES DEC 19 · 2025 3 MIN
POST · 029

Introducing Form Filling: Automatically Fill PDF Forms with AI

Automatically fill PDF forms using AI with our new form filling feature.

PRODUCT UPDATES DEC 17 · 2025 3 MIN
POST · 028

Introducing: Forge Evals

A new tool to evaluate how various settings and modes impact document parsing quality across PDFs, DOCX files, and spreadsheets.

PRODUCT UPDATES DEC 10 · 2025 4 MIN
POST · 027

Datalab Benchmarks + Evals: View Scores & Compare Output On Your Documents

View scores and compare model output across a sample of diverse pages, or get evals on your own documents.

COMPANY UPDATES DEC 9 · 2025 3 MIN
POST · 026

Launch Week - Day 5: Faster Tracked Changes Outputs

Performance improvements for Tracked Changes output parsing, plus native spreadsheet support in Forge.

PRODUCT UPDATES DEC 5 · 2025 3 MIN
POST · 025

Launch Week - Day 4: Spreadsheet Parsing

Native spreadsheet support in the Datalab API for accurate table segmentation and parsing.

PRODUCT UPDATES DEC 4 · 2025 4 MIN
POST · 024

Launch Week - Day 3: Introducing Agni: Solving Multi-Page Section Hierarchy in OCR

Introducing Agni, a new model that maintains consistent section hierarchy across multi-page documents.

PRODUCT UPDATES DEC 3 · 2025 4 MIN
POST · 023

Launch Week - Day 2: Chandra is Faster (Again)

Announcing Chandra Small, a latency-optimized model that achieves 2-3x faster speeds with minimal performance degradation.

PRODUCT UPDATES DEC 2 · 2025 3 MIN
POST · 022

Launch Week - Day 1: Chandra 1.1

Announcing Chandra 1.1 with improved layout, math, tables, and multilingual performance.

PRODUCT UPDATES DEC 1 · 2025 3 MIN
POST · 021

Speeding Up Chandra

How we achieved 3x faster Chandra performance using Eagle3 speculative decoding without sacrificing accuracy.

PRODUCT UPDATES NOV 17 · 2025 5 MIN
POST · 020

Saturating the olmOCR Benchmark

Our analysis reveals fundamental limitations in the olmOCR benchmark and why traditional edit-distance benchmarks provide diminishing returns for OCR evaluation.

PRODUCT UPDATES NOV 13 · 2025 6 MIN
POST · 019

Extract Tracked Changes Metadata from Word Documents into Markdown & HTML

New feature enabling extraction of tracked changes metadata from Word documents for contract review workflows.

PRODUCT UPDATES NOV 12 · 2025 5 MIN
POST · 018

Grounded Intelligence: How High-Fidelity OCR Drives Accurate Structured Extraction

How superior OCR technology enhances the accuracy of structured data extraction from documents.

PRODUCT UPDATES NOV 4 · 2025 5 MIN
POST · 017

Introducing our newest model: Chandra

Meet Chandra, our latest OCR model that tops independent benchmarks and brings powerful new capabilities to document processing.

PRODUCT UPDATES OCT 30 · 2025 3 MIN
POST · 016

View Your API Requests in Our Playground

Visualize your API requests directly within the Playground to audit and debug your document processing.

PRODUCT UPDATES OCT 7 · 2025 3 MIN
POST · 015

Launch Week Day 5: Playground Examples

We've added practical examples to the Playground that you can explore directly.

PRODUCT UPDATES SEP 30 · 2025 2 MIN
POST · 014

Launch Week Day 4: High Accuracy Mode

Introducing High Accuracy Mode, which combines our proprietary models with frontier LLMs to handle challenging documents.

PRODUCT UPDATES SEP 26 · 2025 1 MIN
POST · 013

Launch Week Day 3: Layout Model Updates

Significant upgrades to our OCR and Layout models, achieving state-of-the-art performance on olmOCR-bench.

PRODUCT UPDATES SEP 25 · 2025 4 MIN
POST · 012

Launch Week Day 2: Document Segmentation

New Document Segmentation feature that automatically detects page boundaries within multi-document PDF files.

PRODUCT UPDATES SEP 24 · 2025 4 MIN
POST · 011

Launch Week Day 1: Datalab's New Playground

Unveiling our updated Playground with parsing, refinement, segmentation, and extraction capabilities.

PRODUCT UPDATES SEP 23 · 2025 1 MIN
POST · 010

Cracking Math OCR: How We're Unlocking High-Quality Data for Reasoning Models

How high-quality mathematical text extraction enables training of reasoning-capable language models.

PRODUCT UPDATES SEP 19 · 2025 3 MIN
POST · 009

Structured Extraction with Datalab API, and Handling Long Documents

How to use Datalab's Marker API for structured extraction from PDFs using JSON schemas, plus strategies for processing lengthy documents.

PRODUCT UPDATES SEP 8 · 2025 6 MIN
POST · 008

Citation Needed - Auditable Structured Extraction

Structured Extraction now includes citation support, enabling users to verify information sources within documents.

PRODUCT UPDATES AUG 19 · 2025 4 MIN
POST · 007

Parse PDFs *Just the Way You Want*

Introducing Marker Prompt API and Forge Parse for customizing PDF parsing outputs through prompts.

PRODUCT UPDATES AUG 14 · 2025 5 MIN
POST · 006

Introducing Forge Extract: Turn PDFs into Structured JSON

Extract structured data from PDFs using custom schemas with our new Forge Extract feature.

PRODUCT UPDATES AUG 7 · 2025 4 MIN
POST · 005

Purchaser.ai: Six-figure cost savings and 50% higher retention with Datalab

Manufacturing procurement runs on messy quotes and thousand-row BOMs. Datalab handles the inconsistency so Purchaser.ai's team ships faster purchasing decisions — and keeps the customers who'd otherwise churn on bad parses.

CASE STUDIES JUL 16 · 2025 4 MIN
POST · 004

The Datalab SDK: Transform Document Processing from Hours to Minutes

Announcing the official Datalab SDK for Python with simplified APIs and powerful features.

PRODUCT UPDATES JUL 9 · 2025 4 MIN
POST · 003

RevisionDojo (YC F24): Achieving 10x cost reduction and seamless scalability with Datalab

RevisionDojo's edtech platform serves 310K+ students who need study materials parsed on demand. Retiring their self-hosted screenshot pipeline for Datalab cut processing cost 10x and brought average parse time to ~10 seconds.

CASE STUDIES JUL 1 · 2025 4 MIN
POST · 002

Gamma: Unlocking a 40% Conversion Boost with reliable and accurate PDF conversion

Gamma's 50M users turn PDFs into working presentations. Switching parsers to Datalab lifted PDF-import conversion 40% — over a million PDFs a week now flow through the engine without dropping layout or table fidelity.

CASE STUDIES JUN 17 · 2025 2 MIN
POST · 001

Building better document intelligence for an AI-first world

Why highly accurate document intelligence matters more than ever in an AI-first world.

COMPANY UPDATES JUN 12 · 2025 6 MIN
START

Get started in minutes.

Free tier. No credit card. SOC 2 Type II.