Wire processors into a pipeline.
Pick the processors you need. Compose them in the playground. Promote a versioned pipeline to production.
We are a research lab that trains document intelligence models, trusted by frontier AI labs, museums, and enterprises where accuracy is non-negotiable.








Each processor is rigorously benchmarked and designed for easy composition in your workflows.
Use Chandra to convert PDFs, spreadsheets, slides and more to structured markdown, html, and json. Includes word-level bounding boxes, redlines, and more.
P · 02 CustomizeCustomize processor output to your needs using natural language rules.
P · 03 ExtractFind key information using pre-defined extraction schemas.
P · 04 SegmentIdentify natural boundaries between documents that were originally combined into one file.
P · 05 EvalEvaluate the quality of output, using our criteria or your own rubrics for quality.
New primitives ship every quarter — see what's available now.
Every number on this page cites a public dataset or our own runnable eval.
One pipeline runtime — from a single API call to 100M pages a day.
Pick the processors you need. Compose them in the playground. Promote a versioned pipeline to production.
Managed cloud, in your VPC, or fully air-gapped. We own the models — so you choose the network.
Continuous evals against rubrics and a reference corpus — quality improves as we ship model updates, and any regression gets flagged.
Same engine — different documents. From training corpora to clinical trials to claims forms, every vertical gets the same bbox-precise parsing.
Parse web crawls, scientific papers, multilingual sources, and 1000+ page books into clean, citation-ready training data.
"Working with Datalab has helped us move faster and solve complex edge cases with our documents. We need highly accurate document processing, and Datalab has been a key part of our data infrastructure. We're excited to continue to partner with Datalab."
— S., Member of Technical Staff, Frontier Lab
Burst at scale on our GPUs. We allocate the workers, you pay per page processed. In production at the largest AI labs and Fortune 500 enterprises.
Marker, Surya, and Chandra are our open source models — used by thousands of teams, 67.7k+ stars combined. Our managed platform handles scaling infrastructure, includes better proprietary models, and has enterprise-grade reliability.
Because we own the models, you can run them anywhere — including networks that don't touch the internet.
Sign up, get an API key, ship in minutes. Free tier available.
Datalab runs inside your private cloud. Compliance, residency, control.
For research labs and regulated environments where data cannot leave the network.
Free tier available. No credit card required.