High-precision document intelligence.

We are a research lab that trains document intelligence models, trusted by frontier AI labs, museums, and enterprises where accuracy is non-negotiable.

TRUSTED BY
Sears Catalogue · 1902 — public domain scan 1 HDR2 H13 H14 TEXT5 H16 TEXT7 TEXT8 H19 TEXT10 TEXT11 H112 TEXT13 TEXT14 H115 H116 IMG17 TEXT18 TEXT19 H120 IMG21 TEXT22 TEXT23 TEXT24 TEXT25 H126 TEXT27 H128 TEXT29 TEXT30 H131 IMG32 TEXT33 TEXT34 TEXT35 TEXT36 H137 TEXT38 IMG39 IMG40 H141 TEXT42 TEXT43 TEXT44 TEXT45 H146 TEXT47 H148 TEXT49 H150 H151 TEXT52 TEXT
CHANDRA · PARSE Sears Catalogue · 1902 · 52 blocks
CONFIDENCE 0.0%
  • 1 Page Header Page Header 0.977
  • 2 Section Header Builders' Hardware 0.959
  • 3 Section Header THE GOODS WE LIST 0.984
  • 4 Text In this department are guaranteed for quality. They are made in proper proportions and are not only beautiful… 0.965
  • 5 Section Header THE FREIGHT RATES 0.980
  • 47 more regions detected
chandra · v1.4.2 · 1,287 ms · markdown · json · html
[04] PLATFORM · CORE

Compose. Deploy. Monitor.

One pipeline runtime — from a single API call to 100M pages a day.

01 COMPOSE

Wire processors into a pipeline.

Pick the processors you need. Compose them in the playground. Promote a versioned pipeline to production.

02 DEPLOY

Run our models anywhere.

Managed cloud, in your VPC, or fully air-gapped. We own the models — so you choose the network.

03 MONITOR

Monitor improvements and flag regressions.

Continuous evals against rubrics and a reference corpus — quality improves as we ship model updates, and any regression gets flagged.

[05] VERTICALS · USE CASES

Trusted in countless mission-critical workflows.

Same engine — different documents. From training corpora to clinical trials to claims forms, every vertical gets the same bbox-precise parsing.

Attention Is All You Need · arXiv — public domain scan 1 CAP2 FIG3 CAP4 FIG5 CAP6 TEXT7 TEXT8 EQN9 TEXT10 TEXT11 H112 TEXT13 TEXT14 FN15 FTR
CHANDRA · PARSE Attention Is All You Need · arXiv
15 blocks 0.0%

Training-grade corpora.

Parse web crawls, scientific papers, multilingual sources, and 1000+ page books into clean, citation-ready training data.

  • Multilingual web corpora
  • Scientific papers (PDF + LaTeX)
  • Long-document books (1000+ pp)
  • Math, charts, diagrams

"Working with Datalab has helped us move faster and solve complex edge cases with our documents. We need highly accurate document processing, and Datalab has been a key part of our data infrastructure. We're excited to continue to partner with Datalab."

— S., Member of Technical Staff, Frontier Lab
[06] MANAGED BATCH · NEW

100M+ pages a day. Zero infrastructure.

Burst at scale on our GPUs. We allocate the workers, you pay per page processed. In production at the largest AI labs and Fortune 500 enterprises.

[07] OPEN SOURCE · COMMUNITY

Open Source first.

Marker, Surya, and Chandra are our open source models — used by thousands of teams, 67.7k+ stars combined. Our managed platform handles scaling infrastructure, includes better proprietary models, and has enterprise-grade reliability.

[08] DEPLOYMENT · MODES

Run it on our infrastructure. Or yours.

Because we own the models, you can run them anywhere — including networks that don't touch the internet.

01

Managed Cloud

Sign up, get an API key, ship in minutes. Free tier available.

  • Auto-scaling
  • SOC 2 Type II
  • 99.99% uptime
  • Pay-as-you-go
Try Playground →
02

VPC Deployment

Datalab runs inside your private cloud. Compliance, residency, control.

  • AWS / GCP / Azure
  • Custom BAA / DPA
  • Sub-processor transparency
  • Dedicated support
Talk to sales →
03

On-prem · Air-gapped

For research labs and regulated environments where data cannot leave the network.

  • No internet required
  • Full model weights
  • Custom hardware support
  • White-glove deployment
Talk to sales →
[09] START

Start parsing.

Free tier available. No credit card required.