Radical AI is a New York-based materials science startup using AI-powered autonomous discovery to accelerate the development of next-generation materials. While many companies in this space are focused on simulation alone, Radical AI is taking it further — enabling AI systems to interface with the physical world. By combining simulation with a working materials science lab, the team is compressing experimental cycles that have traditionally taken decades into weeks and creating a new paradigm for how discovery and manufacturing can come together.
The company is currently hyperscaling its team after a $55 million Seed+ round (one of the largest in NY history), and is working with some of the most demanding research environments in the field, tackling everything from advanced alloys to novel materials for semiconductors. With 40+ people spanning AI researchers, materials scientists, physicists, mechatronics engineers, and automation specialists, Radical operates as multiple specialized teams within a single ambitious mission.
The Challenge: Unlocking Scientific Literature at Scale
At the core of Radical’s platform is a knowledge base built from scientific literature — academic papers, research publications, technical reports, and experimental datasets — that underpins the AI systems guiding materials discovery and experimental design.
Building and maintaining this knowledge base created a document processing challenge with several dimensions:
- Complex scientific layouts: Materials science papers contain dense tables, chemical notation, XRD diagrams, and microscopy images that break standard PDF parsers and general-purpose OCR tools.
- Domain ontology complexity: Scientific literature uses highly specialized taxonomies (alloy designations, crystal structures, processing conditions) that require layout-preserving extraction to remain queryable.
- Visual asset extraction for model training: SEM images, XRD diagrams, and other experimental figures aren’t just illustrations — they’re training data for analytical VLMs and need to be extracted, bounded, and stored as discrete assets linked to source context.
- Composition and entity extraction: Materials science requires parsing structured data like chemical compositions, synthesis parameters, and performance metrics, not just free text. Standard parsers collapse these into noise.
- Extreme volume: An initial backfill of ~5 million pages needed to be processed, indexed, and made available for retrieval in a compressed timeframe.
- Ongoing ingestion: Each time Radical acquires a new data source, hundreds of thousands of document pages need to be processed rapidly. The workload is bursty by nature.
Scientific literature ingestion is a key part of how Radical outperforms work once done by hundreds of PhD-level specialists. Before a single experiment runs, before a materials hypothesis is tested in the lab, Radical’s agentic systems need to retrieve and synthesize thousands of relevant papers. The quality of that ingestion layer determines the quality of all downstream inference: which experiments get prioritized, which material compositions get explored, and ultimately how fast performance goals get hit.
That’s why document parsing isn’t a peripheral concern for Radical. It’s foundational infrastructure.
The Solution: Datalab as the Document Intelligence Layer
Radical integrated Datalab’s Chandra model via API to power the ingestion pipeline at the core of its knowledge base. The integration happened within days — the team had already tested Datalab’s API and validated extraction quality on complex academic content before the first commercial conversation began.
“We were quite blown away by the quality of extraction, especially on the type of complex academic papers that we were passing in. We’re very excited to start using the product at scale.” — Rohil Kulkarni, Product Lead
Radical’s pipeline uses Datalab to convert raw scientific PDFs into structured markdown with full layout understanding, feeding directly into indexing and retrieval systems. This enables semantic and keyword search over extracted text while storing charts, figures, and scientific images — XRD diagrams, microscopy images, and more — as discrete assets linked to source documents.
Key reasons Radical chose Datalab over alternatives:
- Layout fidelity for scientific entity extraction: Reliable composition parsing and domain-specific retrieval depends on structural fidelity — tables staying intact, multi-column layouts preserved, figure boundaries correctly identified. Downstream extraction pipelines break when parsing quality degrades.
- Speed at scale: At 2-3 pages per second, Datalab processed Radical’s 5 million-page backfill within the target window and handles ongoing batch jobs without throughput constraints.
- API flexibility: Given Radical’s bursty processing pattern, the API offered the right balance of capacity and cost efficiency without infrastructure overhead.
- Data privacy: With no data retained beyond a temporary processing cache, Radical could feed sensitive research data through the pipeline without compliance concerns.
Key Outcomes
- 16 weeks to hit materials performance goals on a recent campaign, a process that typically requires decades and hundreds of PhD-level specialists
- 5 million+ pages indexed in days
- 2-3 pages/sec throughput via API, enabling large batch jobs to complete in hours
Impact: Speed and Accuracy as a Competitive Advantage
By offloading document parsing to Datalab, Radical’s engineering team focused on the business logic that drives discovery — how to represent materials knowledge, index and retrieve it effectively, and use it to accelerate experimental cycles. Discovery processes that in traditional R&D required multi-year timelines and large specialist teams can now be compressed to weeks.
For Radical, Datalab didn’t just solve a data engineering problem. It became the foundation of a scientific retrieval architecture that goes well beyond document search — enabling composition-aware querying, domain ontology alignment, and visual model training at a scale and quality that wouldn’t be possible without high-fidelity extraction upstream.
Looking Forward
As Radical expands its corpus and explores new materials domains, Datalab will remain the ingestion backbone. With chart understanding capabilities in active development, Radical also sees potential for richer extraction of experimental visualizations — unlocking additional value for materials scientists working with complex figures and diagrams.