BLOG · PRODUCT UPDATES

By Zach Nussbaum 3 mins

Launch Week - Day 2: Chandra is Faster (Again)

Announcing Chandra Small, a latency-optimized model that achieves 2-3x faster speeds with minimal performance degradation.

TL;DR; We trained a latency-optimzed small model that is 2-3x faster with minimal performance degradations.

Chandra Small

We recently shipped a few updates to Chandra and our API (Making Chandra 3x faster and Chandra 1.1). Now, we’re excited to announce Chandra Small, a latency-optimized model found exclusively in the Datalab API.

Shortly after our launch of Chandra, we trained and deployed Chandra Small. In our testing, Chandra Small is 2-3x faster than Chandra with minimal performance degradations. Additionally, we trained Chandra Small using QAT to enable quantization and further reduce latency.

olmOCR Benchmark Scores

Additionally, we found that we can reduce the number of tokens needed for many pages. This led to 30% latency reductions. With Chandra Small, users can expect 2-4 pages/s on an H100.

Try Out Chandra Small

Give Chandra Small a try in the API (via Fast mode).

We’re excited to continue to push the frontier to make Chandra as accurate and fast as possible! Reach out to [email protected] for more information or for access to a self-hosted version of Chandra Small!

START

Get started in minutes.

Free tier. No credit card. SOC 2 Type II.