BLOG · CASE STUDIES

By Datalab Team 4 mins

Learnboost fixes the layer beneath the AI and lifts student retention

How an AI-powered learning platform boosted retention by replacing its document conversion with Marker.

Lift in user retention
~6%
MoM growth absorbed
30-50%
Pages processed
5M+

Learnboost platform interface

Learnboost was growing quickly. Students were signing up, trying the platform, and getting hooked on the idea of uploading your notes and getting back a full set of study tools. Conversion was healthy. However, fewer students stuck around than founder, Leon Oxenfart, knew the product deserved.

The obvious suspect was the AI, but Leon discovered that the real culprit sat one layer underneath it - somewhere most teams never think to look.

A retention opportunity that looked like an AI problem

Learning has historically been a messy process. Materials live in handwritten notes, flashcards, lecture slides, and textbooks, scattered across formats. Learnboost pulls it all together: pick your sources, set a target grade, and the platform builds a personalized study plan to help you get there. After two years as a web app, it launched on iOS, with a Play Store release planned.

The growth was real, but Leon noticed students dropping off after they saw some outputs. For example, summaries with missing text and incorrectly ordered text would lead to inaccuracies in AI outputs. It affected retention by impacting the process of students reaching their learning goals.

The first instinct was to blame the model, but the team dug in instead.

The root cause, one layer down

Upon investigation, the team discovered that underneath the LLMs, document conversion was the bottleneck to their output quality. When the source material was converted incorrectly, every downstream study tool inherited the error.

Leon’s team went deep on the problem, testing simple Python packages and the open-source components of Unstructured. They found missing text, broken hierarchy, unrecognized titles, and bad image captions. When text was split into columns or the structure was ambiguous, the AI would hallucinate headings. Students would read a summary generated from the wrong structure, and notice it was off. Formula recognition was a particular weak point.

Why Datalab: quality and latency, proven through their own testing

Leon found Marker through Datalab’s open-source GitHub page. He didn’t need to schedule a demo or a sales conversation - he tested the Marker API directly himself, and only reached out after both quality and latency had passed his bar.

He made his decision based on two factors: parse quality and latency. Quality held up on his own documents, and Marker returned results in 5-15 seconds per job. This was essential for a consumer-facing product where the user journey relies on speed. AWS Textract was slower, and other paid services did not match Datalab on both dimensions simultaneously. With the evaluation done, Leon integrated Marker via API into Learnboost’s processing pipelines.

Holding up on the hardest documents

To test its limits, Learnboost ran their most complex documents using Marker: mixed images, tables, and text blocks; two-column academic papers with annotations; and scanned files that required OCR. Some files combined all of these at once.

A dense textbook page with Marker-parsed output

Animation of a complex textbook page being parsed by Marker into structured output

Hierarchies came back almost perfectly. Formula parsing, flagged early in the relationship, improved noticeably after Leon shared sample documents with the Datalab team.

“Better extraction quality improved the AI output, and the output was learning material the students care about.” — Leon Oxenfart, Founder, Learnboost

Impact: better accuracy, better retention, hands-free scale

Fixing the layer beneath the AI lifted the quality of every output above it, which in turn improved student retention. The same students who previously quietly churned now got flashcards and summaries that matched their material. To date, Learnboost has processed more than 5 million pages from PDF to Markdown with Marker.

The second benefit was scalability. Running conversion on their own GPUs would have meant fixed infrastructure costs; Marker’s API gives them a variable cost model that tracks their business instead. When Learnboost grew 30% one month and 50% the next, that surge was absorbed with zero engineering attention on conversion. Nobody had to think about it, which freed the team to focus on the rest of the product. As Leon puts it, having document parsing handled is a kind of safety moat.

“When you’re using Marker, you don’t have to care about text extraction for your PDFs.” — Leon Oxenfart

The clearest signal isn’t a metric. When other founders ask him about building anything around PDFs or document pipelines, he tells them to use Marker, because it takes one entire problem off of their plate. It’s a recommendation that he makes unprompted, which has started to appear as a recent uptick in interest from other German startups.

Built with support from Datalab’s startup program

Learnboost is part of Datalab’s startup program, which begins with discounted per-page rates and Team fees, and scales as a startup grows. For Leon, the early support meant that the money saved on infrastructure could go straight into growth. As Learnboost scales, the team is looking to explore deeper personalization and learning profiles that optimize how each individual studies. The layer beneath the AI, by now, is the one thing they’ve stopped worrying about.

START

Get started in minutes.

Free tier. No credit card. SOC 2 Type II.