# Datalab

Document infrastructure for workflow automation

- Team: Vik Paruchuri (co-founder)
- Founded: 2024
- Invested: 2024
- Links: [Website](https://www.datalab.to), [LinkedIn](https://www.linkedin.com/in/vikparuchuri), [X](https://x.com/VikParuchuri), [GitHub](https://github.com/VikParuchuri)
- Field: How software gets made

## The problem: How do you get a computer to read a document the way people do?

### What a PDF actually holds

A PDF is a complete description of a fixed-layout document: the text, the fonts, the graphics and the images needed to show it. The text is stored as instructions to draw characters at certain positions, and each character is a code that maps to a glyph in a font. It is a printout in digital form. It records where the ink goes, and says very little about what the words are doing there.

A scanned page doesn't even have that much. It is a picture, and getting words out of it takes optical character recognition, or OCR: turning images of typed, handwritten or printed text into machine-encoded text. Along the way an OCR system usually does layout analysis, picking out columns, paragraphs and captions, and then finds the lines and words.

Layout analysis is where a page stops being ink and becomes a document. It separates text from everything else, puts the text zones in their correct reading order, and labels the parts that play different roles, such as titles, captions and footnotes.

### Why it is hard

**Built to look right.** PDF's whole point is keeping a document looking the same everywhere, and that very emphasis makes it hard to convert or to pull out text, images, tables and metadata. There is a variant, tagged PDF, that carries structure information precisely so that text can be extracted reliably. Its existence tells you what the ordinary kind is missing.

**Which column next?.** People read a two-column page without thinking. Software has to work it out. Top-down methods cut the page into columns and blocks using white space, which is fast but only works if they make assumptions about the layout. Bottom-up methods build words into lines into blocks from the raw pixels, which assumes nothing and takes time. Both have to cope with noise and with pages scanned slightly crooked.

**Tables fight back.** Tables come in every shape and size, with complicated header arrangements, rows that run over several lines, all sorts of separator lines, and blank cells. Recovering which number belongs to which row and column from an image is, in the dry words of one paper, a non-trivial task. Get it slightly wrong and the numbers are all there, in the wrong places.

**Good, cheap, pick one.** Some content stays hard for even the best tools: formulas, tiny fonts, old scans. Vision language models tend to beat traditional open-source tools on quality. The best of them can also be prohibitively costly, over 6,240 USD per million pages for GPT-4o, and some documents can't be sent to someone else's API at all.

### What Datalab is after

Datalab is a research lab that trains document intelligence models. Its aim is to turn messy, unstructured content into precise data that production systems can depend on, from clean training data for AI models to the documents a business runs on.

The jobs break down into a few pieces: convert a document into structured text, extract the key facts, split apart files that were originally several documents, and check that the output is any good.

### How they go at it

**Built in the open.** Marker, Surya and Chandra are Datalab's own open-source models. Marker converts PDFs, images, slides, Word files and more into markdown, JSON and HTML. It formats tables, forms, equations and code blocks, and strips out running headers and footers.

**Layout, then letters.** Surya is the OCR model underneath, with 650 million parameters. Besides reading text, it does layout analysis with reading order and recognises the rows and columns of tables, and together with Marker it covers more than 90 languages.

**A model for messy pages.** Chandra, the flagship model, converts images and PDFs into structured HTML, markdown or JSON while keeping the layout information. It is built for the messiest documents: complex tables, forms and handwriting.

**Pipelines that check themselves.** On the managed platform these pieces are wired together into versioned pipelines, with continuous evaluations against rubrics and a reference corpus so that any regression gets flagged. Because Datalab owns its models, they can also run inside a customer's own cloud or on machines with no internet at all.

### Still open

Does every page need a giant model? Maybe not. One open-source team fine-tuned a vision language model of 7 billion parameters to turn PDFs into clean text in natural reading order, and it converts a million pages for only 176 USD, against thousands of dollars for the biggest commercial models.

How do you know a parse is right? Mostly by testing against labelled datasets, and several exist for PDF conversion and extraction. One recent benchmark gathers 1,400 PDFs full of the content types that still trip up the best tools, which is a polite way of saying the problem is not solved.

### Words used here

- **glyph**: The drawn shape of a character in a particular font.
- **OCR**: Optical character recognition: software that turns pictures of text into text a computer can search and edit.
- **layout analysis**: Working out which parts of a page are text, images or tables, what role each plays, and how they fit together.
- **reading order**: The sequence a person would read a page's blocks in, which may not match their position top to bottom.
- **tagged PDF**: A PDF that also stores the document's structure, such as headings and paragraphs, alongside how it looks.
- **Vision language models**: AI models that take images as well as text as input, so they can read a page by looking at it.
- **markdown**: A simple plain-text format for headings, lists and tables that both people and language models read easily.

## About Datalab

Datalab is building document infrastructure that allows companies to understand and automate workflows built on human-readable documents. The platform learns the structure of a company's documents and the workflows that operate on top of them, creating a data layer that can be reliably used by AI systems.

Most enterprise knowledge lives in documents — contracts, reports, PDFs, internal wikis — that are readable by humans but poorly understood by machines. Datalab goes beyond OCR or point extraction, focusing instead on modeling document structure and workflow context so AI can be applied in a repeatable, production-grade way.

The platform is used by developers, enterprises, and AI labs that need reliable document understanding as a foundation for automation.

## Sources

1. [PDF](https://en.wikipedia.org/wiki/PDF), Wikipedia
2. [Optical character recognition](https://en.wikipedia.org/wiki/Optical_character_recognition), Wikipedia
3. [Document layout analysis](https://en.wikipedia.org/wiki/Document_layout_analysis), Wikipedia
4. [TableFormer: Table Structure Understanding with Transformers](https://arxiv.org/abs/2203.01017), arXiv
5. [olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models](https://arxiv.org/abs/2502.18443), arXiv
6. [High-Precision Document Intelligence](https://www.datalab.to/), Datalab
7. [marker](https://github.com/datalab-to/marker), Datalab (GitHub)
8. [surya](https://github.com/datalab-to/surya), Datalab (GitHub)
9. [chandra](https://github.com/datalab-to/chandra), Datalab (GitHub)
