← PortfolioHow software gets made

Datalab

Document infrastructure for workflow automation

Team
Vik Paruchurico-founder
Founded
2024
Invested
2024
Links
The problem

How do you get a computer to read a document the way people do?

Point at a block on the page or a line of the output to see where it went. Hold the button down to read the page straight across instead, the way a naive program would.A two-column page as a PDF or a scan holds it: marks at positions, with little to say what each part is doing. Layout analysis splits it into blocks, labels what each one is, and works out the order a person would read them in. The text on the right is written out in that order. Read straight across instead and the two columns come out interleaved.An illustration, not real data.
What a PDF actually holds

A PDF is a complete description of a fixed-layout document: the text, the fonts, the graphics and the images needed to show it. The text is stored as instructions to draw characters at certain positions, and each character is a code that maps to a in a font. It is a printout in digital form. It records where the ink goes, and says very little about what the words are doing there.

A scanned page doesn't even have that much. It is a picture, and getting words out of it takes optical character recognition, or : turning images of typed, handwritten or printed text into machine-encoded text. Along the way an OCR system usually does , picking out columns, paragraphs and captions, and then finds the lines and words.

Layout analysis is where a page stops being ink and becomes a document. It separates text from everything else, puts the text zones in their correct , and labels the parts that play different roles, such as titles, captions and footnotes.

Further reading PDF (Wikipedia)Optical character recognition (Wikipedia)Document layout analysis (Wikipedia)

Why it is hard
  1. i.

    Built to look right

    PDF's whole point is keeping a document looking the same everywhere, and that very emphasis makes it hard to convert or to pull out text, images, tables and metadata. There is a variant, , that carries structure information precisely so that text can be extracted reliably. Its existence tells you what the ordinary kind is missing.

  2. ii.

    Which column next?

    People read a two-column page without thinking. Software has to work it out. Top-down methods cut the page into columns and blocks using white space, which is fast but only works if they make assumptions about the layout. Bottom-up methods build words into lines into blocks from the raw pixels, which assumes nothing and takes time. Both have to cope with noise and with pages scanned slightly crooked.

  3. iii.

    Tables fight back

    Tables come in every shape and size, with complicated header arrangements, rows that run over several lines, all sorts of separator lines, and blank cells. Recovering which number belongs to which row and column from an image is, in the dry words of one paper, a non-trivial task. Get it slightly wrong and the numbers are all there, in the wrong places.

  4. iv.

    Good, cheap, pick one

    Some content stays hard for even the best tools: formulas, tiny fonts, old scans. tend to beat traditional open-source tools on quality. The best of them can also be prohibitively costly, over 6,240 USD per million pages for GPT-4o, and some documents can't be sent to someone else's API at all.

Further reading PDF (Wikipedia)Document layout analysis (Wikipedia)TableFormer: Table Structure Understanding with Transformers (arXiv)olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models (arXiv)

What Datalab is after

Datalab is a research lab that trains document intelligence models. Its aim is to turn messy, unstructured content into precise data that production systems can depend on, from clean training data for AI models to the documents a business runs on.

The jobs break down into a few pieces: convert a document into structured text, extract the key facts, split apart files that were originally several documents, and check that the output is any good.

Further reading High-Precision Document Intelligence (Datalab)

How they go at it
  1. Step 1: Built in the open

    Marker, Surya and Chandra are Datalab's own open-source models. Marker converts PDFs, images, slides, Word files and more into , JSON and HTML. It formats tables, forms, equations and code blocks, and strips out running headers and footers.

  2. Step 2: Layout, then letters

    Surya is the OCR model underneath, with 650 million parameters. Besides reading text, it does layout analysis with reading order and recognises the rows and columns of tables, and together with Marker it covers more than 90 languages.

  3. Step 3: A model for messy pages

    Chandra, the flagship model, converts images and PDFs into structured HTML, markdown or JSON while keeping the layout information. It is built for the messiest documents: complex tables, forms and handwriting.

  4. Step 4: Pipelines that check themselves

    On the managed platform these pieces are wired together into versioned pipelines, with continuous evaluations against rubrics and a reference corpus so that any regression gets flagged. Because Datalab owns its models, they can also run inside a customer's own cloud or on machines with no internet at all.

Further reading High-Precision Document Intelligence (Datalab)marker (Datalab (GitHub))surya (Datalab (GitHub))chandra (Datalab (GitHub))

Still open
  • Does every page need a giant model?

    Maybe not. One open-source team fine-tuned a vision language model of 7 billion parameters to turn PDFs into clean text in natural reading order, and it converts a million pages for only 176 USD, against thousands of dollars for the biggest commercial models.

  • How do you know a parse is right?

    Mostly by testing against labelled datasets, and several exist for PDF conversion and extraction. One recent benchmark gathers 1,400 PDFs full of the content types that still trip up the best tools, which is a polite way of saying the problem is not solved.

Further reading olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models (arXiv)PDF (Wikipedia)

About Datalab

Datalab is building document infrastructure that allows companies to understand and automate workflows built on human-readable documents. The platform learns the structure of a company's documents and the workflows that operate on top of them, creating a data layer that can be reliably used by AI systems.

Most enterprise knowledge lives in documents — contracts, reports, PDFs, internal wikis — that are readable by humans but poorly understood by machines. Datalab goes beyond OCR or point extraction, focusing instead on modeling document structure and workflow context so AI can be applied in a repeatable, production-grade way.

The platform is used by developers, enterprises, and AI labs that need reliable document understanding as a foundation for automation.

Words used here
glyph
The drawn shape of a character in a particular font.
OCR
Optical character recognition: software that turns pictures of text into text a computer can search and edit.
layout analysis
Working out which parts of a page are text, images or tables, what role each plays, and how they fit together.
reading order
The sequence a person would read a page's blocks in, which may not match their position top to bottom.
tagged PDF
A PDF that also stores the document's structure, such as headings and paragraphs, alongside how it looks.
Vision language models
AI models that take images as well as text as input, so they can read a page by looking at it.
markdown
A simple plain-text format for headings, lists and tables that both people and language models read easily.
Sources