← PortfolioHow software gets made

Kepler

Deterministic infrastructure for trustworthy AI

Team
Vinoo Ganeshco-founder
John McRavenco-founder
Founded
2025
Invested
2025
The problem

How do you make an AI's numbers something you can defend to an auditor?

Click any figure in the filing. The top answer names it and is filled in from its line; the retyped copy below can only match digits.A filing where every tagged figure has an ID, and two answers to the same question. In the top one the model names the figure and the system fills in the filed value, linked back to its line. In the bottom one the model retypes the digits, which match every figure that shares them, and a number with no reference is refused.An illustration, not real data.
How a language model answers

A large language model writes one at a time. At each step it works out a probability for every possible next token, then picks one. Always taking the likeliest, called , tends to give dull, repetitive text, so most systems roll a weighted die instead. A popular version, , draws only from the smallest set of likely tokens whose probabilities add up past a threshold.

So ask the same question twice and you may get two answers. Setting the to zero should take the die away, yet LLM APIs still aren't deterministic in practice. One lab sampled 1,000 answers to a single prompt at temperature zero and got 80 different completions. The cause is mundane. Computers round when they add decimals, so (a + b) + c need not equal a + (b + c), and your answer can shift with how many strangers' requests share the server with yours.

Then there is : output that contains false or misleading information presented as fact. Chatbots can tuck plausible-sounding falsehoods, like invented citations, into otherwise sensible answers. In one New York personal-injury case, a lawyer filed a brief citing six fake precedents that ChatGPT had made up.

Further reading Top-p sampling (Wikipedia)Defeating Nondeterminism in LLM Inference (Thinking Machines Lab)Hallucination (artificial intelligence) (Wikipedia)

Why it is hard
  1. i.

    Trained to guess

    Pre-training teaches a model to predict the next word, which rewards it for having a go even when it lacks the information. A student who never leaves an exam question blank scores better on average, and is sometimes confidently wrong.

  2. ii.

    Filings are hard homework

    Financial documents are where this bites. A benchmark of questions about public company filings was meant as a minimum standard, with clear-cut, straightforward questions. A leading model with a retrieval system still answered wrongly or refused on 81% of them, and every model tested showed weaknesses such as hallucination.

  3. iii.

    Same digits, different fact

    A number in a filing is mostly context: the currency, the scale in the table header, the date in the column. One LinkedIn quarterly report tags the same figure, $1,322,500,000, three times as three different facts. Search for the digits and you match whichever one you meant.

  4. iv.

    No copy button

    Language models don't copy; they retype, and the retyping is usually right. At one slip in 1,000 figures, a 400-figure workbook carries a wrong number every two or three runs, and nothing says which. Better models fail more quietly, with the right digits from the wrong period or unit.

Further reading Hallucination (artificial intelligence) (Wikipedia)FinanceBench: A New Benchmark for Financial Question Answering (arXiv)Better language models won't fix hallucination. They'll make it quieter. (Kepler)

What Kepler is after

Finance runs on numbers someone has to stand behind. Every figure in a regulatory filing, a deal pitch or a research report has to be checkable against its source, and the traditional tools can pull data but still leave the checking to analysts. As Kepler puts it, "one wrong figure in 1,000 means someone checks all 1,000."

Kepler wants AI whose every number traces back to the original filing, page and line item, where the same input gives the same output every time. It starts with finance, where SEC filings, earnings transcripts and market data come already indexed, and is building the same machinery for other industries' own data.

Further reading How Kepler built verifiable AI for financial services with Claude (Anthropic)Better language models won't fix hallucination. They'll make it quieter. (Kepler)AI that proves it's right (Kepler)

How they go at it
  1. Step 1: Language here, numbers there

    The AI interprets what you asked and breaks it into an execution plan. It never retrieves data or computes numbers itself. Deterministic code does that, and the two layers talk through structured contracts, never raw model output.

  2. Step 2: One word, one number

    Underneath sits an , a map of financial concepts and how they relate, so "revenue", "top line" and "total sales" all resolve to the same verified number. Common workflows, such as enterprise value across a messy capital structure, are built as skills designed so the same input always gives the same output.

  3. Step 3: Name it, don't type it

    Every tagged figure in a filing becomes a fact with an ID. The model doesn't write the digits; it names the fact, and the system fills in the filed value, linked to its line. An unreferenced number is treated like a compile error. It's the same fix software found for : keep the data out of the command string.

  4. Step 4: Choices written down

    Some questions have several true answers, like whether debt counts at principal or at its balance-sheet value. Kepler makes that choice a named setting in the output, so the next reader can see it and flip it.

Further reading AI that proves it's right (Kepler)How Kepler built verifiable AI for financial services with Claude (Anthropic)Better language models won't fix hallucination. They'll make it quieter. (Kepler)

Still open
  • Can inference itself be made repeatable?

    Yes, with work. Rewriting the core GPU routines so results don't depend on batch size made all 1,000 completions identical, and a first, unoptimised version ran slower, though not disastrously. Identical is still not the same as correct.

  • Will better models make the problem go away?

    Hallucination rates are falling, but not as fast as they look. In one release each claim became 23% more likely to be right, yet only 3% fewer responses contained an error, because the model was making more claims per response.

  • Does a citation mean the model used the source?

    Not necessarily. Planting a few words of a model's answer into a document it hadn't cited moved a production retrieval model's citation there up to 57% of the time, with the answer unchanged.

Further reading Defeating Nondeterminism in LLM Inference (Thinking Machines Lab)Better language models won't fix hallucination. They'll make it quieter. (Kepler)

About Kepler

Kepler is building the provenance and traceability infrastructure layer for AI in regulated environments. The platform separates AI reasoning from deterministic data retrieval, ensuring every output is traced to its source document, page, and line item. Starting with financial services, where the product is live and in production across 950K+ SEC filings and 14K+ companies. Based in New York City.

Words used here
token
A chunk of text, often a word or part of one, that a language model reads and writes one at a time.
greedy search
Always picking the single most likely next token, with no randomness.
top-p sampling
Choosing the next token at random from only the most likely candidates, whose combined probability just passes a set threshold.
temperature
A setting that controls how much randomness goes into picking tokens; at zero the model should always take the likeliest.
hallucination
AI output that states something false as though it were fact.
ontology
A structured map of the concepts in a field, their names and how they relate to each other.
SQL injection
An attack where text typed into a form is run as a database command because data and code travel in one string.
Sources