# Kepler

Deterministic infrastructure for trustworthy AI

- Team: Vinoo Ganesh (co-founder), John McRaven (co-founder)
- Founded: 2025
- Invested: 2025
- Links: [Website](https://kepler.ai), [LinkedIn](https://www.linkedin.com/in/vinoo-ganesh)
- Field: How software gets made

## The problem: How do you make an AI's numbers something you can defend to an auditor?

### How a language model answers

A large language model writes one token at a time. At each step it works out a probability for every possible next token, then picks one. Always taking the likeliest, called greedy search, tends to give dull, repetitive text, so most systems roll a weighted die instead. A popular version, top-p sampling, draws only from the smallest set of likely tokens whose probabilities add up past a threshold.

So ask the same question twice and you may get two answers. Setting the temperature to zero should take the die away, yet LLM APIs still aren't deterministic in practice. One lab sampled 1,000 answers to a single prompt at temperature zero and got 80 different completions. The cause is mundane. Computers round when they add decimals, so (a + b) + c need not equal a + (b + c), and your answer can shift with how many strangers' requests share the server with yours.

Then there is hallucination: output that contains false or misleading information presented as fact. Chatbots can tuck plausible-sounding falsehoods, like invented citations, into otherwise sensible answers. In one New York personal-injury case, a lawyer filed a brief citing six fake precedents that ChatGPT had made up.

### Why it is hard

**Trained to guess.** Pre-training teaches a model to predict the next word, which rewards it for having a go even when it lacks the information. A student who never leaves an exam question blank scores better on average, and is sometimes confidently wrong.

**Filings are hard homework.** Financial documents are where this bites. A benchmark of questions about public company filings was meant as a minimum standard, with clear-cut, straightforward questions. A leading model with a retrieval system still answered wrongly or refused on 81% of them, and every model tested showed weaknesses such as hallucination.

**Same digits, different fact.** A number in a filing is mostly context: the currency, the scale in the table header, the date in the column. One LinkedIn quarterly report tags the same figure, $1,322,500,000, three times as three different facts. Search for the digits and you match whichever one you meant.

**No copy button.** Language models don't copy; they retype, and the retyping is usually right. At one slip in 1,000 figures, a 400-figure workbook carries a wrong number every two or three runs, and nothing says which. Better models fail more quietly, with the right digits from the wrong period or unit.

### What Kepler is after

Finance runs on numbers someone has to stand behind. Every figure in a regulatory filing, a deal pitch or a research report has to be checkable against its source, and the traditional tools can pull data but still leave the checking to analysts. As Kepler puts it, "one wrong figure in 1,000 means someone checks all 1,000."

Kepler wants AI whose every number traces back to the original filing, page and line item, where the same input gives the same output every time. It starts with finance, where SEC filings, earnings transcripts and market data come already indexed, and is building the same machinery for other industries' own data.

### How they go at it

**Language here, numbers there.** The AI interprets what you asked and breaks it into an execution plan. It never retrieves data or computes numbers itself. Deterministic code does that, and the two layers talk through structured contracts, never raw model output.

**One word, one number.** Underneath sits an ontology, a map of financial concepts and how they relate, so "revenue", "top line" and "total sales" all resolve to the same verified number. Common workflows, such as enterprise value across a messy capital structure, are built as skills designed so the same input always gives the same output.

**Name it, don't type it.** Every tagged figure in a filing becomes a fact with an ID. The model doesn't write the digits; it names the fact, and the system fills in the filed value, linked to its line. An unreferenced number is treated like a compile error. It's the same fix software found for SQL injection: keep the data out of the command string.

**Choices written down.** Some questions have several true answers, like whether debt counts at principal or at its balance-sheet value. Kepler makes that choice a named setting in the output, so the next reader can see it and flip it.

### Still open

Can inference itself be made repeatable? Yes, with work. Rewriting the core GPU routines so results don't depend on batch size made all 1,000 completions identical, and a first, unoptimised version ran slower, though not disastrously. Identical is still not the same as correct.

Will better models make the problem go away? Hallucination rates are falling, but not as fast as they look. In one release each claim became 23% more likely to be right, yet only 3% fewer responses contained an error, because the model was making more claims per response.

Does a citation mean the model used the source? Not necessarily. Planting a few words of a model's answer into a document it hadn't cited moved a production retrieval model's citation there up to 57% of the time, with the answer unchanged.

### Words used here

- **token**: A chunk of text, often a word or part of one, that a language model reads and writes one at a time.
- **greedy search**: Always picking the single most likely next token, with no randomness.
- **top-p sampling**: Choosing the next token at random from only the most likely candidates, whose combined probability just passes a set threshold.
- **temperature**: A setting that controls how much randomness goes into picking tokens; at zero the model should always take the likeliest.
- **hallucination**: AI output that states something false as though it were fact.
- **ontology**: A structured map of the concepts in a field, their names and how they relate to each other.
- **SQL injection**: An attack where text typed into a form is run as a database command because data and code travel in one string.

## About Kepler

Kepler is building the provenance and traceability infrastructure layer for AI in regulated environments. The platform separates AI reasoning from deterministic data retrieval, ensuring every output is traced to its source document, page, and line item. Starting with financial services, where the product is live and in production across 950K+ SEC filings and 14K+ companies. Based in New York City.

## Sources

1. [Top-p sampling](https://en.wikipedia.org/wiki/Top-p_sampling), Wikipedia
2. [Defeating Nondeterminism in LLM Inference](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/), Thinking Machines Lab
3. [Hallucination (artificial intelligence)](https://en.wikipedia.org/wiki/Hallucination_(artificial_intelligence)), Wikipedia
4. [FinanceBench: A New Benchmark for Financial Question Answering](https://arxiv.org/abs/2311.11944), arXiv
5. [AI that proves it's right](https://kepler.ai), Kepler
6. [Better language models won't fix hallucination. They'll make it quieter.](https://kepler.ai/research/engineering/quieter-hallucination), Kepler
7. [How Kepler built verifiable AI for financial services with Claude](https://claude.com/blog/how-kepler-built-verifiable-ai-for-financial-services-with-claude), Anthropic
