# Augment Code

AI agents that understand your entire codebase

- Team: Igor Ostrovsky (Chief Architect), Guy Gur-Ari (CEO)
- Founded: 2022
- Invested: 2022
- Links: [Website](https://www.augmentcode.com), [LinkedIn](https://www.linkedin.com/in/igoro/), [X](https://x.com/igoro), [GitHub](https://github.com/guygurari), [Scholar](https://scholar.google.com/citations?user=RFMJY4IAAAAJ&hl=en)
- Field: How software gets made

## The problem: How does an AI find its way around millions of lines of code?

### How coding agents see code

A large language model can only look at so much at once. That limit is its context window, and it is measured in tokens, the chunks of text a model reads. Anything outside the window the model simply can't see, unless it is summarized, retrieved or handed over again.

A coding agent is a model wrapped in a loop with tools. A separate program watches what the model writes for a special tool-calling syntax, runs the tool, and feeds the output back in. The tools can fetch information or run code.

The standard way to bring in knowledge the model wasn't trained on is retrieval-augmented generation, or RAG. The material is turned into embeddings, numerical representations in a large vector space, and kept in a vector database. When a question comes in, a retriever picks the most relevant pieces and hands them to the model along with it.

### Why it is hard

**Big window, poor memory.** Windows have grown fast, and long-context systems now report hundreds of thousands to millions of tokens. That doesn't mean the model uses all of it equally well. In the "Lost in the Middle" study, models did best when the relevant information sat at the beginning or end of the input, and much worse when it was buried in the middle. Think of a reader who remembers the first and last chapters of a book and skims the rest.

**Fixes that sprawl.** Real bugs rarely live in one place. SWE-bench is a test built from real GitHub issues across 12 popular Python repositories, and resolving them frequently requires coordinating changes across multiple functions, classes and even files at the same time. When it came out, the best model could solve a mere 1.96% of the issues.

**Grep, open, repeat.** Most coding agents build context the brute-force way: grep for a word, open broad chunks of files, and replay everything they found back into the model. They spend budget before they know what matters, and every miss becomes another tool call and another expensive turn.

**Which codebase, exactly?.** A codebase is a moving target. Switching git branches, a search-and-replace or an automatic reformat can change hundreds of files within a second. Retrieve from the wrong branch and the function you are working on may not even exist there, and models are likely to hallucinate when faced with names that aren't defined in the context.

### What Augment Code is after

Augment wants agents that understand a whole codebase, from side projects to enterprise monorepos, and that do more real engineering work on the same model. Its pitch is blunt: "Same frontier models. More solved work per budget."

The thinking is that what a model gets shown matters as much as the model itself. Put the right slice of code in front of it and it spends fewer turns searching and more of them making correct changes.

### How they go at it

**An index that keeps up.** Augment keeps a real-time index of the codebase for each user, so it follows your branch and not just the main one. The aim is to update that index within a few seconds of any change. It uses its own custom models to make embeddings in place of generic ones, and processes many thousands of files per second, so a branch switch is handled almost instantly.

**Search by meaning.** The Context Engine indexes code semantically and maps the relationships between hundreds of thousands of files. Ask it to add logging to payment requests and it traces the whole path: React app, Node API, payment service, database and webhook handlers. It also draws on commit history, docs, tickets and design decisions, which explain why the code is the way it is.

**Send less, not more.** It doesn't dump the entire codebase into the prompt. It ranks what it finds by relevance to the task, compresses it and sends only what matters. On one public benchmark with the same underlying model, Augment's own agent spent 33% less than a rival agent while solving tasks at effectively the same rate.

**Context for other agents.** The engine isn't tied to Augment's own tools. Other coding agents can plug into it through the Model Context Protocol, and developers can build on it with an SDK.

### Still open

Do benchmarks say much about real work? Most coding benchmarks use self-contained tasks that need no prior context and are graded automatically, which may make models look better than they are. They can also make models look worse, because an agent may fail on a small bottleneck a person would fix in seconds. Turning a score into impact in the wild is hard.

Will bigger context windows make retrieval unnecessary? Windows keep growing, and some models have been tested on retrieval at up to 10 million tokens. But size is not the same as use, and newer benchmarks probe skills beyond simple retrieval, including understanding a whole code repository.

### Words used here

- **context window**: The most text a language model can take in at once while producing an answer.
- **tokens**: The small chunks, often pieces of words, that a model reads and writes text in.
- **RAG**: Retrieval-augmented generation: finding relevant documents first and handing them to the model with the question.
- **embeddings**: Lists of numbers that represent meaning, so similar pieces of text or code sit close together.
- **grep**: A classic command-line tool that searches files for an exact word or pattern.
- **monorepos**: Single repositories that hold the code for many projects or services at once.
- **Model Context Protocol**: An open standard for plugging tools and data sources into AI agents.

## About Augment Code

Augment Code builds AI agents designed to work across entire, real-world codebases, not just individual files or snippets. The system integrates directly into existing developer workflows — from IDE to CLI to code review — allowing AI to participate where engineers already work.

In large enterprises, software systems often consist of millions of lines of code spread across thousands of files, accumulated over many years. These codebases encode business logic, operational assumptions, and institutional knowledge that is difficult for humans — and most AI tools — to navigate. Augment is built to maintain deep, persistent codebase context, enabling use cases that go beyond autocomplete or one-off code generation.

The team initially focused on a coding assistant and has since expanded toward agentic workflows — AI that can take on entire engineering tasks, not just suggest the next line.

## Sources

1. [Context window](https://en.wikipedia.org/wiki/Context_window), Wikipedia
2. [Large language model](https://en.wikipedia.org/wiki/Large_language_model), Wikipedia
3. [Retrieval-augmented generation](https://en.wikipedia.org/wiki/Retrieval-augmented_generation), Wikipedia
4. [Lost in the Middle: How Language Models Use Long Contexts](https://arxiv.org/abs/2307.03172), arXiv
5. [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770), arXiv
6. [The Context Engine](https://www.augmentcode.com/context-engine), Augment Code
7. [A real-time index for your codebase: Secure, personal, scalable](https://www.augmentcode.com/blog/a-real-time-index-for-your-codebase-secure-personal-scalable), Augment Code
8. [Context Services overview](https://docs.augmentcode.com/context-services/overview), Augment Code
9. [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/), METR
