← PortfolioHow software gets made

Augment Code

AI agents that understand your entire codebase

Team
Igor OstrovskyChief Architect
Guy Gur-AriCEO
Founded
2022
Invested
2022
Links
The problem

How does an AI find its way around millions of lines of code?

Click any file to put a task there. The engine follows what it is related to across the codebase and sends the model only the few pieces that matter.A codebase from above, with faint threads between files that are about the same thing. A task lands on one file, and the engine follows those threads across services to everything the change will touch. It ranks what it finds and sends the model one compressed piece from each place, and leaves the rest out.An illustration, not real data.
How coding agents see code

A large language model can only look at so much at once. That limit is its , and it is measured in , the chunks of text a model reads. Anything outside the window the model simply can't see, unless it is summarized, retrieved or handed over again.

A coding agent is a model wrapped in a loop with tools. A separate program watches what the model writes for a special tool-calling syntax, runs the tool, and feeds the output back in. The tools can fetch information or run code.

The standard way to bring in knowledge the model wasn't trained on is retrieval-augmented generation, or . The material is turned into , numerical representations in a large vector space, and kept in a vector database. When a question comes in, a retriever picks the most relevant pieces and hands them to the model along with it.

Further reading Context window (Wikipedia)Large language model (Wikipedia)Retrieval-augmented generation (Wikipedia)

Why it is hard
  1. i.

    Big window, poor memory

    Windows have grown fast, and long-context systems now report hundreds of thousands to millions of tokens. That doesn't mean the model uses all of it equally well. In the "Lost in the Middle" study, models did best when the relevant information sat at the beginning or end of the input, and much worse when it was buried in the middle. Think of a reader who remembers the first and last chapters of a book and skims the rest.

  2. ii.

    Fixes that sprawl

    Real bugs rarely live in one place. SWE-bench is a test built from real GitHub issues across 12 popular Python repositories, and resolving them frequently requires coordinating changes across multiple functions, classes and even files at the same time. When it came out, the best model could solve a mere 1.96% of the issues.

  3. iii.

    Grep, open, repeat

    Most coding agents build context the brute-force way: for a word, open broad chunks of files, and replay everything they found back into the model. They spend budget before they know what matters, and every miss becomes another tool call and another expensive turn.

  4. iv.

    Which codebase, exactly?

    A codebase is a moving target. Switching git branches, a search-and-replace or an automatic reformat can change hundreds of files within a second. Retrieve from the wrong branch and the function you are working on may not even exist there, and models are likely to hallucinate when faced with names that aren't defined in the context.

Further reading Context window (Wikipedia)Lost in the Middle: How Language Models Use Long Contexts (arXiv)SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (arXiv)The Context Engine (Augment Code)A real-time index for your codebase: Secure, personal, scalable (Augment Code)

What Augment Code is after

Augment wants agents that understand a whole codebase, from side projects to enterprise , and that do more real engineering work on the same model. Its pitch is blunt: "Same frontier models. More solved work per budget."

The thinking is that what a model gets shown matters as much as the model itself. Put the right slice of code in front of it and it spends fewer turns searching and more of them making correct changes.

Further reading The Context Engine (Augment Code)

How they go at it
  1. Step 1: An index that keeps up

    Augment keeps a real-time index of the codebase for each user, so it follows your branch and not just the main one. The aim is to update that index within a few seconds of any change. It uses its own custom models to make embeddings in place of generic ones, and processes many thousands of files per second, so a branch switch is handled almost instantly.

  2. Step 2: Search by meaning

    The Context Engine indexes code semantically and maps the relationships between hundreds of thousands of files. Ask it to add logging to payment requests and it traces the whole path: React app, Node API, payment service, database and webhook handlers. It also draws on commit history, docs, tickets and design decisions, which explain why the code is the way it is.

  3. Step 3: Send less, not more

    It doesn't dump the entire codebase into the prompt. It ranks what it finds by relevance to the task, compresses it and sends only what matters. On one public benchmark with the same underlying model, Augment's own agent spent 33% less than a rival agent while solving tasks at effectively the same rate.

  4. Step 4: Context for other agents

    The engine isn't tied to Augment's own tools. Other coding agents can plug into it through the , and developers can build on it with an SDK.

Further reading A real-time index for your codebase: Secure, personal, scalable (Augment Code)The Context Engine (Augment Code)Context Services overview (Augment Code)

Still open
  • Do benchmarks say much about real work?

    Most coding benchmarks use self-contained tasks that need no prior context and are graded automatically, which may make models look better than they are. They can also make models look worse, because an agent may fail on a small bottleneck a person would fix in seconds. Turning a score into impact in the wild is hard.

  • Will bigger context windows make retrieval unnecessary?

    Windows keep growing, and some models have been tested on retrieval at up to 10 million tokens. But size is not the same as use, and newer benchmarks probe skills beyond simple retrieval, including understanding a whole code repository.

Further reading Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR)Context window (Wikipedia)

About Augment Code

Augment Code builds AI agents designed to work across entire, real-world codebases, not just individual files or snippets. The system integrates directly into existing developer workflows — from IDE to CLI to code review — allowing AI to participate where engineers already work.

In large enterprises, software systems often consist of millions of lines of code spread across thousands of files, accumulated over many years. These codebases encode business logic, operational assumptions, and institutional knowledge that is difficult for humans — and most AI tools — to navigate. Augment is built to maintain deep, persistent codebase context, enabling use cases that go beyond autocomplete or one-off code generation.

The team initially focused on a coding assistant and has since expanded toward agentic workflows — AI that can take on entire engineering tasks, not just suggest the next line.

Words used here
context window
The most text a language model can take in at once while producing an answer.
tokens
The small chunks, often pieces of words, that a model reads and writes text in.
RAG
Retrieval-augmented generation: finding relevant documents first and handing them to the model with the question.
embeddings
Lists of numbers that represent meaning, so similar pieces of text or code sit close together.
grep
A classic command-line tool that searches files for an exact word or pattern.
monorepos
Single repositories that hold the code for many projects or services at once.
Model Context Protocol
An open standard for plugging tools and data sources into AI agents.
Sources