# Cognee

Memory infrastructure for AI agents

- Team: Vasilije Markovic (founder)
- Founded: 2024
- Invested: 2025
- Links: [Website](https://cognee.ai)
- Field: Reasoning and trust

## The problem: How do you give an AI agent a memory that gets better with use?

### How AI memory works

A language model can only use what sits in its context window: all the text it can see while it writes a reply. That window is its working memory, and it is separate from everything the model learned in training. It also gets wiped. Model APIs are stateless, so every request has to carry the whole conversation again. Close the chat and, as far as the model is concerned, you never met.

The standard patch is retrieval-augmented generation, or RAG. Documents are turned into embeddings, lists of numbers that stand for meaning, and kept in a vector database. When a question arrives, a retriever picks the most relevant pieces and adds them to the prompt. A vector database finds records that are similar in meaning to the question, where an ordinary database looks for exact matches.

A knowledge graph stores things another way, as entities (objects, events, situations, ideas) and the relationships between them. A vector search asks what sounds like this. A graph can also ask what is connected to this.

### Why it is hard

**More is not better.** Bigger context windows sound like the answer, but more context isn't automatically better. As the token count grows, accuracy and recall degrade, which has earned the name context rot. Models also use information best when it sits at the start or end of their input, and worst when it is buried in the middle, even models built for long contexts. Think of a desk piled so high you can no longer find the pen.

**Close is not relevant.** Similarity is a rough guide to usefulness. Vector search can miss key facts needed to answer a question, and one fix is to bolt old-fashioned keyword search on top. Plain RAG also fails on questions about a whole collection, such as "What are the main themes?", because answering those means summarising everything, and no single passage holds the answer.

**Right source, wrong story.** Fetching the right document doesn't guarantee the right answer. A model can lift a statement out of its context and reach the wrong conclusion. When sources disagree it may blend outdated and updated facts into one confident, misleading reply. RAG does not stop hallucination, and without specific training a model may answer when it should admit it isn't sure.

### What Cognee is after

Cognee is going after the forgetting. An agent without memory re-solves a task it already solved last week, burning tokens and time. Inside a company the answer to a question is often spread across a doc, a ticket and a chat session. And every organisation runs on rules no model was trained on, so agents make up their own.

The goal is persistent, long-term memory for agents that survives from one session to the next: documents, code and conversations turned into a knowledge graph that agents can search and reuse. The core is free and open source, and runs locally.

### How they go at it

**Three stores, one memory.** Cognee keeps three kinds of storage side by side. A relational store tracks documents, chunks and where each came from. A vector store holds the embeddings. A graph store holds entities and the relationships between them. The idea is data that is searchable by meaning and connected by relationships at the same time.

**Reading into a graph.** As data comes in, a pipeline splits documents into smaller pieces, uses LLMs to pick out entities and relationships, writes a summary of each chunk and embeds the lot. It can run again as a dataset grows and skips what it has already processed. With no LLM key set, a small local model does the extraction instead.

**Asking the graph.** When an agent asks something, Cognee's recall runs graph retrieval by default rather than plain embedding similarity, and each result is tagged so you can tell where it came from.

**Learning from use.** A separate improve step raises or lowers the importance of parts of the graph, based on feedback on answers that used them. It also moves useful short-term session memory into the permanent graph. An optional check compares newly touched facts with the ones already stored nearby and records each conflict as its own link.

### Still open

How do you tell whether a memory works? Benchmarks such as BEAM test conversational memory with synthetic long conversations and a language model as the judge. Tuning studies of graph-based retrieval find gains that are consistent but not uniform, varying by dataset and metric, which says as much about the metrics as about the systems.

Will bigger context windows make memory layers unnecessary? So far, longer isn't the same as better: recall degrades as context grows and facts in the middle get lost. That makes curating what goes into the window as important as how much of it there is, which is the job a memory layer is trying to do.

### Words used here

- **context window**: All the text a language model can see at once while writing a reply, including the reply itself.
- **stateless**: Keeping nothing between requests, so each one must bring along everything the model needs.
- **retrieval-augmented generation**: Looking up relevant documents first and pasting them into the prompt before the model answers.
- **embeddings**: Lists of numbers that represent a piece of text so that similar meanings end up close together.
- **vector database**: A database that stores embeddings and finds the ones nearest to a query.
- **knowledge graph**: A store of facts as entities joined by named relationships, like a map of who and what connects to what.
- **context rot**: The drop in a model's accuracy and recall as its context grows longer.

## About Cognee

Cognee builds memory infrastructure for AI agents. The platform transforms data into a living knowledge graph that agents can query, update, and reason over — replacing traditional retrieval systems that degrade over time with a memory layer that actually learns from feedback and builds new knowledge continuously.

Cognee supports 38+ data types, integrates with major AI frameworks, and works across vector and graph databases. It sits behind agents as a persistent reasoning substrate, enabling more accurate, context-aware responses without the recall and accuracy limitations of static retrieval.

## Sources

1. [Context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows), Anthropic
2. [Using the Messages API](https://platform.claude.com/docs/en/build-with-claude/working-with-messages), Anthropic
3. [Retrieval-augmented generation](https://en.wikipedia.org/wiki/Retrieval-augmented_generation), Wikipedia
4. [Vector database](https://en.wikipedia.org/wiki/Vector_database), Wikipedia
5. [Knowledge graph](https://en.wikipedia.org/wiki/Knowledge_graph), Wikipedia
6. [Lost in the Middle: How Language Models Use Long Contexts](https://arxiv.org/abs/2307.03172), arXiv
7. [From Local to Global: A Graph RAG Approach to Query-Focused Summarization](https://arxiv.org/abs/2404.16130), arXiv
8. [Cognee - Open-Source Agent Memory Platform](https://www.cognee.ai/), Cognee
9. [Cognee README](https://raw.githubusercontent.com/topoteretes/cognee/main/README.md), GitHub (topoteretes/cognee)
10. [Core Concepts Overview](https://docs.cognee.ai/core-concepts/overview), Cognee Documentation
11. [Cognify](https://docs.cognee.ai/core-concepts/main-operations/legacy-operations/cognify), Cognee Documentation
12. [Recall](https://docs.cognee.ai/core-concepts/main-operations/recall), Cognee Documentation
13. [Improve](https://docs.cognee.ai/core-concepts/main-operations/improve), Cognee Documentation
14. [Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning](https://arxiv.org/abs/2505.24478), arXiv
