Cognee
Memory infrastructure for AI agents
How do you give an AI agent a memory that gets better with use?
A language model can only use what sits in its All the text a language model can see at once while writing a reply, including the reply itself.: all the text it can see while it writes a reply. That window is its working memory, and it is separate from everything the model learned in training. It also gets wiped. Model APIs are Keeping nothing between requests, so each one must bring along everything the model needs., so every request has to carry the whole conversation again. Close the chat and, as far as the model is concerned, you never met.
The standard patch is Looking up relevant documents first and pasting them into the prompt before the model answers., or RAG. Documents are turned into Lists of numbers that represent a piece of text so that similar meanings end up close together., lists of numbers that stand for meaning, and kept in a A database that stores embeddings and finds the ones nearest to a query.. When a question arrives, a retriever picks the most relevant pieces and adds them to the prompt. A vector database finds records that are similar in meaning to the question, where an ordinary database looks for exact matches.
A A store of facts as entities joined by named relationships, like a map of who and what connects to what. stores things another way, as entities (objects, events, situations, ideas) and the relationships between them. A vector search asks what sounds like this. A graph can also ask what is connected to this.
Further reading Context windows (Anthropic)Using the Messages API (Anthropic)Retrieval-augmented generation (Wikipedia)Vector database (Wikipedia)Knowledge graph (Wikipedia)
- i.
More is not better
Bigger context windows sound like the answer, but more context isn't automatically better. As the token count grows, accuracy and recall degrade, which has earned the name The drop in a model's accuracy and recall as its context grows longer.. Models also use information best when it sits at the start or end of their input, and worst when it is buried in the middle, even models built for long contexts. Think of a desk piled so high you can no longer find the pen.
- ii.
Close is not relevant
Similarity is a rough guide to usefulness. Vector search can miss key facts needed to answer a question, and one fix is to bolt old-fashioned keyword search on top. Plain RAG also fails on questions about a whole collection, such as "What are the main themes?", because answering those means summarising everything, and no single passage holds the answer.
- iii.
Right source, wrong story
Fetching the right document doesn't guarantee the right answer. A model can lift a statement out of its context and reach the wrong conclusion. When sources disagree it may blend outdated and updated facts into one confident, misleading reply. RAG does not stop hallucination, and without specific training a model may answer when it should admit it isn't sure.
Further reading Context windows (Anthropic)Lost in the Middle: How Language Models Use Long Contexts (arXiv)Retrieval-augmented generation (Wikipedia)From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv)
Cognee is going after the forgetting. An agent without memory re-solves a task it already solved last week, burning tokens and time. Inside a company the answer to a question is often spread across a doc, a ticket and a chat session. And every organisation runs on rules no model was trained on, so agents make up their own.
The goal is persistent, long-term memory for agents that survives from one session to the next: documents, code and conversations turned into a knowledge graph that agents can search and reuse. The core is free and open source, and runs locally.
Further reading Cognee - Open-Source Agent Memory Platform (Cognee)Cognee README (GitHub (topoteretes/cognee))
- Step 1: Three stores, one memory
Cognee keeps three kinds of storage side by side. A relational store tracks documents, chunks and where each came from. A vector store holds the embeddings. A graph store holds entities and the relationships between them. The idea is data that is searchable by meaning and connected by relationships at the same time.
- Step 2: Reading into a graph
As data comes in, a pipeline splits documents into smaller pieces, uses LLMs to pick out entities and relationships, writes a summary of each chunk and embeds the lot. It can run again as a dataset grows and skips what it has already processed. With no LLM key set, a small local model does the extraction instead.
- Step 3: Asking the graph
When an agent asks something, Cognee's recall runs graph retrieval by default rather than plain embedding similarity, and each result is tagged so you can tell where it came from.
- Step 4: Learning from use
A separate improve step raises or lowers the importance of parts of the graph, based on feedback on answers that used them. It also moves useful short-term session memory into the permanent graph. An optional check compares newly touched facts with the ones already stored nearby and records each conflict as its own link.
Further reading Core Concepts Overview (Cognee Documentation)Cognify (Cognee Documentation)Cognee README (GitHub (topoteretes/cognee))Recall (Cognee Documentation)Improve (Cognee Documentation)
How do you tell whether a memory works?
Benchmarks such as BEAM test conversational memory with synthetic long conversations and a language model as the judge. Tuning studies of graph-based retrieval find gains that are consistent but not uniform, varying by dataset and metric, which says as much about the metrics as about the systems.
Will bigger context windows make memory layers unnecessary?
So far, longer isn't the same as better: recall degrades as context grows and facts in the middle get lost. That makes curating what goes into the window as important as how much of it there is, which is the job a memory layer is trying to do.
Further reading Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning (arXiv)Cognee README (GitHub (topoteretes/cognee))Context windows (Anthropic)Lost in the Middle: How Language Models Use Long Contexts (arXiv)
Cognee builds memory infrastructure for AI agents. The platform transforms data into a living knowledge graph that agents can query, update, and reason over — replacing traditional retrieval systems that degrade over time with a memory layer that actually learns from feedback and builds new knowledge continuously.
Cognee supports 38+ data types, integrates with major AI frameworks, and works across vector and graph databases. It sits behind agents as a persistent reasoning substrate, enabling more accurate, context-aware responses without the recall and accuracy limitations of static retrieval.
- context window
- All the text a language model can see at once while writing a reply, including the reply itself.
- stateless
- Keeping nothing between requests, so each one must bring along everything the model needs.
- retrieval-augmented generation
- Looking up relevant documents first and pasting them into the prompt before the model answers.
- embeddings
- Lists of numbers that represent a piece of text so that similar meanings end up close together.
- vector database
- A database that stores embeddings and finds the ones nearest to a query.
- knowledge graph
- A store of facts as entities joined by named relationships, like a map of who and what connects to what.
- context rot
- The drop in a model's accuracy and recall as its context grows longer.
- 1Context windows · Anthropic
- 2Using the Messages API · Anthropic
- 3Retrieval-augmented generation · Wikipedia
- 4Vector database · Wikipedia
- 5Knowledge graph · Wikipedia
- 6Lost in the Middle: How Language Models Use Long Contexts · arXiv
- 7From Local to Global: A Graph RAG Approach to Query-Focused Summarization · arXiv
- 8Cognee - Open-Source Agent Memory Platform · Cognee
- 9Cognee README · GitHub (topoteretes/cognee)
- 10Core Concepts Overview · Cognee Documentation
- 11Cognify · Cognee Documentation
- 12Recall · Cognee Documentation
- 13Improve · Cognee Documentation
- 14Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning · arXiv