← PortfolioReasoning and trust

Cognee

Memory infrastructure for AI agents

Team
Vasilije Markovicfounder
Founded
2024
Invested
2025
Links
The problem

How do you give an AI agent a memory that gets better with use?

Click any point in the graph to ask about it, then mark the answer useful or not and watch the links it used get heavier or lighter.A doc, a ticket and a chat are cut into pieces and read into one graph, and every piece remembers where it came from. A question lands on the closest match in meaning, then recall follows the heaviest links out from there, so the answer draws on all three sources. Feedback on the answer makes the links it used heavier or lighter, and that changes the next walk.An illustration, not real data.
How AI memory works

A language model can only use what sits in its : all the text it can see while it writes a reply. That window is its working memory, and it is separate from everything the model learned in training. It also gets wiped. Model APIs are , so every request has to carry the whole conversation again. Close the chat and, as far as the model is concerned, you never met.

The standard patch is , or RAG. Documents are turned into , lists of numbers that stand for meaning, and kept in a . When a question arrives, a retriever picks the most relevant pieces and adds them to the prompt. A vector database finds records that are similar in meaning to the question, where an ordinary database looks for exact matches.

A stores things another way, as entities (objects, events, situations, ideas) and the relationships between them. A vector search asks what sounds like this. A graph can also ask what is connected to this.

Further reading Context windows (Anthropic)Using the Messages API (Anthropic)Retrieval-augmented generation (Wikipedia)Vector database (Wikipedia)Knowledge graph (Wikipedia)

Why it is hard
  1. i.

    More is not better

    Bigger context windows sound like the answer, but more context isn't automatically better. As the token count grows, accuracy and recall degrade, which has earned the name . Models also use information best when it sits at the start or end of their input, and worst when it is buried in the middle, even models built for long contexts. Think of a desk piled so high you can no longer find the pen.

  2. ii.

    Close is not relevant

    Similarity is a rough guide to usefulness. Vector search can miss key facts needed to answer a question, and one fix is to bolt old-fashioned keyword search on top. Plain RAG also fails on questions about a whole collection, such as "What are the main themes?", because answering those means summarising everything, and no single passage holds the answer.

  3. iii.

    Right source, wrong story

    Fetching the right document doesn't guarantee the right answer. A model can lift a statement out of its context and reach the wrong conclusion. When sources disagree it may blend outdated and updated facts into one confident, misleading reply. RAG does not stop hallucination, and without specific training a model may answer when it should admit it isn't sure.

Further reading Context windows (Anthropic)Lost in the Middle: How Language Models Use Long Contexts (arXiv)Retrieval-augmented generation (Wikipedia)From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv)

What Cognee is after

Cognee is going after the forgetting. An agent without memory re-solves a task it already solved last week, burning tokens and time. Inside a company the answer to a question is often spread across a doc, a ticket and a chat session. And every organisation runs on rules no model was trained on, so agents make up their own.

The goal is persistent, long-term memory for agents that survives from one session to the next: documents, code and conversations turned into a knowledge graph that agents can search and reuse. The core is free and open source, and runs locally.

Further reading Cognee - Open-Source Agent Memory Platform (Cognee)Cognee README (GitHub (topoteretes/cognee))

How they go at it
  1. Step 1: Three stores, one memory

    Cognee keeps three kinds of storage side by side. A relational store tracks documents, chunks and where each came from. A vector store holds the embeddings. A graph store holds entities and the relationships between them. The idea is data that is searchable by meaning and connected by relationships at the same time.

  2. Step 2: Reading into a graph

    As data comes in, a pipeline splits documents into smaller pieces, uses LLMs to pick out entities and relationships, writes a summary of each chunk and embeds the lot. It can run again as a dataset grows and skips what it has already processed. With no LLM key set, a small local model does the extraction instead.

  3. Step 3: Asking the graph

    When an agent asks something, Cognee's recall runs graph retrieval by default rather than plain embedding similarity, and each result is tagged so you can tell where it came from.

  4. Step 4: Learning from use

    A separate improve step raises or lowers the importance of parts of the graph, based on feedback on answers that used them. It also moves useful short-term session memory into the permanent graph. An optional check compares newly touched facts with the ones already stored nearby and records each conflict as its own link.

Further reading Core Concepts Overview (Cognee Documentation)Cognify (Cognee Documentation)Cognee README (GitHub (topoteretes/cognee))Recall (Cognee Documentation)Improve (Cognee Documentation)

Still open
  • How do you tell whether a memory works?

    Benchmarks such as BEAM test conversational memory with synthetic long conversations and a language model as the judge. Tuning studies of graph-based retrieval find gains that are consistent but not uniform, varying by dataset and metric, which says as much about the metrics as about the systems.

  • Will bigger context windows make memory layers unnecessary?

    So far, longer isn't the same as better: recall degrades as context grows and facts in the middle get lost. That makes curating what goes into the window as important as how much of it there is, which is the job a memory layer is trying to do.

Further reading Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning (arXiv)Cognee README (GitHub (topoteretes/cognee))Context windows (Anthropic)Lost in the Middle: How Language Models Use Long Contexts (arXiv)

About Cognee

Cognee builds memory infrastructure for AI agents. The platform transforms data into a living knowledge graph that agents can query, update, and reason over — replacing traditional retrieval systems that degrade over time with a memory layer that actually learns from feedback and builds new knowledge continuously.

Cognee supports 38+ data types, integrates with major AI frameworks, and works across vector and graph databases. It sits behind agents as a persistent reasoning substrate, enabling more accurate, context-aware responses without the recall and accuracy limitations of static retrieval.

Words used here
context window
All the text a language model can see at once while writing a reply, including the reply itself.
stateless
Keeping nothing between requests, so each one must bring along everything the model needs.
retrieval-augmented generation
Looking up relevant documents first and pasting them into the prompt before the model answers.
embeddings
Lists of numbers that represent a piece of text so that similar meanings end up close together.
vector database
A database that stores embeddings and finds the ones nearest to a query.
knowledge graph
A store of facts as entities joined by named relationships, like a map of who and what connects to what.
context rot
The drop in a model's accuracy and recall as its context grows longer.
Sources