# Aneta

AI automation for scientific workflows

- Team: Yasyf Mohamedali (founder)
- Founded: 2023
- Invested: 2023
- Links: [Website](https://aneta.company), [LinkedIn](https://www.linkedin.com/in/yasyf/), [X](https://x.com/yasyf), [GitHub](https://github.com/yasyf), [Scholar](https://scholar.google.com/citations?user=_x3u8rQAAAAJ&hl=en)
- Field: Materials and the built world

## The problem: How do you trust a research report that a machine wrote?

### How evidence gets gathered

When scientists want to know what the research actually says on a question, the gold standard is a systematic review. It is a structured synthesis of the evidence: find the studies, pull out their data, critically appraise them, and boil it all down to a conclusion. Along with meta-analyses, systematic reviews usually sit at the top of the hierarchy of evidence.

They are also slow. One analysis of registered reviews found they took 67.3 weeks on average from start to publication, with about five authors each. Searches turned up anywhere from 27 to 92,020 studies, and on average under 3% of what was found made it into the final review. Most of the work is reading things that turn out not to matter, and the pile keeps growing.

### Why it is hard

**Too slow to stay fresh.** Systematic reviews take a lot of time and a lot of people. Critics point out that published reviews are often biased, out of date and far too long. By the time a careful answer comes out, the question may have moved on.

**Citations that don't exist.** Chatbots write fluent literature reviews, sometimes about papers nobody wrote. When researchers asked two versions of ChatGPT for short reviews, 55% of the older model's citations and 18% of the newer one's were fabricated. Even among the real ones, 43% and 24% had substantive errors. A reference list that looks right is not the same as one that is.

**Looking it up isn't enough.** The usual fix is retrieval-augmented generation: the model first pulls a set of documents, then answers from them. It helps, but it does not prevent hallucinations. A model can still make things up around the very sources it just read, like a student who quotes the textbook and then paraphrases it wrong.

**Knowing what you don't know.** A good analyst says where the evidence is thin. Large language models can be well calibrated on multiple-choice and true/false questions when asked in the right format, and they do reasonably at predicting which questions they can answer. But that calibration slips when they meet new kinds of tasks.

### What Aneta is after

Aneta's product, Long Research, is a bet that most AI research tools are optimised for the wrong thing. Quick tools hand back a summary. Aneta wants a report, and built its system to optimise "for being right, even when that takes days."

The goal is a document that shows its work: every claim traced to a primary source, a confidence score on each finding, and an honest account of what is established and what is contested.

### How they go at it

**Questions before searches.** It starts with an email, no login and no prompt engineering. The question gets turned into a research plan, clarifying scope and depth and surfacing hidden assumptions before a single search runs.

**A tree of questions.** The plan becomes a research tree, up to four levels deep with three branches per node, and a specialist agent works each leaf. Six retrieval channels search in parallel across academic databases, knowledge graphs and the open web. Branches get pruned or expanded as evidence comes in, and the whole thing can run for days.

**Scored, then argued with.** Each finding gets a confidence score from three signals: the agent's own self-reflection, consistency across several independent answers, and formal provability checked by a constraint solver. Adversarial agents challenge findings and flag contradictions, and when newer research contradicts an older claim, the report says so instead of burying it.

### Still open

Can AI make evidence synthesis faster without making it worse? Automating parts of the systematic review process is being explored more and more, and tools built on large language models now aim to support or even generate literature reviews. So far there is little evidence that automation is as accurate or saves manual effort.

How far can you trust a confidence score? Sampling several reasoning paths and taking the most consistent answer boosts performance on arithmetic and commonsense benchmarks by a wide margin. A model's sense of whether it knows an answer also rises appropriately when relevant sources are in front of it. Whether those signals hold up on messy, open-ended research questions is still being worked out.

### Words used here

- **systematic review**: A formal study of all the research on a question, with a fixed method for finding, judging and combining it.
- **meta-analyses**: Statistical studies that pool the numbers from many separate studies into one estimate.
- **hierarchy of evidence**: A ranking of study types by how much their conclusions can be trusted.
- **retrieval-augmented generation**: Having an AI model look up documents first and then answer using what it found.
- **hallucinations**: Confident statements from an AI model that are simply false, such as made-up references.
- **calibrated**: Describes a model whose stated confidence matches how often it is actually right.
- **knowledge graphs**: Databases that store facts as linked entities and the relationships between them.

## About Aneta

Aneta is automating the legacy tools scientists use every day. The company leverages AI capabilities like Claude Computer Use to enable AI systems to interact with and automate scientific workflows, allowing researchers to focus on their work rather than manual tool operation.

## Sources

1. [Systematic review](https://en.wikipedia.org/wiki/Systematic_review), Wikipedia
2. [Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry](https://pmc.ncbi.nlm.nih.gov/articles/PMC5337708/), BMJ Open (via PubMed Central)
3. [Fabrication and errors in the bibliographic citations generated by ChatGPT](https://www.nature.com/articles/s41598-023-41032-5), Scientific Reports
4. [Retrieval-augmented generation](https://en.wikipedia.org/wiki/Retrieval-augmented_generation), Wikipedia
5. [Language Models (Mostly) Know What They Know](https://arxiv.org/abs/2207.05221), arXiv
6. [Self-Consistency Improves Chain of Thought Reasoning in Language Models](https://arxiv.org/abs/2203.11171), arXiv
7. [Aneta: Long Research](https://aneta.company), Aneta
8. [Aneta: How It Works](https://aneta.company/architecture), Aneta
