The case study
We tested Atelos against Claude. The first time, we lost.
We wanted to know one thing. If you work with Atelos for a few weeks, does it remember what you told it better than Claude does? Here is what happened, what we changed, and what happened when we tested it again.
The test
A few weeks of client work, then 30 questions.
Atelos and Claude each started with the same 117 files from a real consulting practice. Then we ran the same six working sessions through each one, about a made-up client project: planning, meeting notes, weekly status updates, and decisions that changed along the way. Some things were written into documents. Some were only ever said in conversation.
Then we asked each one the same 30 questions about the work. The people grading the answers didn't know which tool had written them.
Round one
Claude won, easily.
On anything written in a document, all three tied. The whole gap came from one place: things that were only said in conversation. On those, Atelos scored zero.
Why
Atelos kept every conversation, and never looked at one.
Every conversation was saved in the project, with the missing answers in it. Nothing told Atelos the answers might be there, so it searched the documents, found nothing, and said so.
With decisions that changed, it did worse than miss. The documents still held the original plan, so Atelos reported the old plan as current, with confidence. In one session the client switched from a live data connection to a nightly export. Two sessions later, Atelos still said “live.”
Claude did well for a reason we hadn't expected: it quietly keeps its own notes as you work. So “Claude forgets and Atelos remembers” turned out not to be true.
The fix
Three small changes.
It writes down what was decided.
When a conversation ends, Atelos saves the decisions and facts from it as notes in your folder. You see what it kept, and you can undo any of it.
It looks there before it answers.
Atelos now checks those notes alongside your documents. If neither has the answer, it checks your past conversations before telling you it isn't on record.
It keeps track of what changed.
When a later conversation changes a decision, the old note is marked as replaced. Ask about it and you get the new answer, along with what it replaced.
Round two
A new client, new questions, and the fix held.
A fix that works on questions you've already seen proves very little. So we started over with a different made-up client, new working sessions, and 30 new questions written by someone who never saw the first round. The fix was locked in before the questions were written.
Every decision that changed during the project got the current answer from the fixed Atelos: 5 out of 5.
What it cost
And it had to read far less to get there.
Every time an AI answers, it reads through your files first, and on your own AI key that reading is what you pay for. Atelos knows where things are, so it reads much less.
Atelos with the fix read the least of anything we tested, in both rounds, and less than Atelos did before the fix.
Across the six working sessions themselves, the gap was wider: Claude read about 30 times as much as Atelos. We don't turn this into a dollar figure, because some of what Claude reads is billed at a discount. But less reading means a smaller bill.
Honestly
What this doesn't prove.
- It was one test: one made-up client project, six working sessions and 30 questions. A different kind of work could come out differently.
- Atelos's lead over Claude is real but not huge. Question by question, they tied on 20, Atelos won 8, and Claude won 2.
- Six sessions is not six months. We haven't yet tested what happens after months of notes, and that is our next test.
- We tested the improved Atelos before it was released, and we are running the same test again on the released version.
- Claude ran the way it comes out of the box, plus a set of hand-written rules in one setup. Someone who has tuned Claude for months might do better.
- "Claude" here means Claude working directly in the project folder (Anthropic's Claude Code), on the same AI model as Atelos. Not the Claude chat app.
All three changes are in the latest version of Atelos. The full report, with every question and every score, is available on request.