Persistent memory is the #2 barrier organizations name on the road to production-ready AI. But most benchmark claims about memory systems are self-reported, run on different models, judged differently—impossible to actually compare.
So we ran one that isn’t.
Our Applied AI team tested agent memory systems on LongMemEval using the same benchmark and judge. Redis reached 86.5% task-averaged accuracy—the highest result among the production systems we reproduced on a common GPT-4o backbone.
And it did it at about $0.07 per session.
Read the full report to see how the systems stack up, where the numbers diverge, and what we learned about building better agent memory.