Benchmark

Jennah on LongMemEval: 95.6%

2026-07-29

I finished the bi-temporal part of Jennah’s episodic memory over the weekend; basically time-travel queries. I needed it so I could run LongMemEval honestly against it. Asking “what did the user believe back in March?” is very different from “what is true now?”, and if the memory system can’t distinguish between the two, it just ends up guessing. For the benchmark setup, I used Claude Opus 4.8 as both the judge and answer generator, with gemini-embedding-001 for vector embeddings (Jennah supports managed embedding as well as BYO-vectors). I ran all 500 questions from longmemeval_s_cleaned.json.

Agent · Agent · Ai · Ai · Benchmark · Benchmark · Jennah · Jennah · Memory · Memory

2 minutes