The one number this project can quote is Ontology2SQL on BIRD Mini-Dev. Nothing measures the thing it is actually built around: answering a question as of a date, and getting the answer right when the fact changed.
The existing memory benchmarks do not test this. LongMemEval and DMR measure conversational recall, which is Zep's home ground, and neither asks "who ran this project in March 2023" over a corpus where the answer changed twice. Reporting a mid-table score on someone else's benchmark says less than publishing the benchmark that is missing.
What it needs: a corpus where facts change on known dates, questions whose answers depend on the as-of point, and a scoring script anyone can re-run. The e2e setup already builds this kind of data by hand for every temporal fix — the Phoenix succession, the salary correction, the merge-then-reconcile cases. Turning that into a fixed, versioned set makes the regressions visible too, which is worth more than the headline figure.
The one number this project can quote is Ontology2SQL on BIRD Mini-Dev. Nothing measures the thing it is actually built around: answering a question as of a date, and getting the answer right when the fact changed.
The existing memory benchmarks do not test this. LongMemEval and DMR measure conversational recall, which is Zep's home ground, and neither asks "who ran this project in March 2023" over a corpus where the answer changed twice. Reporting a mid-table score on someone else's benchmark says less than publishing the benchmark that is missing.
What it needs: a corpus where facts change on known dates, questions whose answers depend on the as-of point, and a scoring script anyone can re-run. The e2e setup already builds this kind of data by hand for every temporal fix — the Phoenix succession, the salary correction, the merge-then-reconcile cases. Turning that into a fixed, versioned set makes the regressions visible too, which is worth more than the headline figure.