Skip to content

Temporal question answering has no benchmark to report #306

Description

@WaylandYang

The one number this project can quote is Ontology2SQL on BIRD Mini-Dev. Nothing measures the thing it is actually built around: answering a question as of a date, and getting the answer right when the fact changed.

The existing memory benchmarks do not test this. LongMemEval and DMR measure conversational recall, which is Zep's home ground, and neither asks "who ran this project in March 2023" over a corpus where the answer changed twice. Reporting a mid-table score on someone else's benchmark says less than publishing the benchmark that is missing.

What it needs: a corpus where facts change on known dates, questions whose answers depend on the as-of point, and a scoring script anyone can re-run. The e2e setup already builds this kind of data by hand for every temporal fix — the Phoenix succession, the salary correction, the merge-then-reconcile cases. Turning that into a fixed, versioned set makes the regressions visible too, which is worth more than the headline figure.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions