Agent memory

LongMemEval open model results

Long-term conversational memory benchmark covering information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention.

Scores are from Table 3 of the paper. Both rows use round-level values, V + user-fact key expansion, Stella V5 retrieval, and top-10 recalled items. These are memory-system results, not raw model scores.
SystemModelScoreEvaluation scopeEvidence
V + user-fact memory (top-10)Round values; V + fact keys; Stella V5 retriever; top-10 Llama 3.1 70B InstructOpen weights 68.2% LongMemEval_M Paper reported
V + user-fact memory (top-10)Round values; V + fact keys; Stella V5 retriever; top-10 Llama 3.1 8B InstructOpen weights 57.2% LongMemEval_M Paper reported

Primary sources

Official benchmark Original paper Evaluation repository Data or result file