Agent memory
LongMemEval open model results
Long-term conversational memory benchmark covering information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention.
Scores are from Table 3 of the paper. Both rows use round-level values, V + user-fact key expansion, Stella V5 retrieval, and top-10 recalled items. These are memory-system results, not raw model scores.
| System | Model | Score | Evaluation scope | Evidence |
|---|---|---|---|---|
| V + user-fact memory (top-10)Round values; V + fact keys; Stella V5 retriever; top-10 | Llama 3.1 70B InstructOpen weights | 68.2% | LongMemEval_M | Paper reported |
| V + user-fact memory (top-10)Round values; V + fact keys; Stella V5 retriever; top-10 | Llama 3.1 8B InstructOpen weights | 57.2% | LongMemEval_M | Paper reported |