Agent memory finally has a leaderboard and a new leader. past.dev's frontier memory, out oct 6,…
agent memory finally has a leaderboard and a new leader. past.dev's frontier memory, out oct 6, took the top spot on beam, the largest public memory benchmark, where the previous best sat at 68.0%.
remembering is becoming a measurable model skill.
Context
past.dev's release of October 6, 2026 says its memory API scores 85.03% on BEAM at 10 million tokens against a previous best of 68.0%. It lists 92.08% at 100K tokens, 89.63% at 500K and 90.65% at 1M, against previous best published results of 76.9%, 71.1% and 75.0%.
The release says the previous best is the best competing score published as of October 2, 2026, counting only runs on every question at that size with a public, reproducible harness. It says past.dev opened its evaluation harness. past.dev's page describes BEAM as Beyond a Million Tokens (Tavakoli et al., ICLR 2026), 2,000 questions over 100 conversations.
The 68.0% previous best and the October 6 date match the release. The #1 claim is past.dev's own, made on its own page and release, and the previous best is defined by its own inclusion rules, so it is a vendor result, not an independent ranking.
'The largest public memory benchmark' is past.dev's wording. 'Took the top spot' is accurate to the release, which says it leads at every history size where results were published.
'Remembering is becoming a measurable model skill' is the author's line. The note's comparison is to 68.0%, which applies at 10M tokens only.
Related work
- past.dev Introduces Frontier Memory for AI Agents: #1 on BEAM (PR Newswire, October 6, 2026) ↗Source for the scores, the previous best figures and the comparison rules.
- BEAM benchmark results (past.dev) ↗past.dev's BEAM results page with the benchmark description.
Watch next
- Run or find an independent reproduction using the opened harness. Read the BEAM paper for how the questions are scored.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 10:14 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →