The context window just stopped being the limit. a memory layer called galahad-kv, out oct 9, holds…
the context window just stopped being the limit. a memory layer called galahad-kv, out oct 9, holds a 50-million-token stream on one h100 while cutting gpu energy up to 12.3x versus recompute.
the whole codebase now fits in a single call.
Context
Verified source: the paper 'Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute' (arXiv 2610.10845) describes a memory layer released as the public package galahad-kv, which saves the key-value state of each block of text and restores it later, byte-exact, without recomputing.
In the paper's test on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100 with Gemma 4 12B and 31B, all 100 probed blocks were restored without recompute at depths from 0 to 50M tokens. A restore was 2.8x to 4.3x faster and used 8.8x to 12.3x less GPU energy than recomputing the block, with GPU memory flat over the whole stream.
The 50-million-token stream, the single H100 and the 12.3x energy figure match the paper; 12.3x is the top of an 8.8x to 12.3x range, so 'up to' is right. The post says 'out oct 9'; the arXiv ID places it in October 2026 but the exact day was not seen, so that date is not seen in the sources read, so unsupported here, not refuted. The comparison is a block restore against recompute, not a whole-codebase single call, so that part is the author's reading.
'the context window just stopped being the limit' is the author's opinion.
Related work
- arXiv: Real Long-Term Memory for AI (2610.10845) ↗The paper.
- PyPI: galahad-kv ↗KV-cache layer for vLLM and SGLang.
- arXiv: Working Around the Compute Ceiling (2609.39358) ↗Earlier Galahad paper.
Watch next
- Check whether the restore results hold on models other than Gemma 4.
Sources
- arXiv, Real Long-Term Memory for AIarxiv.org
- PyPI, galahad-kvpypi.org
- arXiv, Working Around the Compute Ceilingarxiv.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 09:18 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →