← Founder Notes
Archive

The cheapest model is a trap about a third of the time. a stanford-led study found cheaper ai…

Yethikrishna ROriginal on Threads

the cheapest model is a trap about a third of the time. a stanford-led study found cheaper ai models cost more in 32% of comparisons, once quality, retries and reasoning tokens are priced in, per oct 5.

sticker price is the least reliable part of an llm bill.

Context

A study led by Lingjiao Chen of Stanford, 'The Price Reversal Phenomenon', tested 8 frontier reasoning models across 12 tasks and found the model with the lower listed price cost more in 106 of 336 head-to-head comparisons, about 32%.

Coverage cites a case where Gemini 3 Flash was listed 80% cheaper than GPT-5.4 but cost 38% more in practice. The study attributes this to differences in how many reasoning tokens models use.

How it compares

The paper is on arXiv from March 2026, with a v2 dated May 28. News coverage ran October 5, so 'per Oct 5' is the press date.

The tests price reasoning tokens. Retries and quality adjustments are not part of the 32% in the sources read, so that part of the note is unsupported here, not refuted.

'Sticker price is the least reliable part of an LLM bill' is the author's opinion.

Related work

Watch next

  • Read the paper's method to see exactly which costs are counted.

Sources

  1. arXiv 2603.23971 v2, May 28, 2026arxiv.org
  2. The Clarity, Oct 5, 2026theclarity.today
  3. Implicator, Oct 5, 2026implicator.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 19:09 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-cheapest-model-is-a-trap-about-a-DeRorrBDEtQ" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The cheapest model is a trap about a third of the time. a stanford-led study found cheaper ai…"></iframe>

More notes