The jailbreak just got 93 percent cheaper. blindbias, out oct 2, cuts the cost of attacking an llm…
the jailbreak just got 93 percent cheaper. blindbias, out oct 2, cuts the cost of attacking an llm by 93 percent without touching model weights or logits.
the guardrail economics just changed.
Context
Verified dates: the paper Controlled Decoding Attacks on Black-Box LLMs, which introduces BlindBias, is arXiv 2609.36956 dated Sep 29, 2026. AI Weekly covered it on Oct 2, 2026, saying researchers from Adobe, Amazon and the University of Southern California built a decoding attack that needs no weights or token probabilities, only sampled text.
AI Weekly reports the attack cuts the average jailbreak query cost from 4,000 to 283 API calls, a 92.9 percent drop, tested on four commercial endpoints: GLM-5, Gemini-3.5-Flash, Qwen3-32B and Kimi-K2.5.
The 93 percent matches the 92.9 percent drop, but it measures API calls per jailbreak, not dollars, and the paper is dated Sep 29, with Oct 2 being the AI Weekly coverage date. No weights or logits matches the paper's black-box setting. A Pith review says it counted 20 of 24 table cells where BlindBias leads. 'The guardrail economics just changed' is the author's opinion; whether guardrail defenses change was not seen in the sources read, so unsupported here, not refuted.
Related work
- AI Weekly: BlindBias attacks black-box LLMs with 92.9% fewer API calls ↗Oct 2, 2026.
- AGI Hunt: Blind Bias jailbreaks text-only LLM APIs ↗Paper summary; names Weak-to-Strong, Emulated Disalignment and JULI as earlier decoding-time attacks.
- TesterArmy e2e ↗Oct 1, 2026.
Watch next
- Find the paper's discussion of defenses against sample-only decoding attacks.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 15:40 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →