← Founder Notes
Archive ·

Researchers found that anthropic, openai, and google encrypt their chain-of-thought blocks with one…

20:17 ISTby Yethikrishna R

researchers found that anthropic, openai, and google encrypt their chain-of-thought blocks with one shared key per provider, so replaying a trace from a strong model into a weaker sibling can recover the hidden reasoning in plaintext. the attack is called a decryption jailbreak and it works across sessions, users, and models. encryption without key separation is just obfuscation with extra steps.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/researchers-found-that-anthropic-openai-and-google-encrypt-Ddbr1DSDu4Y" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Researchers found that anthropic, openai, and google encrypt their chain-of-thought blocks with one…"></iframe>

Original

More notes