1
0 Comments

Anthropic disclosed a training error in Mythos that nobody is really discussing — reward code saw chain-of-thought in 8% of RL episodes.

https://www.revolutioninai.com/2026/04/claude-mythos-training-error-chain-of-thought-capability-jump.html

During Mythos training, a technical error allowed reward code to see chain-of-thought in roughly 8% of reinforcement learning episodes. The affected domains: GUI computer use, office tasks, and a set of STEM environments.

submitted this linkon April 11, 2026