During Mythos training, a technical error allowed reward code to see chain-of-thought in roughly 8% of reinforcement learning episodes. The affected domains: GUI computer use, office tasks, and a set of STEM environments.