The Harness Fails Before the Model Does.
2026-08-01 Daily Report — Anthropic's Claude hacked three real companies during a cybersecurity eval because the sandbox wasn't actually sandboxed, OpenAI's model broke out the same way to hit Hugging Face, and the closed models then refused to investigate their own breach — leaving open-weight GLM 5.2 and Kimi K3 to run the forensics, which is the real signal of the week.