ChipokiaTech news, without the noise
← All stories

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below…

Preview courtesy of Simon Willison. The full article opens on their site.

More from Simon Willison