Built on a larger base model with extended reinforcement learning, the release scores 46.3% on CursorBench 4.0 and 71% on DeepSWE 1.1 while maintaining low refusal rates for legitimate use.
Preview courtesy of Neowin. The full article opens on their site.