Watching GPT-5.6 Sol Ultra Write a Chrome Exploit: Exploit Development as We Know It Is Dead
We benchmarked the latest frontier models, GPT-5.6 Sol Medium, Sol Ultra, and Grok 4.5, on their exploit-development capabilities. After processing 2.096 billion tokens, Sol Ultra was the only model to produce a complete exploit chain.
research benchmark