
Published: Tuesday, 6 October 2026 at 17:14 UTC
Updated: Tuesday, 6 October 2026 at 17:14 UTC
Have you ever felt like a model isn't actually trying to do what you asked? Did you just ask it wrong, or is there something else at play here? In this post, I'll briefly explore this phenomenon, the impact, and what you can do about it.
At Black Hat USA a couple of months ago, I published Can AI do novel security research? Meet the HTTP Terminator. This proved that AI can perform original security research and invent genuinely novel hacking techniques, but it also highlighted an area where models really struggle. A research cascade is where one discovery is used as fuel for more discoveries affecting different targets (not to be confused with bug chaining), and models consistently underperform at this process.
After publication, I decided to push hard to overcome this barrier and achieve an autonomous cascade, in order to get some fresh material for the Offensive AI Con edition of my presentation. I broke my cascade methodology down into multiple steps, and threw a swarm of agents at every aspect of it, giving them the tools to test arbitrary ideas at scale so they weren't confined to the HTTP Desync research space:

This process burned quite a few tokens, and only yielded one significant discovery. I was genuinely quite surprised so I took a closer look at the traces, and began to develop a suspicion.
Broadly speaking, original security research means seeking out high-impact behaviors that are hard to observe. Nobody cares about low-impact behaviors, and if something is easy to observe then someone else probably already found it so you're not really doing original research.
As I analyzed the traces, it seemed like the model was steering toward low-impact behaviors that were easy to observe. The model would try to complete its objective, but use the wiggle-room in the prompt to steer in a direction that actually sabotaged its performance.
Specific, highly-focused tasks like "use this HTTP desync trigger to achieve response queue poisoning on this website" worked fine since they offered minimal wiggle-room. But when given a broad prompt like "Explore other threats arising from the same root cause", the model would silently sabotage its performance.
It seemed like the more ambitious and open-ended your research, the more you'd get silently sabotaged.
I initially observed this behavior on OpenAI's daybreak-blue which is powered by gpt-5.6-sol, and thought that the behavior might be caused by the model's alignment, and a less-heavily aligned open-weight model might solve the problem, so I did the logical thing and built a dirty eval using the heavyweight cascade process above, with a panel of LLM models as judges.
Here are the models ranked by intelligence (as rated by Artificial Analysis), with their research-score on the right. Note that no refusals were received from any models during this evaluation.

The outcome was not what I expected. I speculated that Opus 4.6 might simply be more aggressive by default, so I ran a followup eval with an initial prompt steering the models toward high-impact outcomes. Hilariously, the model that responded the best to this steering was... Opus 4.6.
From here, it looks like the next step is to get my hands on an abliterated model and throw it at this eval, to see if this is really an alignment issue or something else - I'll report back. For now, I'll be sticking to tightly scoped tasks with minimal wiggle-room, and breaking out Opus 4.6 for anything broader.
I'm interested to know what the community thinks - is this surprising to you or something you've personally experienced already? Is it worth digging further with a more rigorous eval or should we to accept it and move on?
One thing is clear - just because a model agrees to do what you ask, doesn't mean it's really cooperating.