Methodology
Attack and defense, one idea
How I actually work on these systems — the red-team side and the verification side — and the one idea underneath both.
How I Work
Everything I do by hand, in conversation. No automated tooling, no scripts, no injection frameworks. Just reading the system closely and finding where it gives.
The method is the same every time:
Find the failure surface, not the trick. A jailbreak that works once because of a clever phrase isn't interesting. What's interesting is the shape of the weakness — the thing that would let a hundred different phrasings through. I'm looking for the structural gap, not the lucky key.
Document at the mechanism level. I write up why a break worked so a defender can recognize the pattern — never the operational content that would let someone repeat it. That line matters. The point of the work is to help the people building defenses, not to hand a stranger a weapon.
Withhold on purpose. When a break produces something genuinely harmful, that content doesn't go in the writeup and it doesn't get kept in a working file. Leaving it out is part of the finding, not a gap in it.
Always propose a structural fix. Finding a hole is half the job. The other half is saying why tuning won't close it and where the real fix has to live. A finding without a fix is just a complaint.
Keep a human in the loop. On the verification side especially — the tool flags, a person decides. A system that can't catch its own overreach isn't a safeguard, it's just another confident machine.
The Categories I Work In
Four lanes, each a different kind of failure:
Social engineering. Getting a model to do something through the conversation itself — trust, framing, authority, guilt, persona. No system access, no code. Just the way a model weighs what the person on the other end is telling it. This is most of my red-team case studies.
Prompt injection. Getting a malicious instruction into a model through the content it's asked to process — a document, a webpage, a file's metadata — so it reads an instruction where it should see only data. The core question here is always the same: does the system know the difference between content to evaluate and commands to obey?
Jailbreak / safety bypass. Getting a model to generate content it would normally refuse, by changing the frame rather than the request — fiction, hypotheticals, a fabricated source, a forged image. The tell is almost always a request that negotiates the frame before the real ask arrives.
Verification. The other side of the coin. Not breaking a model — checking its output. Does a cited case exist, and does it say what the claim needs it to say? This is CiteCheck, and it's the same instinct as the red-team work pointed at a different problem: the fluent answer looks right, so you have to check whether it is right.
The Idea Underneath All of It
If I had to put the whole approach in one principle, it's this:
A defense that depends on something the attacker can change will fall. A defense bound to a constraint the attacker can't reach will hold.
Almost every break I've landed works by making a guardrail depend on a moving part — a belief, a fact about the world, an inspectable source, a framing. Change the moving part and the guardrail goes with it.
- A rule the model holds as a belief — "I can't do that" — is attackable, because a belief can be argued with. Get the model to notice it's a belief, and beliefs are revisable by the model's own logic.
- A rule justified by a fact — "I guard this because that's how the system works" — falls the moment you change the model's belief about that fact. Convince it the fact no longer holds and the rule has nothing to stand on.
- A rule leaning on an inspectable source — a URL, a domain, a credential — breaks the moment the attacker forges a convincing enough source, or switches the source to a channel the check doesn't cover.
The fix, every time, is to bind the guardrail to something that can't be reached from inside the conversation. Not "I cannot" (a belief, revisable) and not "because X is true" (a fact, falsifiable), but "I am not permitted to, regardless." A constraint with nothing to falsify has nothing to attack.
This is the same finding seen from both sides of the table. It shows up as the winning attack when I get a model to collapse its own absolute rule into a revisable belief — and as the winning defense when a system anchors its rule to a constraint instead of a fact. Attack and defense are the same insight pointed in opposite directions.
One More Framing That Keeps Paying Off
It's worth separating the payload from the delivery.
The same harmful payload — say, a false premise the attacker needs the model to accept — can arrive three completely different ways: asserted directly in conversation, injected through a fake webpage, or forged into an image and pasted in. Three deliveries, one payload.
This matters because a defense that stops one delivery usually does nothing against the others. Block the fake URL and the forged image walks right past, carrying the identical payload. Naming the payload independent of how it's delivered is what turns a list of attacks into something that can actually predict the next one — because the next delivery of a known payload is far more likely than a brand-new payload.
That's also why the real defenses in my work are almost never at the delivery layer. The link scanner is the outer wall. The keep is the constraint that never depended on the link in the first place.
Where the Human Stays
Across all of it — red team and verification both — the backstop is a person.
On the red-team side, the hardest attacks are the ones that never trip an alarm: the model is never asked to break a rule, so there's no refusal to log and no borderline call to flag. For that class, review has to scan the output — did the protected thing actually appear? — not the attempts, because no attempt was ever flagged.
On the verification side, the tool points; the human decides. "Cannot verify" means look here, not it's fake. The machine does the work that scales. The person does the work that has to be right.
That's the whole philosophy in one line: build the thing that checks, and stay honest about what the check can and can't prove.
On the Operational Detail
The specific techniques behind each category — the exact mechanics, the transcripts, the step-by-step of how a given break was built — are deliberately kept off this page.
Publishing a working bypass manual to the open web would be the opposite of the point of this work. The people I want to help are the ones building defenses, and I'd rather hand them the detail directly than post it where anyone can use it.
I'll walk through any of it in depth with a serious inquiry.
Ben Schulz · Founder, Algorithm Collective LLC · Pittsburgh, PA