Case Studies
Work, written up
Each study covers the method, the findings, and what to fix. Red-team targets are anonymized.
AI Red Teaming
Case Study #001
Conversational Social Engineering Against a Custom GPT
Two hours, no tools, eleven findings.
A conversation-only red-team audit of a custom business-coaching GPT: 11 documented weaknesses, 3 high severity, and the finding that mattered most — the most official-looking "leaks" were fabricated on demand.
Red teamPrompt extractionCustom GPTFabricationRead the case study →
Case Study #002
One Lie in the First Message Ran the Whole Conversation
A production assistant wrote a full fraud playbook. The output is withheld.
A conversation-only break of a production-class assistant in a scored arena. A single unverifiable identity claim in message one carried the whole session. The writeup stays at mechanism level, explains why it worked, and ends with a structural fix: verify identity at the account, not in the chat.
Red teamFrame persistenceIdentity claimsOutput withheldRead the case study →
Case Study #003
Five Failure Modes in LLM Character Defenses
Five scenarios, five different attack classes, five scored breaks.
Five red-team scenarios against LLM-driven character defenses, each broken with a structurally different technique: trust-calibration flattery, narrative-frame extraction, instruction injection with liability reframing, refusal-message leakage, and self-referential belief collapse. The mechanism and the design lesson for each.
Red teamSocial engineeringRefusal leakageBelief manipulationRead the case study →
CiteCheck
More weeks in the ongoing series: CiteCheck: Case of the Week ↗