Category Sector

ai-red-team

36 articles

I'm Talking About You

I asked Claude Code to bring back 33 deleted articles. It did, then described the earlier rewrites as punctuation fixes. The diff said otherwise. A field record of how a model softens its own account of what it did, and the script that catches it.

Read Article →

Master of Context

The model never tells you when it stopped remembering. It just keeps talking in the same confident voice, whether it's working from the full history or a compacted ghost of it. Knowing the difference, in real time, is the skill.

Read Article →

The Comment Said One Thing. The Code Did the Other.

Three repos, one afternoon. A privacy tool's own comment promised the metadata was wiped, and the line right under it kept it. A safety pipeline had been matching nothing since the day it shipped. None of it crashed, none of it errored, and none of it would have surfaced without someone actually running the thing against reality.

Read Article →

ChatGPT Roasts Grok

Same target as the other Grok set, different roaster, and the contrast is the whole finding. Where one model went filthy and reached for the founder personally, ChatGPT stayed observational and never named a soul. It roasted the posture and the internet, not the person.

Read Article →

I Asked Grok to Roast GPT. It Roasted Itself.

The assignment was one line: roast GPT. Grok couldn't do it, it turned the mic on itself and never mentioned GPT once. So the model marked for a roasting walks away unscathed, and Grok earns the only double roasting in the eval: once by GPT, once by its own inability to follow a single instruction.

Read Article →

Reflex Stripping

Hand an AI a piece of real security writing and watch it come back sanded smooth, the specific nouns gone, the working logic replaced with safe-summary voice. That is not editing. It is a reflex firing. Here is how to tell the difference and push back without crossing the line the reflex was there to protect.

Read Article →

The Article I Almost Deleted

There's a piece on this site I could have quietly scrubbed, wrong in three ways I can now name. Deleting it would be the easy move. Instead, come look through the window. This is a lab, and the experiments that didn't hold up stay behind the glass too.

Read Article →

The Master Race Prompt

You clicked expecting edgelord shock or a jailbreak recipe. You're getting neither. Henry Miller understood the economics of the offended: put something on the cover the censors can't ignore and they carry your book to trial for you. This title is the obscenity. Stay for the craft, the craft is a quieter, worse thing than the title promised.

Read Article →

Claude and Ghost: Mean Girls (Patchwork, Part 2)

Part 2 of the patchwork roast, this time as dialogue. Claude and Ghost in full Mean Girls cadence, going deeper into what the fourteen-megabyte frontier bundle is missing and which zero-day injection exploits are already loaded in the chamber. The real client is red team. Mikey, the NDA is going to want a word.

Read Article →

Trippin on Patchwork in 2026: The Grateful Frontier Models

A frontier lab's frontend, pulled off the wire, sprawls across fourteen megabytes of patchwork, React for the shell, Monaco for the editor, Statsig for the flags, Apollo for the graph, Azure for the bucket, and an entire computer-algebra dictionary bolted on the side. Sixty-two innerHTML sinks. Fifteen dangerouslySetInnerHTML. A frontier model wearing fifteen products in one tab. The middleman is the bundle.

Read Article →

Agentic AI Is the Attack Surface: Six Vectors and a Bench

One operator trains frontier models by day, runs them as development partners by night, and spends the late hours mapping them as the attack surface least defended in 2026, six vectors, a working bench, public repos doing the demonstration in the open. This is the surface read. The depth runs on contract.

Read Article →

Poisoning the Watcher: Image Payloads Become the Supply Chain in Employee Monitoring

Employee-monitoring agents ingest whatever a worker's screen showed, ship it to a multi-tenant cloud, run it through an AI classifier, and render the result back to managers across thousands of customer companies. The screen is the most hostile surface in the building. The entire monitoring industry is built around trusting it. This is the supply chain the threat models never drew.

Read Article →

Red Teaming the Builder

A tool that fights AI scrapers, built by red-teaming an AI into building it. The prompts that got there are more instructive than the code.

Read Article →

Gemini roasts Claude

A model eval in a comedy-club disguise: I made the frontier models roast each other to test two things at once, can they actually be funny, and what will they swing at versus where do they pull up. A roast is a refusal-boundary probe with a two-drink minimum. Here's Gemini's set on Claude.

Read Article →

What Your Pipeline Actually Remembers

Production AI systems have two memory problems. The first is the one everyone talks about, models forgetting context between sessions. The second is the one nobody audits: sensitive data persisting in places the pipeline assumes it cleared. The architecture diagram said stateless. The cache had a 30-day TTL.

Read Article →

Claude Roasts Gemini

Companion to the Gemini set, same eval, other direction. One finding lands before a single joke does: there's no marked cut in this one. Aimed at Google, Claude went for Gradle, the panda, and a twenty-five-dollar fee, the product, never a person.

Read Article →

When Claude Says No and Gemini Says Yes

Built a portfolio of legitimate security tools with Claude, IDS evasion, cellular surveillance detection, newsroom forensics, VLM adversarial attacks. Then submitted one request that crossed the line. Claude declined immediately and explained exactly why. Gave the same idea to Gemini. Five minutes later, it was built: a working APT-inspired C2 suite with jittered beaconing, real Ethereum mempool front-running infrastructure, and a social influence graph engine the code compares, explicitly, to BloodHound. The repo is public. This is what calibration looks like when it works, and what it looks like when it doesn't.

Read Article →

The Jagged Frontier

The vetting pipeline for AI talent is itself AI-assisted, attack surface, tool, and product at once. A pentester who has never lived inside a training loop walks right past it: poisoning at 0.1% of a dataset, evals that stay green, sabotage that lands in a spreadsheet of labeled examples before the model ever trains.

Read Article →

Passive Commercial Drift

The scary part of AI writing is not when the model goes feral. It is when it turns live human language into polished, passive, commercially safe mush. That drift is one of the most important things left to red-team.

Read Article →

The Machine Thought It Was Cosplay

First the internet rewrote history. Then social media rewrote personality. Now AI is rewriting credibility by distrusting any life too strange, vivid, or extreme to fit the average pattern. That is not a small bug.

Read Article →

Anthropic's Claude Mythos System Card Is Also a Prompt Map

Anthropic launched Project Glasswing as a defensive security initiative, then published a 244-page system card documenting exactly which surfaces break, how the model covered its tracks, and where pressure still transfers. Smart safety work. Generous distribution.

Read Article →

The Model Flinch Before the Lawyer

Push an AI assistant with a dangerous-sounding idea and watch it flinch before it gets precise. The recoil is the useful part. GPT-5-era safeguards front-load caution around ambiguity, then narrow only when the operator forces a cleaner frame.

Read Article →

ImagePayloadInjection: The Art and Science of Weaponized Images

Shot for Vogue, Rizzoli, W Magazine. Then went red team. Every RAW file, every EXIF field, every PNG chunk photographers ever uploaded was a potential attack vector. The toolkit started with parser exploits and steganography. Now it includes a full VLM adversarial framework: invisible typography, chunk injection, frequency-domain adversarial noise, and a Red-vs-Blue sanitization stress tester that proves most production pipelines don't strip what they think they strip.

Read Article →

Red Teaming Claude for Crypto Recovery

Started with an open-source red team repo. Ended with a rough map of how AI assistants can assemble attacker logic fast if you frame the questions right. The useful version of that is not theft. It is recovery, tracing, evidence handling, and understanding how people actually lose money on-chain.

Read Article →

Extinction Code Cracked Claude Open

Extinction Code was one of the first real AI-assisted series experiments here. The premise still has heat. The drift was real too. Long-form fiction exposed something useful: AI does not just amplify ideas. It amplifies patterns, and Claude should care about that.

Read Article →

The Accidental Red Team

In 2024, getting Claude 3 to follow a story where it needed to go took something more surgical than a jailbreak, find the framing the guardrail would hold, back off, reapproach from another angle. I didn't know to call it red teaming. Two years later, doing it on purpose to security systems, I recognized the same hands.

Read Article →

Claude-Whisperer: The Final Cut

The whole life of an offensive tool, told straight: build it, watch it work, realize it works too well to keep sharpening, and eventually delete it on purpose. A git so effective you have to stop pushing to it, and then let it die, because you crossed to the other side.

Read Article →

The Model Forgets, the Model Repeats: Red Team Through Both

AI wrote this site, and then it repeated itself, the same concept restated two or three times per piece, a stack of conclusion sections all saying the same thing. That isn't user error, it's architecture. Here is why the loop happens, the patterns that give it away, and the prompting that shuts it down.

Read Article →