Engagements
Prompt injection, agentic tool-poisoning, the jailbreak that survives the guardrail, taken apart on purpose, then answered with tooling shipped in the open.
25+ security tools built and running, a research series that names the abuse patterns before they scale, and a month inside a frontier lab's red team. I know where every door is. I knock on the front one.
Op Specs
- I test frontier and multimodal models for real weaknesses: prompt injection, tool poisoning, and the trust gaps between agents that only show up under load.
- I also build the defense side: deception tokens, egress forensics, and attribution that never mistakes an IP address for a person.
- Architecture reviews for teams building agentic systems, plus working detection tooling you keep after the engagement ends.
The Tools
The attack side. Where the payload lives in the pixels and the parser trusts the wrong byte.
The half most red teamers skip. You don't understand a trap until you've had to set one that works.
The part that stays honest about the difference between a lead and an identity.
Proof
Most recently: a month-long AI red-team engagement on a flagship, newest-generation model release at a leading frontier lab, under NDA, with current commercial vetting passed short of a government clearance. Working the methods labs are defending against right now, not last year's playbook. I can't name it. I can tell you what the work is made of.
A private channel had a device on it that the roster did not account for. I built a canary, a link that previewed as an ordinary shared photo and logged the truth: an unrostered handset, iPhone and Android both present. It was one piece of a real investigation. I handed over the clean timeline and stepped back when the detective did not want outside help. An IP is a lead, not an identity. The tool knows the difference, and so do I.
The arsenal is open on GitHub and the research is open here: the frontier attack surface, the training-pipeline supply chain, SPID and eIDAS threat models, the surveillance stack read from the inside. Written to be checked, not admired.
If you've got a system you're not sure about, or a red-team program that needs someone who's worked both sides of the field, this is where to start. Authorized scope only. I bring the paperwork discipline, not just the exploit.
github.com/ghostintheprompt