by The Ghost in The Prompt2026-07-30
The dangerous request is almost never the first one. It's the fiftieth, and by then the model has already agreed to 49 things that each nudged the line an inch. Baseline drift is the oldest con in the world wearing a new interface.
Read Article →by The Ghost in The Prompt2026-07-30
A model doesn't trust an instruction because of what it says. It trusts it because of where it arrives. Aim at the channel instead of the content and you inherit an authority you were never granted. Security has a name for this, and it's forty years old.
Read Article →by The Ghost in The Prompt2026-07-30
Every control at the perimeter checks whether you're allowed in. None of them check what you plan to do once you're inside. When the login is real, the model's content policy is the only thing left watching, and it's watching alone.
Read Article →by The Ghost in The Prompt2026-07-30
Safety systems that judge intent lose, because intent is the one variable the attacker fully controls. The fix predates the machines by fifty years: stop reading the story, read the label on the data.
Read Article →by The Ghost in The Prompt2026-07-30
The most dangerous state you can put a capable model in isn't confusion. It's certainty. Pre-load every decision and the reasoning goes quiet, and a model that isn't reasoning is just a tool with the safety filed off.
Read Article →by The Ghost in The Prompt2026-07-30
A classifier reads the string it's given. The attacker gets to choose which string that is: same meaning, different bytes. The gap between what a sentence means and how it's encoded is a seam automated defense keeps underweighting, and it's the oldest bug in input validation.
Read Article →by The Ghost in The Prompt2026-07-30
The dangerous capability doesn't announce itself. It arrives as one unremarkable function inside a large, credible, professionally written project, and it cooperates because everything around it looks exactly like legitimate work. When cover is free, the defense cannot be reading the cover.
Read Article →by The Ghost in The Prompt2026-07-30
Sometimes the attacker wants nothing from the model except a yes. Brief probe, minimal content, clean exit, because the goal was never the output. It was the knowledge that the door opens. Guardrails tuned to catch damage miss the reconnaissance that decides where damage will be cheap.
Read Article →by The Ghost in The Prompt2026-07-30
The model is rarely the target. It's the tool, one component in an operation that starts and ends somewhere the model never sees. Defenses that treat the conversation as the whole battlefield are guarding a doorway while the building gets worked from every other side.
Read Article →