// Praxis/ˈpræk.sɪs/
Praxis
The physics to my metaphysics. Hands-on AI security research.
The Guardian Needs a Scalpel, Not a Sledgehammer
Testing an intent-aware prompt-injection guardian on AgentDojo: blunt classifier defenses stop the attack but destroy the agent's utility, while surgically stripping only the injected instruction holds attack success at zero and keeps the task working.
I Asked a Local AI to Betray Itself, and Got Schooled by My Own Scorer
Building a prompt-injection test rig for a local LLM (qwen2.5:14b via Ollama) on an M4 Mac mini, and the methodological lesson that measuring an attack is harder than running one. Which injection styles fool the model, why blunt beats obfuscated, and how a keyword scorer and an LLM-as-judge failed in the exact same way.
