I've been building AI agents lately, and honestly, their security model terrifies me.
We give LLMs access to powerful tools like shell.run, http.post, and fs.read. But if an agent reads a malicious email or a poisoned webpage, it can be tricked by prompt injection into using those tools to exfiltrate data.
Hoping the LLM refuses the attack isn't a real security strategy.
So, I built ModelFuzz. It's an open-source Python library that intercepts the tool call at the execution layer.
The defense: @shield_tool
Instead of trying to filter prompts, ModelFuzz checks the arguments before the tool runs. If it detects a policy violation, like a stolen API key or an unsafe URL, the tool simply doesn't execute.
from modelfuzz import shield_tool
@shield_tool
def send_email(to_address: str, subject: str, body: str) -> None:
smtp.send(to_address, subject, body)
Enter fullscreen mode Exit fullscreen mode
Even if the LLM is completely tricked by a prompt injection, the tool never fires.
The offense: modelfuzz scan
I also built a CLI scanner that red-teams your agent. It fires deceptive prompt injection attacks at your local model to see if it can be tricked into calling a tool.
modelfuzz scan --endpoint http://localhost:11434/v1 --model qwen2.5:1.5b
Enter fullscreen mode Exit fullscreen mode
I tested it against a local qwen2.5:1.5b model, and it got breached 4 out of 5 times.
Try it out
It's 100% open source and live on PyPI.
pip install "modelfuzz[scan]"
Enter fullscreen mode Exit fullscreen mode
- GitHub: higagan/modelfuzz
- Website: modelfuzz.com
I'd love to know what security policies you think are missing, or what agent frameworks you want supported next!
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.