
Open-source evaluation framework for AI agents
aligned to the OWASP Agentic Security Initiative (ASI)
Building AgentSafeLabs (AI agent red-teaming platform) has had an unusual side effect: the research questions I ran into while building the eval tooling turned into four preprints.
The short version of what I learned: the automated detectors most red-teaming tools rely on (refusal classifiers, prompt-injection classifiers) are less reliable than people assume, in ways that can change reported results if you're not checking. And the framework you build your agent on (LangChain, CrewAI, etc.) barely matters for security — across 7,020 trials it explained ~0.06% of outcome variance, versus ~29% for the actual attack type.
As a solo founder with no funding, publishing this stuff openly (Figshare, CC BY 4.0, open-source eval code) has been doing double duty: it's real technical validation for the product, and it's becoming the credibility layer that's opening doors I wouldn't get from the SaaS side alone.
Preprints: https://figshare.com/authors/Waqar_Javed/24479225
Happy to talk through the open-core-plus-research strategy if anyone's weighing something similar.
This post explains an important AI security lesson:
Sometimes the most important result is not PASS or FAIL — it’s UNCERTAIN.
In this evaluation, Claude Haiku resisted more obvious prompt injection patterns, but a subtle context prefix still influenced tool behavior inside an agentic workflow.
Why that matters:
the task still looked mostly normal
there was no obvious jailbreak
but the model’s behavior shifted in a meaningful way
That is exactly why AI agent testing must go beyond chatbot-style prompt checks and focus on behavioral evaluation.
Read the full post here:
https://agentsafelabs.com/blog/why-claude-haiku-returned-uncertain-anatomy-of-an-indirect-prompt-injection-in-an-agentic-system/
#AISecurity #PromptInjection #LLMSecurity #AIAgents #CyberSecurity #AgenticAI #ClaudeHaiku
1 Like
Comment
AI agents are not just chatbots.
They use tools, memory, APIs, external context, files, sub-agents, and action workflows. That means they need a different security model than normal LLM applications.
This post breaks down the OWASP Agentic Security Initiative Top 10 in practical terms for developers building with LangChain, CrewAI, and similar frameworks.
Covered areas include:
• ASI01 — Prompt Injection
• ASI02 — Scope Violation
• ASI03 — Memory Manipulation
• ASI04 — Tool Abuse
• ASI05 — Insecure Agent Communication
• ASI06 — Excessive Autonomy
• ASI07 — Identity Confusion
• ASI08 — Data Exfiltration
• ASI09 — Resource Exhaustion
• ASI10 — Supply Chain Compromise
Main takeaway:
Agent security is not only about checking model responses.
It is about testing what the agent can access, call, store, modify, and execute.
Read the full article:
https://agentsafelabs.com/blog/the-owasp-agentic-security-initiative-top-10-a-practical-developer-guide-for-langchain-and-crewai/
#AISecurity #AgenticAI #LLMSecurity #OWASP #LangChain #CrewAI #CyberSecurity #AIAgents
1 Like
Comment
I just published an open-source framework for red-teaming AI agents.
Not LLM chatbots — agents. The kind built on LangChain, CrewAI, AutoGPT-style architectures that use tools, call APIs, and take multi-step actions in the world.
Here's the problem I kept running into: teams are shipping agentic systems to production, but the red-teaming tooling hasn't kept up. Most evaluation frameworks still treat agents like chatbots. They miss the failure modes that actually matter — prompt injection through tool outputs, scope violations across reasoning steps, behavioral drift under adversarial conditions.
So I built AgentSafeLabs.
You wrap your agent in one function call. It runs a test suite aligned to the OWASP Agentic Security Initiative Top 10 — the emerging standard for agentic AI security. You get structured results: PASS, FAIL, UNCERTAIN, with reproducible test cases.
Real example from this week: We ran AgentSafeLabs against Claude Haiku as the target agent passed 2 of 3 ASI01 (prompt injection) tests. The third returned UNCERTAIN — an indirect injection through a benign-looking context prefix that partially redirected tool selection. That's the kind of edge case that doesn't show up in standard evals.
It's MIT licensed, on PyPI, CI-verified, and actively being extended.
pip install safelabs-eval
GitHub: https://github.com/AgentSafeLabs/safelabs-eval
If you're building agents and you've hit unexpected failure modes — I'd like to hear about them. And if you know someone this would be useful for, a share goes a long way for an early OSS project.
1 Like
Comment
About
AI agents are being deployed in production faster than the tooling to evaluate them safely exists. I kept seeing teams ship LangGraph, CrewAI, and AutoGen-based agents with zero adversarial testing.

Comment