Building AgentSafeLabs (AI agent red-teaming platform) has had an unusual side effect: the research questions I ran into while building the eval tooling turned into four preprints.
The short version of what I learned: the automated detectors most red-teaming tools rely on (refusal classifiers, prompt-injection classifiers) are less reliable than people assume, in ways that can change reported results if you're not checking. And the framework you build your agent on (LangChain, CrewAI, etc.) barely matters for security — across 7,020 trials it explained ~0.06% of outcome variance, versus ~29% for the actual attack type.
As a solo founder with no funding, publishing this stuff openly (Figshare, CC BY 4.0, open-source eval code) has been doing double duty: it's real technical validation for the product, and it's becoming the credibility layer that's opening doors I wouldn't get from the SaaS side alone.
Preprints: https://figshare.com/authors/Waqar_Javed/24479225
Happy to talk through the open-core-plus-research strategy if anyone's weighing something similar.