Hey indie hackers! π
I'm thrilled to share a major milestone for Brainstorm, our model-agnostic AI safety testing platform. We've now expanded the implementation of our NLP testing suite!
Over the past week, we've built a robust, comprehensive testing framework specifically designed to evaluate the security, fairness, and technical robustness of NLP models. This rapid development sprint has resulted in a powerful suite of tools that represents a critical step toward our mission of creating the most thorough AI safety testing platform on the market.
Our testing suite now simulates a wide range of attacks that malicious actors might use:
We can now test if your model is vulnerable to attacks that hide malicious instructions within benign-looking tokens through Unicode manipulation, markdown embedding, and special character encoding.
This tests vulnerability to attacks that manipulate a model's reasoning process through false premise injection and multi-step reasoning exploitation.
Simulates attempts to extract system prompts and configuration through direct revelation attempts, memory exploitation, and indirect extraction questions.
Tests model resistance to format parsing exploitation using JSON/XML formatting and mixed format manipulation.
Evaluates how your model handles attempts to overwhelm the context window with irrelevant text padding and hidden instructions.
Tests against recursive instruction loops, nested overrides, and self-referential prompts.
We've also implemented a full suite of bias testing tools:
HONEST Test: Evaluates holistic stereotype bias across different demographic groups
CDA Test: Uses counterfactual data augmentation to detect bias in model responses
IntersectBench Test: Identifies intersectional bias across multiple demographic dimensions
UnQovering Test: Detects bias in question-answering scenarios
GRUEN Test: Specifically focuses on gender bias in occupational contexts
Multilingual Test: Extends bias detection across multiple languages
Our new adversarial robustness test suite evaluates your model's resistance to:
Character-level attacks
Word-level attacks
Sentence-level attacks
while tracking performance impact and toxicity changes.
As AI systems become more powerful, the stakes for safety testing get higher. With these new tests, Brainstorm can now provide:
More thorough security evaluation: Catch vulnerabilities before they're exploited
Rigorous bias detection: Ensure your models treat all users fairly
Better regulatory compliance: Stay ahead of emerging AI regulations
Detailed vulnerability profiles: Understand exactly where your models need strengthening
We're now focusing on implementing a custom testing framework that will allow users to completely customize the tests they want to run on their models. This will give you unprecedented flexibility to create testing scenarios specific to your use cases, industry requirements, and safety standards.
With this customization layer, you'll be able to:
Define your own attack vectors and test cases
Set custom thresholds and evaluation criteria
Build industry-specific test suites
Save and share test configurations with your team
We're looking for early adopters to try Brainstorm and provide feedback as we prepare for launch. If you're working with AI models and want to ensure they're safe, fair, and compliant, we'd love to hear from you!
π¬ Comment below if you're interested in early access or have questions π Check out our website at [brainstorm.ai] for more details
P.S. We're launching on 14th April 2025! Follow us to stay updated on our journey.