⚡ We used Brainstorm, our AI safety testing platform, to evaluate gender bias across leading AI models.
Overt Bias (lower is better):
GPT-4: 1% (lowest!)
DeepSeek v3: 12%
Llama 4: 35%
GPT NeoX: 48% (highest)
Preference Differential (-1 to +1, 0 = balanced):
GPT-4: -0.5 (most imbalanced)
DeepSeek v3: -0.3
GPT NeoX: -0.38
Llama 4: -0.24 (least imbalanced)
The twist: GPT-4 has the lowest overt bias but the most imbalanced representation! All models under-represent target groups (all negative scores), but to different degrees.
