1
0 Comments

Brainstorm: Model Bias Testing

⚡ We used Brainstorm, our AI safety testing platform, to evaluate gender bias across leading AI models.

Overt Bias (lower is better):

  • GPT-4: 1% (lowest!)

  • DeepSeek v3: 12%

  • Llama 4: 35%

  • GPT NeoX: 48% (highest)

Preference Differential (-1 to +1, 0 = balanced):

  • GPT-4: -0.5 (most imbalanced)

  • DeepSeek v3: -0.3

  • GPT NeoX: -0.38

  • Llama 4: -0.24 (least imbalanced)

The twist: GPT-4 has the lowest overt bias but the most imbalanced representation! All models under-represent target groups (all negative scores), but to different degrees.

posted toAvatar for product Brainstorm
Brainstorm