2
0 Comments

Public LLM Bencmarks are BS. You need your own

I've been following a bunch of different benchmarks to help me decide which LLM model should I use for GrandpaCAD. It turns out models are very jagged - meaning they excel at certain task (refactoring a code base for example), and suck at others (should I walk or take a car to a car wash).

I wrote more about it here: https://grandpacad.com/en/blog/public-benchmarks-misled-me-opus-4-7

The point is, build your benchmarks/evals as soon as possible, otherwise you may be under performing or overpaying or both for same results.

posted toAvatar for product GrandpaCAD
GrandpaCAD