NotTrack AI Open chat

What a score does and does not measure

2026-08-19

Benchmarks measure a fixed set of tasks with a fixed prompt, scored automatically. Your work is a different task, with your own wording, judged by you. Four specific gaps separate the 2, and none of them closes as scores improve.

Scores are the most quotable thing in the category and the least predictive. This is not an accusation of dishonesty; it follows from what a benchmark has to be in order to be a benchmark.

The tasks are not your tasks

Benchmarks favour problems with checkable answers: exams, puzzles, code that runs. Most real work is judged rather than checked, and no benchmark can score judgement automatically. The measurable subset is not a sample of your day.

The prompt is not your prompt

Scores are produced with carefully written prompts, often tuned for the test. You will write your question in 1 line while thinking about something else. The gap between an optimised prompt and a real question is larger than the gap between 2 leading models.

The wrapper is not measured

Benchmarks score the model, and you use a product: instructions, filter, interface, limits. A high-scoring model inside a cautious product produces a worse experience than a lower-scoring model inside a well-built one.

Contamination is a real and quiet problem

Popular test sets leak into training data, which raises scores without raising ability. It is difficult to detect from outside and it affects the most cited benchmarks most, because popularity is what causes it.

What people ask before switching

Are benchmarks useless then?

They are useful for tracking the field over years and for spotting large gaps. They are poor at ranking close competitors, which is exactly what people use them for.

What about blind human preference tests?

Closer to useful, and they measure average preference on average questions. Your questions are not average and yours is the only preference that matters to you.

Should I ignore scores entirely?

Use them to build a shortlist and never to pick from it. Twenty minutes with your own questions decides better than any table.

Skip the table and run your own 3 questions here.

Open the chat

Choosing between AI chats