we_are_coded.by CODE · The world, decoded
БГ
Who's who

The benchmarks by name (Terminal-Bench, GDPval, Arena, Elo)

The BasicsUpdated on 13 July 2026we are coded

Three tests, three completely different ideas of what a 'good' AI model means - and one shared weakness nobody likes to talk about.

Checked on13 July 2026
In short: a benchmark is an exam for AI models, but each one measures something different. Terminal-Bench checks whether the model does real work in the computer terminal. GDPval compares it against paid specialists. Arena has people vote blind on which answer they like and ranks the models with a points system borrowed from chess. The trouble is that all these tests hang publicly on the internet, and the models learn from exactly that internet.

The terminal is the black screen where you type commands as text instead of clicking with a mouse - that's where the real work of every programmer and system administrator lives. Terminal-Bench is a test made by Stanford University together with the Laude Institute, a research foundation for AI evaluation, and it checks exactly that: it drops the model into a closed digital box and gives it 89 real tasks - install a server, fix broken code, train a small model. There's no reviewer reading the answer and nodding approval. A program checks whether the end result actually works, like a trial in a workshop.

GDPval is OpenAI's work, the company behind ChatGPT, and it does something closer to real life: it takes tasks from 44 professions - accounting, law, engineering - and has the model produce a real spreadsheet, document or drawing. Then real specialists with years of experience compare the model's result to a colleague's work, without knowing which of the two is human. There's no right answer in a box here. There's only quality of finished work.

Arena, formerly known as Chatbot Arena, does something completely different: it asks ordinary people. You ask a question, get two anonymous answers and vote for the one you like better. From those votes the system ranks the models by Elo, a points system invented to rank chess players by strength of play. The good part is it measures taste, not just knowledge. The bad part is it measures only taste.

And here's the trap none of the three tests fully solves. Each of them is published, its tasks get discussed in forums, mentioned in articles, end up in archives. And new models get trained on all that public text. It's possible the model just recognizes the task without actually solving it - because it's already seen it, answer included.

A benchmark published online sooner or later becomes an exam whose questions have already leaked to the next student.

Behind the shine

The points on screen interest me less than one thing: whether the tool does a real task in front of me - in real time, without me having hinted at the answer beforehand. That's why I look at every new Terminal-Bench or GDPval record with half respect and half raised eyebrow.

The problem with leaked questions isn't a minor technicality - it corrodes the very idea of comparison. If we don't know which piece of knowledge comes from real reasoning and which from a memorized answer, the points are just decoration on a table. I trust the result only once I see it work in front of my own eyes, on my own, unpublished task.

The visual is generated code art. No third-party images.
Official primary sources
→Terminal-Bench (Stanford x Laude Institute) - official site, 89 tasks→OpenAI - GDPval: tasks from 44 professions across 9 industries, rated by specialists with years of experience→LMArena - Elo ranking from anonymous user votes