A number in the headline doesn't prove intelligence. It only proves memory.
The AI model is the program behind the chats you talk to. Before it reaches you, it has "read" a huge amount of text and drawn from it a way of answering. To check whether such a program is any good, you can't just ask it "are you smart". So researchers put together an exam: a collection of questions in math, logic, reading comprehension, writing. They run the model through them and count how many it got right.
So far this sounds reasonable. The problem is where these questions come from. Most exam collections sit published online, so any researcher can use them. And the model learns from exactly that internet - from articles, forums, books, including the very exam sheets with the answers already in them. When it's later tested on the same questions, it isn't solving the task from scratch. It's simply recognizing something it has already seen.
Imagine reading a headline that says "our assistant scored excellently on the legal knowledge test" and deciding to trust it with an important document at home. The exam result only tells you the model does well on THAT particular exam - not that it understands your real, messy, off-the-textbook case. The two things aren't the same, no matter how close they sound in the headline.
Companies know this very well and still publish benchmark leaderboards, because the trick works as marketing - a number sounds more convincing than an explanation. You see a leaderboard, you see who's "first", and you decide to trust it. Nobody in the headline tells you whether the model had seen the exam beforehand.
Here's what I do instead
Every time I see a headline that says "model X is first on benchmark Y", my first reaction is to skip it. Not because the number is a lie - it usually is true. But because it doesn't answer the question that actually matters to me: will it do the job on MY task. So I never pick a tool by leaderboard. I run it through something real from my own work and see what happens.
My advice to you is the same. If you read that some AI is "the best" on some test, take it as a PR sentence, not a guarantee. The only exam that actually matters is your own task - the email you want to write, the photo you want explained. Try it on that before you believe the headline.