Anthropic published results from a lab, not a simulation. External evaluators produced and tested the model's designs. The share that worked is between 22 and 35 percent, against a usual 10 to 15.
- Claude designed small proteins against 15 targets and succeeded against 14 of them.
- The designs were produced and tested by external labs, not scored on a computer.
- On chemical analysis, another model did work that needs a specialist in 23 and 19 minutes.
The number I am watching is not 14 out of 15. It is the share of designs that worked: between 22 and 35 percent, against a usual 10 to 15.
That is where it sits, because the first number tells you whether it works at all, and the second tells you how much work you throw away.
And the important part: this is not a score from a computer. The designs went out to external labs, were manufactured and were physically tested.
Where the caveat is
The numbers are Anthropic's, and the targets are well known ones used in design competitions. Two were picked as new on purpose, so they would not have been seen during training. That is an honest move, but the check is still theirs.
The other thing they say plainly: life science tasks are blocked in their most capable model, and access for scientists is still being prepared. So they are publishing a result from a tool you cannot touch.
The second experiment tells me more than the first. A chemist gets raw data from a lab and has to read it. Specialist work, dull and slow. The model did it in about twenty minutes from two lines of instruction, with the same answer.
That is the shape in which these things actually enter work. Not a brilliant discovery, but waiting time removed from the middle of a process.
If you are in a trade where you wait on a specialist for something routine, the next two years will be interesting for you.