1,221 people from the community checked the claims in a third of the conference. In 496 papers at least one claim failed. 266 reproduced fully.
- Between 15 July and 2 August Hugging Face and alphaXiv ran the largest open audit of an ML conference so far.
- 35,908 scientific claims from 2,226 papers were assessed, which is 34 per cent of ICML.
- 51% of papers had at least one verified claim, 23% had at least one falsified.
One thousand two hundred and twenty one people sat down and started checking other people's papers. Not reading them. Running them again and seeing whether the same thing comes out.
The number I stop at is not 23 per cent. It is 266. That is how many papers reproduced fully, out of two thousand two hundred and twenty six. The rest are not fake, most simply cannot be repeated by an outsider with what is available.
That is the crux. In our trade we call this documentation, and anyone who has taken over someone else's project knows what its absence costs. In the science around AI the price is the same, the bill just arrives later.
They are opening the data and the agent traces too. So next time someone will check the check itself. That is how it should be.