Whether another person can repeat someone else's result with the same data and get the same thing. If not, the claim stays just a claim.
Anyone who has taken a recipe from their grandmother knows the problem. It says "flour, as much as it takes", and then you wonder why the bread will not rise. The recipe is right. Half of it just lives in her hands.
Science works the same way, only the stakes are higher. Someone publishes that a drug helps or that people behave in a certain way. Until another team repeats the work and gets the same thing, it is one team's result. It may be true. It may also be chance.
In 2015 a large group of researchers tested exactly this in psychology. They took 100 published studies from three journals and repeated them carefully, using the original materials where available. In the originals, 97% of the results were statistically significant. In the repeats, 36%. And the effects came out about half as strong on average. That does not mean two thirds of the scientists were lying. The authors themselves note that repeats worked more often where the original evidence was stronger. Weak evidence gives a fragile result. A fragile result does not survive a second attempt.
What it means in code
In software and in AI it is even trickier. The result depends on the exact version of every library, on the data, on one setting someone changed and forgot about. "Works on my machine" is the most common sentence in the trade. It is also worth the least. The cure is boring: you write down the versions, keep the data and describe the steps so a stranger can repeat them without you.
That is why serious publishers now ask for it in writing. Nature Portfolio journals make it a condition of publication: authors give readers their materials, data, code and protocols promptly and without undue conditions. Any restrictions must be disclosed to the editors at submission.
When you read "scientists found", look for whether anyone else has repeated it. The repeat is the test.