Skip to content
Novapedia

Science

How a scientific claim earns the right to be called true

Scientists do not certify truth in one stroke; they move claims through a rough pipeline of testing, quantifying and replication.

Published 24 July 2026 · 7 min read

Barbara McClintock in her laboratory in 1947, standing beside equipment and notes
Barbara McClintock in her laboratory in 1947, standing beside equipment and notesSmithsonian Institution/Science Service; Restored by Adam Cuerden, Public domain

In summary

Scientific claims become trustworthy by surviving a sequence of checks rather than by winning one decisive experiment. Cases from continental drift to H. pylori, the OPERA neutrino error and the 2015 reproducibility project show both the strength and the limits of that process.

Someone notices something

The pipeline usually starts with an irritation: a fact that does not sit comfortably with the rest. Alfred Wegener, a German meteorologist, looked at the Atlantic coastlines and the fossil record and argued in 1912 that the continents had once been joined. South America and Africa seemed to match. Identical fossil plants and reptiles turned up on landmasses now separated by oceans. Mountain belts lined up across continents. The observations were real enough; what he lacked was a convincing mechanism.

Geologists did not reject continental drift because they were stupid or hidebound. They rejected it because Wegener could not show how granite-rich continents could plough through oceanic crust. On the evidence available in the 1910s and 1920s, that was a serious defect. A good scientific idea has to explain more than an impression. It has to fit with physics, with measurements and with neighbouring fields. Drift therefore spent roughly half a century in the wilderness, not because the data were absent, but because the account was incomplete.

That delay is worth keeping in mind. Science is often pictured as a machine for rapidly converting observations into facts. In practice it is slower and more suspicious. Novel claims are cheap. Plausible claims that can survive contact with existing knowledge are rarer. The first stage, then, is not truth but notice: somebody says that something odd is there.

Somebody tries to break it

Once a claim exists, the next serious move is adversarial. If it can be broken quickly, so much the better. The history of Helicobacter pylori is a tidy example. For much of the twentieth century, stomach ulcers were blamed chiefly on stress, excess acid and lifestyle. In the late 1970s and early 1980s, the Australian clinicians Robin Warren and Barry Marshall noticed curved bacteria in stomach biopsies from patients with gastritis and ulcers. The idea that bacteria could colonise the acidic stomach met open scepticism.

Scepticism was reasonable. Many findings disappear when someone looks harder. Marshall and Warren had to show not merely association but causation. They cultured the organism, linked it to disease and eventually, in 1984, Marshall drank a broth containing H. pylori, developed gastritis and then treated himself with antibiotics. Self-experiment is not a model for modern ethics review, and it was not by itself decisive. But it dramatized a harder point: the old picture was failing. Over time, randomised trials showed that antibiotics could heal ulcers and prevent recurrence far better than acid suppression alone.

Trying to break a claim does not mean simple hostility. It means designing tests that give the idea every chance to fail. Some claims survive this stage and become stronger for it. Others collapse almost at once. In 2011, the OPERA collaboration reported that neutrinos sent from CERN to the Gran Sasso laboratory in Italy appeared to arrive about 60 nanoseconds earlier than light would over the same baseline of roughly 730 kilometres. If true, special relativity was in trouble. Physicists did the proper thing and swarmed the result. Within months, the anomaly was traced to instrumental problems, including a faulty fibre-optic connection and an oscillator issue in the timing system. The claim was withdrawn. That was not science malfunctioning. That was science working under pressure.

The number gets an error bar

A mature claim is not just a story; it comes with a number and a statement of uncertainty. This sounds technical, but it is the hinge between hunch and knowledge. A measurement is only useful if others know how much trust to place in it. How large is the effect? Under what conditions? What are the likely sources of bias? What does the apparatus miss?

Physics is unusually explicit about this. A result may be reported as a value plus or minus a statistical and systematic uncertainty. The first reflects the scatter expected from limited data. The second comes from the instrument and method: calibration, temperature drift, background noise, software choices. In the OPERA case, the lesson was brutal. The result was striking because the quoted uncertainty was small enough to make a 60-nanosecond lead seem meaningful. Once the timing system itself proved unreliable, the dramatic claim evaporated. Precision without trustworthy calibration is an expensive way to be wrong.

Other sciences make the same move in different language. Medicine uses confidence intervals, effect sizes, pre-registered outcomes and increasingly detailed statistical plans. Geology eventually rescued continental drift by quantifying seafloor spreading and palaeomagnetism in the 1950s and 1960s. Magnetic minerals in oceanic crust recorded reversals in Earth’s magnetic field, producing symmetrical stripes on either side of mid-ocean ridges. That was not a suggestive resemblance but a measurable pattern. Harry Hess’s seafloor-spreading proposal and later plate tectonics supplied the mechanism Wegener lacked. The claim had acquired a machine underneath it, and numbers attached to the moving parts.

Error bars are social as well as mathematical

Uncertainty is not only a property of instruments. It is also a property of judgement. Which variables were controlled? Which analyses were tried and discarded? Was the sample representative? Those questions cannot be answered by a decimal place. They depend on reporting standards, data sharing and the willingness of journals to publish awkward detail. A clean graph can hide a messy chain of choices. Scientists know this, which is why methods sections, supplementary files and code repositories matter more than outsiders often suppose.

Other labs try it

Replication is the stage where a finding leaves the care of its parents. A result produced once, by one group, with one set of students, reagents or software, may be promising. It is not yet sturdy. Other laboratories need to reproduce the effect, ideally with independent materials and slightly different methods. If the claim is real, it should not depend on the original team’s touch.

This is where entire fields discover how much they have been relying on fragile results. In 2015, the Open Science Collaboration published a large replication effort in Science, attempting to repeat 100 psychology studies from high-profile journals. The headline number was sobering. Depending on the criterion used, only a minority of effects replicated clearly, and the average effect sizes in the replications were about half those originally reported. That did not mean psychology was fraudulent as a whole. It meant that publication bias, small samples, flexible analysis and preference for surprising positives had produced a literature with more noise than many researchers had admitted.

Other disciplines have had similar reckonings, though the details differ. Particle physics uses large collaborations and blind analyses to reduce some risks, but it can still be bitten by instrument error, as OPERA showed. Biomedicine struggles with cell lines, antibodies and animal models that behave differently across labs. Climate science, by contrast, often relies less on direct replication of a single experiment and more on converging lines of evidence: thermometers, satellites, ocean heat content, ice cores, physics of greenhouse gases and model intercomparison. Replication is not always a photocopy. Sometimes it is independent routes arriving at the same place.

Peer review helps, but it is not the gate at the end

Peer review is often mistaken for the gold seal of truth. It is not. It is a limited pre-publication check carried out by busy experts who usually do not rerun the experiment, inspect the raw instrument or reproduce the statistics from scratch. It catches some errors and nonsense; it misses others. A paper can be peer-reviewed and wrong, as countless literatures show. Replication by outsiders is a sterner test because it exposes the claim to people with less investment in its survival.

It gets folded into textbooks

Consensus arrives late, and it arrives unevenly. Scientists do not hold a vote on truth. A claim becomes settled when it proves useful, repeatedly survives criticism and begins to support other work. At that point it migrates from specialist papers into review articles, clinical guidelines, undergraduate courses and textbook chapters. The shift can be hard to date because it is cultural as much as evidential.

Plate tectonics is a classic case. By the late 1960s, with seafloor spreading, magnetic reversals, global earthquake patterns and subduction zones all fitting together, continental drift had been transformed from a rejected speculation into the core framework of geology. The fifty-year wait looks less mysterious from that vantage. A mechanism had been found, rival observations were incorporated and the theory generated fresh predictions. Once that happens, students are taught the framework first and the old objections later.

The ulcer story followed a different timetable but the same logic. H. pylori moved from heresy to routine clinical practice when evidence accumulated across pathology, microbiology and trials, and when treatment changed outcomes in ordinary hospitals, not just in one Australian clinic. The 2005 Nobel Prize in Physiology or Medicine awarded to Marshall and Warren marked recognition, not proof. Proof, in science, is usually too strong a word outside mathematics. What science offers is a hierarchy of confidence, with some conclusions so well tested that betting against them becomes irrational.

That leaves a final complication. Textbook status does not make a claim sacred. Newtonian mechanics still sits in textbooks because it works extremely well for bridges, planets and artillery, despite being superseded by relativity and quantum mechanics in domains where it fails. Scientific truth is therefore not a medal pinned on forever. It is a standing achieved by the best available explanation after hard use. Some pieces of that explanation are later refined. A few are discarded. The striking thing is not that revision occurs, but that so much survives it.

Key takeaways

  • Scientific claims earn trust by passing through stages: observation, hostile testing, quantified uncertainty, independent replication and eventual integration with other knowledge.
  • The delay between Wegener’s continental drift and plate tectonics shows that good observations are not enough; mechanism and measurement matter.
  • Marshall and Warren’s H. pylori work succeeded because later trials and clinical outcomes backed the idea, not because one dramatic self-experiment settled the matter.
  • The OPERA neutrino episode is a model correction story: a spectacular claim met scrutiny, instrumentation faults were found and the result was withdrawn.
  • The 2015 reproducibility project showed that published, peer-reviewed findings can be less stable than they appear, which is why consensus rests on bodies of evidence, not single papers.

Frequently asked questions

Do scientists ever prove something is true?

Usually not in the mathematical sense. Outside formal logic, science deals in degrees of confidence based on evidence, uncertainty and repeated testing. Some conclusions become extraordinarily secure, such as the existence of atoms or the role of microbes in many diseases, but they remain open in principle to revision if better evidence appears.

Is peer review the same thing as scientific truth?

No. Peer review is a screening process before publication. Reviewers assess whether the methods and claims seem sound enough to enter the literature, but they rarely replicate the work themselves. Many flawed results pass peer review. The stronger test comes later, when other researchers try to reproduce the finding or challenge it with new evidence.

Why can a correct idea take decades to be accepted?

Because evidence has to do more than look suggestive. Wegener’s continental drift was resisted for decades not because the observations were invisible, but because he lacked a convincing physical mechanism and sufficiently precise supporting data. Acceptance often waits until multiple lines of evidence connect and the claim helps explain other phenomena better than its rivals.

What is the difference between reproducibility and replication?

The terms are used differently across fields, but a common distinction is this: reproducibility means getting the same result from the original data and analysis, while replication means collecting new data and seeing whether the result appears again. Reproducibility checks the computation; replication checks whether the phenomenon travels beyond the first study.

If science changes its mind, why trust it?

Because the capacity to change is one of its safeguards. Science is trustworthy not when it never errs, but when it exposes errors and corrects them. The OPERA neutrino claim was dramatic and wrong, yet the correction was public and fast. A method that can reveal its own mistakes is stronger than one that merely defends old answers.

Keep reading on Novapedia

Further reading

Authoritative external sources for readers who want the primary material.

Share this article

Newsletter

One considered article every Sunday

No filler, no tracking pixels. Just the week's best explainer and why it matters.