Science
The Questions That Tell You Whether a Study Holds Up
Reliable studies are usually less dramatic than headlines suggest, and easier to recognise once you know which questions to ask.

In summary
Reliability in science rests less on a paper’s prestige than on design choices that limit bias and random error. The strongest clues are often plain to see: who was studied, whether outcomes were specified in advance, how large the effect is, and whether independent groups can reproduce it.
How many people, and how were they chosen?
Start with the denominator. A study of 40 volunteers can be perfectly useful for measuring a blood biomarker in controlled conditions; it is usually a poor basis for sweeping claims about public health, memory, diet or cancer risk. Reliability depends not just on the number enrolled but on who they are, how they were recruited, who dropped out, and whether the sample resembles the people the authors want to talk about.
A reliable paper tells you the flow of participants: screened, excluded, randomised, followed up, analysed. It also gives a reason for the sample size, often through a power calculation. Power is the probability that a study will detect an effect of a given size if that effect is really there. Conventionally researchers aim for 80% or 90% power. Underpowered studies do not merely miss true effects; they also tend to produce unstable, exaggerated estimates when they do find significance.
Selection can distort results before any statistics are run. Volunteers recruited through online adverts are not a random slice of humanity. Hospital patients at a tertiary centre are often sicker than typical patients. A treatment that looks good in a tightly selected trial may perform less well in routine care.
A good example: the Women’s Health Initiative corrected a dangerous assumption
For years, observational studies suggested that hormone replacement therapy might protect postmenopausal women against heart disease. Those studies were large, but women who chose the therapy also differed in income, education, healthcare access and other habits. The Women’s Health Initiative randomised participants instead. More than 16,000 women were enrolled in the combined oestrogen-progestin trial, and in 2002 the trial reported increased risks of breast cancer, stroke and thromboembolic disease, overturning a confident clinical belief. The lesson was not that large observational studies are worthless. It was that selection effects can make them look more certain than they are.
A bad example: small psychology studies in the priming boom
A great deal of social priming research in the 1990s and 2000s relied on tiny samples, often undergraduate students in groups of a few dozen. Some famous effects could not be reproduced when larger teams repeated the work with more participants and tighter protocols. The problem was not psychology alone; many fields tolerated small, noisy studies for too long. But priming became a clear case of what low power looks like in practice: unstable effects, broad claims, and replication failures.
Was anyone blinded, and what was the comparison?
The next question is what the treatment was compared against, and who knew what. Randomisation helps distribute known and unknown differences between groups. Blinding helps stop expectations from leaking into behaviour, assessment and reporting. In a double-blind trial, neither participants nor investigators know who gets the active treatment. In a single-blind design, one side does. Some interventions cannot be fully blinded, such as surgery or lifestyle coaching, but then the paper should explain how bias was reduced in other ways.
Choice of comparator matters. A new drug tested against placebo may beat doing nothing while still being no better than the standard treatment already on the market. Surrogate endpoints deserve caution too. Lowering a lab value is not identical to making patients live longer or feel better. Reliable studies prefer hard clinical outcomes where possible: death, fracture, hospitalisation, objectively verified infection.
A good example: large blinded vaccine trials during Covid-19
The pivotal mRNA vaccine trials in 2020 randomised tens of thousands of participants and used placebo controls with observer blinding. The primary outcomes were prespecified symptomatic Covid-19 cases confirmed by testing. No trial is perfect, and later questions arose about duration of protection, rare adverse events and generalisability to some groups. Still, as trial architecture goes, these were strong studies: large samples, clear comparators, planned analyses and clinically meaningful endpoints.
A bad example: Vioxx and the hazards of selective emphasis
Merck’s painkiller rofecoxib, sold as Vioxx, was approved in 1999 and later withdrawn in 2004 after evidence of increased cardiovascular risk. One crucial episode was the VIGOR trial, published in 2000 in the New England Journal of Medicine, which compared rofecoxib with naproxen. The trial was not designed around a placebo and involved complex interpretation about whether naproxen was protective or rofecoxib harmful. Subsequent scrutiny, including concerns about how cardiovascular events were reported and framed, became a textbook case in how design, comparator choice and presentation can shape the apparent balance of benefit and harm.
Was the outcome decided before the data, and do the numbers mean much?
A common way to make weak evidence look persuasive is to decide what matters after looking at the data. Researchers test several outcomes, slice the sample into subgroups, adjust the model this way and that, and then report the most flattering result. This is sometimes called p-hacking. Preregistration is the main institutional fix. Before data are analysed, the investigators record the primary outcome, secondary outcomes, sample size, exclusions and analysis plan in a public registry or time-stamped protocol.
Then come the statistics. A p-value is the probability of obtaining data at least as extreme as those observed if the null hypothesis were true. It is not the probability that the claim itself is true. A p-value just under 0.05 can arise from a modest real effect, from chance, or from a great many researcher degrees of freedom. That is why effect size and confidence intervals matter. The effect size tells you how big the difference is; the confidence interval gives a range of values consistent with the data, usually at 95%. A tiny but statistically significant effect in a huge sample may be trivial in practice. A wide confidence interval signals uncertainty even when the p-value passes a threshold.
Relative risk and absolute risk are where headlines most often mislead. Suppose a drug cuts the rate of a bad outcome from 2 in 1,000 to 1 in 1,000 over a year. That is a 50% relative risk reduction, which sounds dramatic. The absolute risk reduction is 1 in 1,000, or 0.1 percentage points. The number needed to treat is 1,000: on average, 1,000 people would need the drug for one year to prevent one event. Whether that is worthwhile depends on cost, side effects and the severity of the event. Reliable reporting gives both absolute and relative figures.
A good example: trial registries and prespecified endpoints became standard for a reason
ClinicalTrials.gov, launched by the US National Library of Medicine, did not solve bias, but it made one kind easier to spot. When a paper’s published primary outcome differs from the registered one, readers can see the switch. Campaigns such as AllTrials pushed journals and funders to treat registration and protocol transparency as basic hygiene. The improvement has been uneven, but the principle is now widely accepted across medicine: decide first, analyse second.
A bad example: sugar-funded reviews obscured the question being asked
In 2016, Kearns and colleagues reported in JAMA Internal Medicine that the Sugar Research Foundation had funded a 1967 review in the New England Journal of Medicine that downplayed evidence linking sugar to coronary heart disease and steered attention towards saturated fat instead. Those old narrative reviews were not systematic reviews by modern standards, and the problem was not p-values so much as selective framing. But the case remains instructive. When outcomes, inclusion criteria or comparisons are not specified in advance, influential summaries can become vehicles for preference rather than neutral appraisal.
Who paid, who analysed, and can another team inspect the workings?
Funding source is not a magic detector, but it deserves attention. Industry-funded trials are not automatically wrong; public or charity-funded research is not automatically clean. The more useful question is whether incentives and control ran in the same direction. Did the sponsor help write the manuscript? Who held the data? Was the analysis conducted independently? Are the code, protocol and anonymised data available, at least in principle, for reanalysis?
Conflicts of interest can shape a paper long before any fraud. They influence which questions get asked, which comparators look convenient and how uncertain results are narrated. Transparency helps, but disclosure alone is not enough. Strong studies build independence into the process, through data monitoring committees, external statisticians and publication commitments.
A good example: Cochrane’s method was built to reduce discretion
Cochrane reviews are not infallible, and some have been revised or disputed, but the organisation’s influence comes from method rather than prestige. Reviewers are expected to use prespecified protocols, systematic searches, risk-of-bias assessment and explicit rules for combining studies. A well-conducted Cochrane review usually gives a clearer sense of where evidence is firm, where it is thin and where heterogeneity between studies blocks simple answers.
A bad example: when access to the full data is constrained
Several pharmaceutical controversies, Vioxx among them, turned on the fact that outsiders did not initially have a full view of patient-level data or adverse-event coding. This is not unique to industry. Academic groups also sit on data, fail to share code, or publish methods too sketchily for others to reproduce the analysis. A paper that cannot be audited asks for trust it has not earned.
Has it been replicated, or is the paper standing alone?
Science rarely turns on one study, however elegant. A reliable finding survives different teams, samples, instruments and settings. Replication does not mean every repetition lands on the same number. It means the broad pattern reappears often enough, under well-specified conditions, that confidence grows. Meta-analysis can help by pooling data, but only if the underlying studies are comparable and not all biased in the same direction.
This is where fields separate sturdy knowledge from exciting noise. Early studies often report large effects because only striking results are published. As more evidence accumulates, estimates usually shrink. That is not failure. It is self-correction doing its work.
A good example: aspirin’s cardiovascular benefits emerged across many trials, with boundaries
The evidence that aspirin can reduce recurrent vascular events in people with established cardiovascular disease rests on many randomised trials and large overviews, not a single dramatic paper. At the same time, later work showed a more mixed picture for primary prevention in people without prior events, because bleeding risks offset smaller benefits. The result is a reliable, nuanced conclusion rather than a universal one.
A bad example: a lone spectacular result with no durable follow-up
The 1998 Lancet paper by Andrew Wakefield and colleagues linking the MMR vaccine to autism was tiny, uncontrolled and methodologically frail from the start. Larger epidemiological studies in several countries failed to support the claim, investigations uncovered ethical and reporting problems, and the paper was retracted. It remains the clearest modern lesson in why dramatic, isolated findings should not outrun the total evidence.
Key takeaways
- Reliability starts with design: representative sampling, adequate power, randomisation and blinding do more for truth than impressive prose or a famous journal.
- A p-value below 0.05 is not a verdict; effect size, 95% confidence intervals and absolute risk often matter more than nominal statistical significance.
- Preregistration and protocol transparency make outcome switching and p-hacking easier to detect, especially in trials and systematic reviews.
- Conflicts of interest are most revealing when they intersect with control over data, analysis and publication, not merely when funding is disclosed.
- Single studies are provisional. Confidence rises when independent teams reproduce the result or when systematic reviews show consistent evidence across settings.
Frequently asked questions
Is a peer-reviewed study automatically reliable?
No. Peer review can catch obvious flaws, unclear methods and overstatement, but it is not an audit and rarely involves re-running the analysis from raw data. Weak studies are published in good journals, and strong studies appear in modest ones. Reliability depends more on design, transparency and replication than on the peer-review label alone.
What is the easiest red flag for a non-specialist to spot?
A striking claim from a small study with vague methods and no preregistration is the classic warning sign. Other quick checks help: was there a control group, were participants randomised, did the paper report absolute risks, and do the confidence intervals look wide? If the headline is loud and the methods are thin, caution is usually justified.
Why do scientists still use p-values if they are so often misunderstood?
Because they can be useful when treated as one clue among several. A p-value summarises how surprising the data would be under a null model, and that has some value. Trouble starts when it is treated as a pass-fail switch. Most statisticians now favour combining it with effect sizes, confidence intervals, prior evidence and transparency about how the analysis was chosen.
Are observational studies less trustworthy than randomised trials?
Not always, but they are generally more vulnerable to confounding because people are not assigned exposures at random. Observational studies are indispensable when trials would be unethical or impractical, as with smoking or many environmental hazards. The key is whether the design and analysis plausibly address obvious alternative explanations.
What does preregistration actually stop?
It does not stop error or misconduct by itself, but it limits a common source of false positives: deciding after the fact which outcome, subgroup or model looks most exciting. By fixing the primary question and analysis plan in advance, preregistration makes exploratory work easier to separate from confirmatory evidence.
Keep reading on Novapedia
- How do scientists know something is true?
for the broader logic by which evidence accumulates beyond a single paper
- Precision, accuracy and uncertainty: what makes a measurement trustworthy
for the statistical foundations behind error bars, bias and uncertainty
- The difference between a hypothesis, a scientific theory, and a scientific law
for the terms that are often confused when results move from early claim to established explanation
Further reading
Authoritative external sources for readers who want the primary material.
- ClinicalTrials.gov — U.S. National Library of Medicine
- Cochrane — Cochrane
- Vioxx (rofecoxib) questions and answers — U.S. Food and Drug Administration
- Cochrane Handbook for Systematic Reviews of Interventions — Cochrane
- CONSORT — CONSORT Group
