The replication crisis threatens economic policy
At its best, the scientific method is one of the most reliable tools humans have devised for generating knowledge. It has earned the authority and trust it commands because it is built to correct itself continuously. A claim, or “hypothesis,” is proposed, then tested. If it survives the initial test, it invites others to verify or overturn it. That second part, the replication process, is what separates reliable knowledge from mere ideology or anecdote. It also protects those who consume this knowledge from the individual biases of the original researchers, from faulty measurements and skewed calculations, or from ordinary honest mistakes.
This is how science is supposed to work. That ideal has broken down across much of the social sciences. What began as a quiet concern in psychology has become a documented crisis that extends deep into economics – the discipline that most directly shapes tax codes, labor rules, welfare programs, stimulus packages and regulatory regimes. When the studies used to justify those policies cannot be replicated, the costs are not merely academic. Taxpayer money is wasted, public trust erodes and bad ideas gain the false authority of “science.”
‘Extraordinary claims require extraordinary evidence’
The replication crisis refers to the repeated finding that a large share of published results fail when independent researchers attempt to reproduce them. Effect sizes (or measured impact) shrink, statistical significance disappears or the originally identified pattern simply does not reappear when the test is repeated by others.
The problem is not limited to fraud, though fraud does occur. More common practices include manipulating data until a non-significant result becomes statistically significant (p-hacking or data dredging), selectively reporting positive results while ignoring conflicting ones and simply using small sample sizes. The crisis is fueled by the intense career pressure academics face to publish striking findings. Journals and funders reward novelty and positive results far more than accurate ones. Replication efforts are rewarded even less, which means the incentive to check other researchers’ work rather than publish one’s own is weak.
Over the past couple of decades, the crisis has become impossible to ignore in psychology. In 2015, the Open Science Collaboration, involving hundreds of researchers, attempted to replicate 100 studies from three leading psychology journals. Only about 36 percent of the replications produced statistically significant results in the same direction as the originals, and the average effect size was roughly half of what had been claimed.
Earlier warnings existed – like John Ioannidis’ 2005 paper arguing that most published research findings are false – but the 2015 project put hard numbers on the table and forced the field to confront them.
Psychology produced some of the most striking examples. Consider Dan Ariely, the behavioral economist whose best-selling books and TED talks made him a public face of research on dishonesty and irrationality. A 2012 paper he coauthored claimed that simply asking people to sign an honesty pledge at the top of a form, rather than the bottom, sharply reduced cheating. The finding was influential and served as a flagship example of “nudge theory” – governments and companies experimented with the technique. Subsequent work repeatedly failed to replicate it, and researchers found that signing at the top did not reduce dishonesty at all.
A closer look at one of the key datasets revealed evidence of blatant data fabrication, with thousands of data points either invented or altered. Though Mr. Ariely denied that he personally fabricated the data, the irony was almost literary: A leading expert on the psychology of lying coauthored a study on honesty based on dishonest data. The paper was retracted in 2021, but the damage was already done, and resources had already been wasted.
The ‘dismal science’ earns its name
Economics was slower to face the same scrutiny, yet the numbers that eventually emerged were grim. A 2016 effort led by Colin Camerer and colleagues attempted to replicate 18 experimental economics papers published in the American Economic Review and the Quarterly Journal of Economics. Only 61 percent produced a notable effect in the same direction, and the average replicated effect size was about two-thirds of the original.
United States Federal Reserve researchers Andrew Chang and Phillip Li examined a broader set of papers from well-regarded journals and found they could fully reproduce results for only about half, even with the authors’ assistance. More recently, the SCORE project (Systematizing Confidence in Open Research and Evidence), funded by the Defense Advanced Research Projects Agency and involving hundreds of researchers, examined 164 quantitative papers published between 2009 and 2018 across business, economics, education, political science, psychology and sociology. The replication success rate was 49 percent, and effect sizes in the replications were less than half of those originally reported. Economics did not stand out as more reliable than the other disciplines.
One of the most frequently cited examples of contested findings driving policy is the 1994 study by David Card and Alan Krueger on New Jersey’s minimum-wage increase. Using telephone surveys of fast-food restaurants, they concluded that raising the minimum wage from $4.25 to $5.05 produced no employment losses and possibly even slight gains compared with neighboring Pennsylvania. The finding challenged the textbook prediction that higher mandated wages reduce demand for low-skilled labor and was widely cited by advocates of minimum-wage hikes.
Later work that relied on actual payroll records rather than surveys found the opposite. Employment in New Jersey fell relative to the control group and it became clear that the original survey data contained serious errors. The Card-Krueger result had already entered the policy bloodstream and continues to be invoked decades later even after the data problems became clear. A contested finding remains dangerous when the conclusions are politically attractive.
The stakes are higher in economics than in many other fields because the results are used to design real interventions. Minimum-wage laws, tax-rate changes, stimulus packages, welfare policies and regulatory cost-benefit analyses all rely on research that purports to show the interventions are “evidence-based.” By the time it becomes clear that the evidence did not hold up, it is too late.
Taxpayer money is spent, incentives are distorted and unintended consequences emerge. Misguided policies do more than waste resources: They also undermine confidence in both science and the institutions that claim to follow it. Once the public loses trust in the scientific method, the foundation of rational policy is gone, and political tribalism takes over.
Awareness of the problem is now reaching the highest levels of economic policymaking. Federal Reserve Chair Kevin Warsh has called for a return to first principles, a hard look at unreliable models and improvements to the analytical tools the central bank relies on. As he highlighted in a speech at the Hoover Institution titled “Challenging the Groupthink of the Guild”: “The economic brain trust at the Fed is a precious public resource. It ought not to be squandered by repeating outputs of broken models.”
The good news is that serious efforts are underway to address the problem. The Institute for Replication, founded by economist Abel Brodeur, organizes “replication games” in which teams of researchers systematically check influential papers, document errors, test robustness and attempt full replications. Other projects have improved data and code availability requirements at major journals. These are valuable efforts, but the backlog is enormous: Decades of potentially dubious published work remains unexamined. Even if current reforms succeed in raising the quality of new research, the older literature that still shapes academic textbooks, citations and policy arguments will take years to sort through.
Scenarios
Most likely: Biased research outpaces efforts to contain it
The demand for research that confirms preferred policy conclusions will not disappear. Interested parties – lobbying groups, government agencies and political think tanks – will continue to fund studies that deliver the desired answer. Heavily incentivized researchers will still find ways to produce statistically significant results that fit the required narrative. Replication efforts can expose some of it, but they cannot eliminate the underlying incentives that ensure biased research will continue to be produced faster than it can be audited.
This dynamic is exacerbated by the general public’s lack of statistical literacy. Most people consume science exclusively through media outlets that prioritize sensational headlines and strip away caveats and context. Even the few readers who do seek out the original papers rarely read beyond the executive summary, as they cannot evaluate or even understand the methodology and quality of the research.
Less likely: Reform raises the political cost of fragile research
In a more optimistic but less likely trajectory, sustained replication work and the public attention that retracted papers attract will raise the political cost of relying on cherry-picked or non-replicable studies. This could make it harder for politicians to cite a single convenient paper as “the science” when independent teams have shown the result does not hold.
The transparency bar could be raised, and journals that tolerate unreproducible work could lose prestige. Even partial success in this direction would represent real progress – fewer bad policies enacted based on fragile findings and a modest recovery of public confidence in scientific institutions.
























