It was hard to finish this book. In a post from 12 months ago I mentioned buying the book and that I looked forward to sharing some comments on it. it’s not that the book is poorly written, or not interesting, or badly put together. The book is hard to finish because it’s so terribly depressing.

Stuart Ritchie lays out and documents a number of problems that are well-know to anyone in the “meta-science” space. He also points out repeatedly just how much resistance there is to making any improvement. Institutes seek to protect their reputations, journals decline to address irreproducability studies, the inertia of the status quo is immune to most reforms, and the few technical reforms that slip though are easily gamed by this population of decidedly clever people. Ritchie reinforces some of the core statements underlying this project: Nobody is in charge of all this; The only potential gatekeepers have explicit conflicts of interest; The practicing scientists have every incentive to cheat (and little chance of consequences).

Science Fictions is not a comprehensive survey, or formal survey sample, of scientific malpractice, but it offers up many dozens of examples and analyses. Every example and claim is backed up in the voluminous notes section. There are 80 pages of citations, many with added notes, at the end of the 6×9 inch format hardcover copy we bought.

The statistics reported from the many reproducability, bias, and faulty math studies consistently include from significant pluralities to outright majorities of faulty published papers. This author often claims that the literal majority of all published science is junk — this book makes the claim at least plausible. 

The book is organized exactly as presented on the cover (a good lesson for publishers!). After an introduction on what we Ritchie thinks we ought to think science is, and a chapter on the “replication crisis” which has shown that science falls far short of that image, the following chapters talk in order about fraud, bias, negligence, and hype. The third section covers causes (perverse incentives) and a few proposals for improvement. 

The fraud section of course includes a few shocking examples of a researcher propagating a significant and sustained malfeasance. It is not shocking that some few people might go wrong along the way. The parts of the story that should make one angry is how often parties in a position of potential oversight both neglect that responsibility and then try to cover up the problem or deny it and protect the bad actor. What also might be surprising is that “small scale” fraud might be commonplace, as Ritchie  notes from a number of sources. But science is hard on a good day. Clean data is hard to come by, and easy to fudge “a little” if there’s a “small” hole in a series. Ritchie notes how a few statistical tools have been applied to find data that’s a little “too nice”.

The bias section focuses on the problems with a research being married to a hypothesis. This section is the most clinical, as statistical tests and tests-of-tests are discussed (in a very readable way). The infamous “p-value” and practice of p-hacking are covered in appropriate detail. It’s not news that the popular standard for statistical significance is a very low bar, and it’s easy (and tempting) to cheat one’s way under it. New to this reader was the idea of a study power “funnel plot”, where an even distribution of detected effect sizes is expected, converging on a true value with larger study sizes. This may not be seen when many low-power studies remain unpublished — unless they lend support to a pet or popular thesis.

The hype section is shorter, as it should be. Hype simply has no place in science. Academic departments should not have press departments. Journalists should not be twisting pre-print paper abstracts into shocking headlines. Ritchie covers these issues and includes a few high-profile examples where early results trumpeting really caused some trouble, including exacerbating an enduring distrust of science.

Academic publishers of course could correct any of these problems by demanding uniform quality of submitted work. But why would they? Publishers make money on volume, assert importance of their products by fiat, and institutional buyers of the products seem so far to be entirely insensitive to what their money buys. “Peer review” is supposed to be the gold standard of research paper currency, but reviewers are selected by the publishers, are not paid, and can freely choose how much effort to put into review.

Ritchie details some of the clinical efforts to improve the situation, like rating researchers themselves according to a statistic like “h-index”, and immediately demonstrates how these metrics can be gamed and cheated. Trading citations with buddies directly bumps up the h-index, as does submitting lots of small papers instead of one logically organized larger paper. Quality measures for journals themselves have the same problems, at least.

Ritchie stops short of criticizing the common statistical tools, but points out that education in them is largely lacking and most researchers plug their numbers into online calculators. [We aim to point out here that the entirety of 20th-century Fisherian statistical inference is based on fantastic assumptions – randomness, independence, and normality – that vanish outside of a math department.]

The book was published in July of 2020, so the editing was finished in the throws of a global lockdown but the text includes only the briefest mention of the expected torrent of pandemic-related papers (from people who were locked in home offices).   

A number of potential fixes for science publishing are discussed, some of which are being implemented in select contexts. This part of the discussion reinforces the first thing this author tells anyone at the start of a relevant discussion: No One is In Charge of All This. It’s abominable to suggest that a government body should be deciding what gets studied, but this is a perfectly appropriate place for regulation. An important and rent-seeking and fee-charging industry has been selling crap for generations. The industry can be held accountable.

Stuart Ritchie has presented a studiously assembled and terribly important book. We’re not aware of anything comparable that’s come along since in scope and scale. [We can heartily recommend Bernoulli’s Fallacy from Aubrey Clayton, which shows how Bayesian statistics can substitute the “reasonable” for the “rational (but based on impossible assumptions)”.]