Experimental Controls Aren’t What You Think

Why scientists use experimental controls like randomization and blinding – and what these methods are actually designed to do.

Experimental Controls Aren’t What You Think
The hands of a scientist using forceps to extract chorion from Zebrafish embryos. Photo by National Cancer Institute on Unsplash.

Randomization, blinding, and concealment are powerful tools in experimental science, to name only a few. But their purpose is widely misunderstood.

Journalist Gary Taubes, for example, says in the context of nutrition that the goal of science is to “establish reliable knowledge” (12:51); that it uses certain “standards … to establish causality” (12:58). In the context of epidemiology, he says “you’ve got to do randomized controlled trials [to] establish causality.” (46:49)

In an article titled ‘Randomised controlled trials—the gold standard for effectiveness research’, the National Institutes of Health (NIH) claim the purpose of experiment is “to prove causality” (while granting that “no study is likely on its own” to do so). They also claim such trials can “provide true assessment of causality” if “conducted appropriately”. Those are two entirely different claims, as I’ll show below.

There’s a shocking disregard for epistemology in scientific circles, even though we can only know from epistemology what experiment can and can’t do, and whether causality can be ‘established’ or ‘proven’. Without epistemology, we don’t actually know how to do science. The best epistemologist to date is Karl Popper, who explained these matters at great length and in simple terms. But most people have never heard of him, so it can seem like his work was in vain. In the words of Popperian thinker David Deutsch, it’s “as if Popper had never lived.”

Science does not aim to establish knowledge as true or reliable. That’s the opposite of science. In reality, science is a tradition of criticism. It seeks to improve scientific theories rather than to ‘prove’ or ‘establish’ them.

Man is fallible. He makes mistakes, but he can learn from them. He wants to explain the world around him. He forms expectations about it. He doesn’t have direct access to it, but he has indirect access through his rational mind; through fallible theories and observations. Sometimes he notices that his expectations are wrong, showing a gap in his knowledge. He tries to fill it by making guesses that reconcile the difference. Just guesses. The question is, how can he form a rational preference for one guess over another?

The answer is simple. He criticizes his guesses and eliminates errors until hopefully, one guess is left standing that he then has no reason not to adopt. (You can find a formalized, computational implementation of this approach here and here.)

Science is rational thought about nature. It’s an error-correction engine. The purpose of a scientific experiment is to help us form a rational preference between competing scientific theories. It does not and cannot establish knowledge as true. (Importantly, that does not mean a scientific theory can’t be true. It can! We just can’t know 100% for sure that it is.)

That our knowledge is made of fallible guesses has been known since antiquity. The Ancient Greek philosopher Xenophanes writes:

And even if by chance [man] were to utter
The perfect truth, he would himself not know it;
For all is but a woven web of guesses.
— Xenophanes. As quoted in Karl Popper. 2002. Conjectures and Refutations. Routledge. P. 34. Brackets mine.

That we can’t prove theories to be true has been known at the very latest since logician Alfred Tarski. He showed in the 1930s that, owing to Kurt Gödel’s incompleteness, there can be no general criterion of truth in languages rich enough. Popper writes:

… Tarski could prove that, if [some language] is sufficiently rich (for example, if it contains arithmetic), then there cannot exist a general criterion of truth. Only for extremely poor artificial languages can there exist a criterion of truth. (Here Tarski is indebted to Gödel.)
— Karl Popper. 1979. Objective Knowledge: An Evolutionary Approach. Oxford University Press. P. 46

Many scientists and science communicators ignore these crucial findings as if Popper and Tarski had never lived.

We makes guesses; some of these guesses may be correct. But if we can never know which guesses are correct, how can we proceed rationally? Although we can never justify a scientific theory as true, we can sometimes justify our preference for a theory. That’s where scientific experiment comes in.

Say you suspect there’s mold in your sink. An environmental consultant places an air-sampling pump in your sink to measure the amount of mold spores in the air. Let’s say that’s all he does. And let’s say the pump indeed finds elevated levels of mold spores. How can we know the pump is working properly; that it wouldn’t have shown elevated levels in an area we know to be mold-free? Without any additional information, we can’t address that criticism. We can’t answer that question.

That’s where a control comes in. The consultant uses the same air pump to get a second measurement, usually preemptively, this time of the air outside your house. If it tests negative for mold, as it should, then we can address the criticism. If it still shows elevated levels, we may conclude that the pump is broken and repeat the process with a different one.

Other experimental tools such as randomization and blinding serve the same purpose. For example, when testing one medicine for cancer against another, you randomly assign cancer patients to two test groups – one for each medicine. You ideally use a computer to do the random assignment for you. Scientists can be biased, even subconsciously, when doing the assignment themselves. People aren’t good at choosing random numbers anyway. For example, when asked to pick a random number between 1 and 100, too many people pick 37. So the purpose of proper randomization is to preemptively address criticisms around bias.

Likewise, when you recruit people for the experiment in the first place, you don’t already know which group they’ll be assigned to. This is known as concealment.

The same is true for blinding. In double-blind experiments, for example, neither patients nor the staff administering the drugs know which drugs are being administered to whom. All pills are made to look the same, say, and the people who do know the difference don’t interact with the patients directly. Blinding is done, again, to preemptively address the possibility of bias, not to prove a theory.

I expect scientists to disagree. In Popper’s words:

[E]xperiments only seem to be designed to prove [a] theory. In reality, experiments have to be designed to refute it, if possible. This can be bewildering for scientists to hear; they’re usually not aware of this. Consciously, the scientist wants to support his theory. But an experiment can only support a theory if the experiment could have refuted it.
— Karl Popper, https://youtu.be/li0ciaqJ0m0?t=81, 1:24 (translation mine)

In short, the role of randomization, blinding, concealment, and other experimental controls and safeguards is not to establish the truth of a scientific theory. Nothing could, because again, there’s no general criterion of truth. Although scientists may not put it this way, they use these experimental tools to correct and prevent errors, and to preemptively address criticisms that would otherwise undermine their inferences.

With those criticisms addressed, scientists can form a rational preference. But justifying a preference for one theory over another, which is doable, is entirely different from justifying that theory as true, ie ‘establishing’ or ‘proving’ causality, which is impossible.