Can a Single Tiny Difference Ever Prove a Cause?
The One-Difference Detective

Picture a science fair. You have two bean plants in identical pots, same soil, same light. You add a little fertilizer to one pot, but only plain water to the other. After a week, the fertilized plant is taller. Did the fertilizer cause the extra growth? You probably think yes, because the only thing you changed between the two pots was the fertilizer.
Philosophers have thought deeply about this kind of reasoning. In the mid-1800s, John Stuart Mill (1806–1873) captured it in a rule he called the method of difference. If a factor is present in one situation and absent in another, and everything else is kept the same, that factor must be a cause, an effect, or part of the cause of whatever changes.
Biologists use this method constantly. Imagine a scientist wants to know if a new compound kills bacteria. She divides a bacterial broth into many samples. To half the samples she adds the compound dissolved in a buffer; to the other half she adds only the buffer. If the treated samples stay clear and the control samples turn cloudy, she can infer that the compound stops bacterial growth. The experimental trick is making sure the two groups differ only in the presence of the suspected antibiotic. Any other hidden difference is a confounder — an unnoticed variable that could really be the cause. Confounders are why a careful experimenter stirs the broth well, uses the same batch of glassware, and handles every sample the same way. Mill’s method works only if we can trust that no confounders slipped in, which in practice is a very strong assumption.
When One Experiment Decides Everything — Or Does It?

Sometimes a single experiment seems to settle a huge scientific debate. In the 1960s, biochemists were arguing fiercely about how cells charge their energy “batteries.” The process, called oxidative phosphorylation, turns food into the energy molecule ATP. Many researchers thought a mysterious chemical intermediate carried the energy from respiration to ATP production — but they could never find it. Peter Mitchell (1920–1992) proposed a radically different idea: respiration pumps protons across a membrane, creating a gradient, and that gradient directly powers ATP production. No chemical middleman needed.
In 1974, Ephraim Racker and Walter Stoeckenius performed an experiment that many textbooks call “crucial.” They built synthetic membrane vesicles containing only two purified proteins: a light-driven proton pump from a bacterium, and the ATP-making enzyme from mitochondria. When they shined light on the vesicles, ATP was produced. No respiration enzymes, thus no possible chemical intermediate, could have been present. The proton gradient alone was enough. Biochemists mostly accepted Mitchell’s chemiosmotic mechanism, and he later won a Nobel Prize.
But the French philosopher Pierre Duhem (1861–1916) would have raised an eyebrow. He argued that a crucial experiment can never logically prove one hypothesis to be true, because we might not have even thought of the true explanation yet. Even if an experiment kills off a rival idea, some unknown alternative might still lurk in the shadows. The Racker-Stoeckenius experiment eliminated the chemical intermediate, but it didn’t prove that Mitchell’s scheme was the only possible one.
A similar caution appears in another famous case. In 1958, Matthew Meselson and Frank Stahl used heavy nitrogen to show how DNA copies itself. After one round of replication, the DNA settled at an intermediate density — exactly what the correct “semi-conservative” model predicted. But the scientists themselves admitted they could not completely rule out a different, “conservative” model if the DNA fragments had stuck together end-to-end. They didn’t know all the molecular details of their own setup. So even a beautiful experiment leaves tiny logical gaps. What made the Racker-Stoeckenius experiment so powerful was not airtight logic, but the extraordinary control it gave biochemists over a clean, artificial system — confounders were dramatically reduced.
Why a Fruit Fly Can Tell You About Your Own Body

It seems strange: to understand human diseases, biologists spend years breeding tiny flies. Model organisms like the fruit fly Drosophila melanogaster, the roundworm Caenorhabditis elegans, and baker’s yeast have starred in countless discoveries. Thomas Hunt Morgan’s lab in the early 1900s used fruit flies to map genes onto chromosomes. The flies were easy to raise, produced many generations quickly, and threw up a steady stream of mutations to study. Morgan’s fly room became what one historian called a “breeder-reactor” — a living machine that churned out genetic knowledge.
But how do scientists generalize from a fly to a person? A deep puzzle called the extrapolator’s circle says: to infer that a mechanism works the same in two organisms, you need to know they are relevantly similar — but to know they are similar, you often already need to know the mechanism is there. Sometimes this circle is broken by powerful evidence from evolution. The genetic code — the rule that translates RNA triplets into amino acids — is almost identical in all life. It’s an arbitrary, “frozen accident” that arose once and stayed nearly unchanged because altering it would be lethal. So the odds that fly and human codes would match by chance are vanishingly small; common ancestry explains it, and that gives a rational basis to extrapolate. In other cases, biologists compare specific genes — if a DNA sequence is conserved across many species, it’s a strong hint that it plays a similar role. None of this makes extrapolation automatic, but it turns a blind leap into a careful, evidence-based argument.
The Blob That Wasn’t There

For over twenty years, electron microscopists thought they had discovered a new structure inside bacteria. They called it the mesosome — a membrane-bound blob that appeared in many different bacteria under different conditions. Textbooks included neat diagrams of it. But in the 1980s, the mesosome disappeared.
What happened? Biologists gradually realized that mesosomes were not real cell parts. They were artifacts — distortions created when the bacteria were chemically fixed for the microscope. The preservatives used caused membranes to crumple into the shapes that looked so convincing. Once scientists performed experiments that traced the blobs directly to the fixation process, they abandoned the mesosome.
This story highlights the problem of data reliability. One strategy philosophers have studied is robustness: a result is more trustworthy if it shows up using completely different methods, like light microscopy versus electron microscopy, which rely on different physical processes and theoretical assumptions. If two independent methods agree, it’s less likely the result is a fluke of one technique. But in the mesosome case, the evidence was discordant — some labs saw mesosomes, others didn’t. Robustness alone couldn’t decide. Instead, the key was understanding the causal process that generated the misleading images. By showing that the fixatives caused the membrane blobs, researchers turned the experimental tools back on themselves. Judging data reliability, at its best, is just another kind of causal detective work — the same mode of thinking Mill pointed to, now directed at your own instruments.
Your Own Detective Toolkit

Every time you change one thing to see what happens, you’re running a miniature version of the method of difference. Did your computer crash because you installed that new app, or was the battery already dying? Did your shampoo really make your hair shinier, or did you just rinse longer than usual? You instinctively look for confounders and check if the effect shows up more than once.
Philosophers of science have shown that real experimental knowledge is never a simple “smoking gun.” It’s built on careful design, open acknowledgment of assumptions, and a habit of questioning your own tools. When you hear a science headline — a new medicine, a scary chemical — you can ask: Was there a proper control group? Could another difference explain the result? Has anyone reproduced it with a different method? These questions, which grew out of centuries of thinking about experiments, help you navigate a world full of causal claims. The puzzles Mill and Duhem raised aren’t just for labs; they’re for anyone who wants to know what’s really going on.
Think about it
- Suppose you want to test whether a new plant fertilizer works better than plain water. You water one plant with the fertilizer and another with water, but you put the fertilized plant in a sunnier spot. What might be a confounder here, and how could you fix the setup?
- Duhem said we can never be absolutely certain an experiment has proved a theory true, because an unknown idea might explain things even better. Does that mean we should never be confident in any scientific finding? Why or why not?
- Medicines are often tested on mice before humans. What are two reasons a drug might work in mice but not in people, and what kind of extra evidence could make you more confident it will work in humans?





