There are some stories from childhood that seem perfectly sensible when you're seven and become rather strange when you get older. The Princess and the Pea is one of them. When you're little, the idea that a real princess can feel a tiny pea through twenty mattresses and twenty soft feather beds doesn't seem particularly odd. Of course she can. She's a princess. That's obviously what princesses do. They probably have very sensitive skin, sleep on enormous piles of bedding and are able to detect vegetables from several feet away. Nobody thinks to ask whether there's ever been any proper research into this. You simply accept it because that's how fairy tales work.
Then you grow up and start asking awkward questions.
The story begins with a prince who wants to get married. More importantly, he wants to marry a real princess. He travels all over the world looking for one, but whenever he meets someone who claims to be a princess, he can't quite establish that she really is one. Andersen doesn't tell us exactly what was wrong with any of them, which is probably just as well. Perhaps one had terrible table manners. Perhaps another didn't know which fork to use. Perhaps one admitted that she didn't particularly like castles. Whatever the reasons, the prince couldn't be sure that any of them was a genuine princess, so eventually he returned home feeling very sad.
Then, one terrible stormy night, with thunder, lightning and rain pouring down, there was a knock at the city gate. A young woman was standing outside, completely soaked. The rain had run down her hair and clothes, into her shoes and out again at the heels. She looked as though she'd been caught in a very large washing machine. She said that she was a princess, but the old queen wasn't convinced. She wanted to find out whether the young woman really was what she claimed to be.
Now, you might think the queen would ask some sensible questions. She could ask where the young woman came from, who her parents were or which kingdom she was supposed to rule. She might even send somebody to check. But this is a fairy tale, so instead the queen went into the bedroom, took the bedding off the bed and put one tiny pea underneath it. Then she put twenty mattresses on top of the pea, followed by twenty soft feather beds. Twenty mattresses and twenty feather beds is a remarkable amount of bedding. By the time you'd climbed to the top, you'd probably have needed a ladder and possibly some medication to stop you from getting vertigo.
The young woman was given the bed for the night. In the morning, the queen asked her how she'd slept, and the young woman complained that she'd slept terribly. She'd hardly closed her eyes and said that something hard had been pressing against her. She was black and blue all over. The queen was delighted because, in her mind, this proved that the young woman must be a real princess. Only a genuine princess, she decided, could be sensitive enough to feel one tiny pea through all that bedding. The prince was delighted too, so he married her, and the pea was put into a museum, where Andersen tells us that it could still be seen, unless somebody had stolen it.
The traditional moral of The Princess and the Pea is that true nobility reveals itself through sensitivity. The princess's extraordinary sensitivity to the pea is presented as proof that she really is a princess. But that isn't necessarily a very satisfying moral today.
It is, however, a surprisingly good way of starting a conversation about p values.
The thing that always strikes me about the story is that the queen isn't really testing whether there's a pea. She knows perfectly well there's a pea because she put it there herself. The pea isn't the mystery. The young woman is. What the queen is really trying to work out is whether the young woman's reaction tells her something useful about who she is. She wants to know whether there's something unusual about her. More specifically, she'd very much like to know whether she might actually be a princess.
To answer that question, the queen needs somewhere to start. In statistics we call this the null hypothesis, but all that really means is a starting assumption. In this story, we could make the null hypothesis that the young woman isn't a princess. But that isn't something the queen can test directly. She needs to think about what she would expect to see if the young woman really were an ordinary person rather than a princess.
So, for the purposes of our little experiment, we need a testable version of that starting assumption. Let's say that the null hypothesis is that the young woman's reaction is no different from the reaction we'd expect from an ordinary person.”
Perhaps she's simply a traveller caught in a storm. Perhaps she's tired, cold, fed up and uncomfortable. Perhaps she slept badly because she'd been standing outside in the rain all night. Perhaps she just found the bed uncomfortable. Whatever the reason, the queen should begin by assuming that there's nothing particularly special about her reaction.
That doesn't mean the young woman definitely isn't a princess. It means that, before seeing the evidence, the queen starts from the position that her reaction is what we'd expect from an ordinary person. Princesses are people too, after all.
There are lots of reasons why someone might complain about a bed. You might be too hot. You might be too cold. You might be worried about something. You might simply be unused to sleeping on twenty mattresses and twenty feather beds, which sounds less like a luxury and more like an extreme climbing expedition.
So the queen really needs to know how often ordinary people would react in the same way. If everybody else also woke up complaining, then there wouldn't be anything particularly remarkable about the young woman's reaction.
Although we're moving well beyond Andersen's original fairy tale, imagine that the queen had previously tested a random thousand ordinary people. Imagine they'd all slept on exactly the same bed and had all been asked the same question the next morning. Suppose that only ten of them complained as much as the young woman did. Just 1%. (And to the statisticians I know are reading this: yes, I know this isn't exactly how we'd establish a null distribution in a real study. But it will do for the purposes of this article.)
Now the result starts to look a bit unusual.
If our starting assumption is that the young woman's reaction is no different from that of an ordinary person, and only about 1% of ordinary people react this strongly, then her complaint seems rather surprising. Not impossible, but surprising enough to make us stop and think.
And that's where the p-value comes in.
People are often intimidated by p-values because they usually arrive with equations or strangely named statistical tests attached to them, but the basic idea is actually pretty straightforward. All a p-value is really trying to do is answer the question:
If the null hypothesis is true, how surprising is the result we've just seen?
More precisely, it asks how often we'd expect to see a result at least as extreme as the one we observed if the null hypothesis were true, assuming that the statistical model and its assumptions are appropriate.
In our example, about 1% of ordinary people reacted this strongly. So, as a simple illustration, a properly designed statistical test could produce a p-value of around 0.01. That would mean that if the null hypothesis were true, a reaction at least as extreme as the one we'd observed would occur about 1% of the time.
Unfortunately, this is usually the point where confusion arrives.
A p value of 0.01 doesn't mean there's a 1% chance that the null hypothesis is true. It also doesn't mean there's a 99% chance that the young woman is a princess. And it certainly doesn't prove that she's a princess. Unfortunately, the explanation favoured by a lot of people, that there's only a 1% probability that the result is due to chance, is also incorrect. What it means is that if the null hypothesis were true, a result at least this unusual would occur about 1% of the time.
That might sound like a small distinction, but it's a really important one.
Imagine the queen observes something unusual and thinks it proves that the young woman is a princess. But there could be other reasons for what she's seen. Maybe the young woman is simply very sensitive. Perhaps she had a particularly uncomfortable night's sleep. She might have been awake because she was worried about being in a strange castle. The evidence might make the "ordinary person" explanation look less convincing, but it doesn't tell us which other explanation must therefore be true.
This is why rejecting a null hypothesis doesn't automatically prove the alternative explanation. It means that the evidence is difficult to explain if the null hypothesis is true, according to the statistical test and the rule we've chosen. The queen might therefore decide that the idea that the young woman is simply an ordinary person is no longer a convincing explanation. But that still doesn't establish that she's a princess.
The queen's favourite explanation, of course, is that the young woman is a princess. But rejecting the null hypothesis doesn't magically make that explanation true. It simply tells us that the evidence doesn't fit particularly well with our starting assumption.
At this point somebody usually mentions the famous number 0.05.
You'll often hear that a result is "statistically significant" if its p value is below 0.05. For many people this number has taken on an almost mythical quality. It can sometimes sound as though the statistical gods handed it down from a mountain carved into a stone tablet. But, of course, they didn't.
Instead, it's a convention that has been widely used as a threshold for statistical significance. It means that, before looking at the result, we've decided that results this unusual would be sufficiently surprising under the null hypothesis to count as statistically significant. But 0.05 isn't a law of statistics, and it isn't appropriate for every situation.
And notice the important bit: before looking at the result.
Imagine the queen decides beforehand that she'll only be impressed if a reaction at least as strong as the young woman's would occur in fewer than five out of every hundred ordinary people. She then carries out the test and discovers that six people react that way. Her pre-decided rule hasn't been met. She can't then decide that six is close enough simply because she likes the answer better. Well, she can, but she shouldn't.
That might sound obvious, but you'd be surprised how often people start changing the rules after seeing the result. If you decide what counts as interesting after you've seen the data, you can make an ordinary result look much more impressive than it really is.
Things get even messier when lots of questions are being asked.
So suppose the queen doesn't just ask whether the young woman felt the pea. Maybe she asks whether she slept badly, whether she woke up in the night, whether her back hurt, whether she felt tired in the morning, whether she complained about the bedding, and twenty other questions that occur to her afterwards. Even if nothing unusual is really happening, one of those questions might produce an interesting result simply because of chance. The more questions she asks, the more opportunities she gives chance to produce something that looks unusual. This is known as multiple testing.
Here's why that's important. Imagine the queen carries out 20 separate statistical tests, and for each one she decides that a p-value below 0.05 will count as statistically significant. If all 20 tests are genuinely testing situations where the null hypothesis is true, and the tests are independent, there is about a 64% chance of getting at least one statistically significant result just by chance. That's quite a big difference from the 5% we started with.
The 5% threshold applies to each individual test. Once we start doing lots of tests, we're giving chance lots of opportunities to produce something that looks unusual. The more tests we perform, the greater the chance that at least one of them will cross our chosen threshold even though nothing unusual is really happening.
The result itself may be perfectly genuine. The problem is that if we tested 20 things and found one interesting result, we need to know that we tested 20 things. Finding one statistically significant result sounds impressive if you tested one thing. It sounds rather less impressive if you tested twenty different things and only mentioned the one that produced the interesting answer.
That's why we need to know how many questions were asked, how many things were measured, and how many different groups or patterns were examined. Otherwise, we can easily mistake an unusual result for an important discovery when it may simply be the unusual result that appeared because we gave chance so many opportunities to produce one.
There's another problem too, and this one often gets forgotten. Even if the result is statistically significant, does it actually matter?
Imagine that the queen discovered that princesses were, on average, 0.01% more sensitive to peas than ordinary people, and because she tested ten million people, the difference produced a wonderfully small p-value. It might be statistically significant. That doesn't necessarily mean it matters very much. A difference can be statistically convincing while being so tiny that it makes practically no difference to anyone.
Would anybody care?
Probably not. You certainly wouldn't be able to detect this difference by adding another mattress.
That's because statistical significance and practical importance aren't the same thing. Something can be statistically convincing whilst still being so tiny that it makes no meaningful difference in the real world. Statistics can tell us that the observed difference would be unusual under the particular model and assumptions we're using. It can't tell us whether that difference is useful, important, or worth spending money on.
The more I think about Andersen's story, the more I suspect the queen would've been a difficult colleague in a research or analytical department. She started with a sensible question about whether the young woman might be a princess. She then observed some unusual evidence. But instead of considering all the possible explanations, she immediately chose the explanation she liked best. Evidence and explanations aren't the same thing, though. There can be many explanations for the same evidence.
That's really what the p-value is trying to remind us. The p-value doesn't tell us whether the young woman is a princess. It tells us how well the evidence fits with the starting assumption. If the evidence doesn't fit very well, we may reject the null hypothesis. But that doesn't prove the alternative explanation, and it certainly doesn't tell us why the result occurred.
What we do with the evidence in front of us still requires thought. We still need to consider other explanations. We still need to decide whether the result matters. We should also ask whether we've looked at the data fairly and whether we decided what we were testing before we saw the answer. Statistics can help us ask those questions, but it can't answer all of them for us.
And maybe that's why I still like this fairy tale. Hidden underneath twenty mattresses, twenty feather beds and one rather unfortunate green vegetable is a lesson that statisticians are still trying to teach today. We begin with a starting position, or null hypothesis. We collect evidence. We ask whether the evidence fits with our hypothesis. Then we think carefully about what it might, and might not, mean.
So the next time somebody mentions a p-value, don't think about equations, Greek letters, strangely named statistical tests, or complex software. Think about a stormy night, a suspicious queen, an uncomfortable and ridiculously high bed, and a young woman who may not be as ordinary as she appears.
Think about The Princess and the P-Value.