There was a brilliant Two Ronnies sketch based on Mastermind in which Ronnie Corbett is playing the contestant and Ronnie Barker is the host. Corbett's specialist subject is answering questions before they're asked. Barker asks him a question and Corbett gives him the answer to the question Barker is about to ask next. Barker then asks that question and Corbett gives him the answer to the one after that. And so it goes on, with Corbett apparently demonstrating an extraordinary level of knowledge while never actually answering the question in front of him.
It’s a simple joke, but it works because the answers given by Ronnie Corbett are perfectly sensible, and in most cases quite amusing as well. The problem is that, although they sound plausible, they’re answers to the wrong questions.
For some reason I was reminded of that sketch when I was thinking about data, statistics, dashboards and artificial intelligence. Actually, the problem I’ve encountered in real life is slightly different, and arguably worse. In the Two Ronnies sketch, at least there is a question. The contestant simply answers a different one. In the real world, however, its more usual that we produce an answer before we’ve even worked out what the question is.
That’s become a surprisingly common feature of working with data. Someone wants a particular number and only afterwards do we start discussing what the number actually means. Reports can continue for years after everyone has forgotten why they were originally introduced. Dashboards can end up containing every measure that can be extracted from a database, on the assumption that if enough information is put in one place somebody will eventually find it useful. Data linkage can become an objective in itself, simply because we now have the ability to bring datasets together.
Dashboards can end up containing every measure that can be extracted from a database, on the assumption that if enough information is put in one place somebody will eventually find it useful. Data linkage can become an objective in itself, simply because we now have the ability to bring datasets together.
One of the things I’ve learnt over the years is that the difficult part of analysis isn’t the calculation. Computers are very good at calculations, although some forms of AI still manage to get surprisingly confused by them. Give me a well defined question, appropriate well defined data and a suitable method, and I'll find somebody or something capable of doing the arithmetic. The much more difficult part is deciding what we’re actually trying to find out.
It sounds obvious that we should be clear on what we want to answer before we start, but it becomes surprisingly complicated once you start dealing with real organisations and real data. Words that appear to have perfectly clear meanings often don’t, particularly when different people assume they’re talking about the same thing when they’re actually talking about something slightly different.
Take something as simple as the number of people attending an event. You might have a registration system that tells you 1,200 people attended. We’ve got a number, so it can feel as though we’ve made progress. Except what does “attended” actually mean? Does it mean they registered, or that they actually turned up? Does somebody who stayed for ten minutes count? What about speakers and exhibitors? What about somebody who attended online? What if one person registered but didn’t turn up, while somebody else arrived without registering? You can spend an afternoon arguing about whether the answer is 1,200 or 1,300 while completely missing the fact that nobody has agreed what they’re actually counting. The calculation might be flawless, but the question can still be rubbish.
You can spend an afternoon arguing about whether the answer is 1,200 or 1,300 while completely missing the fact that nobody has agreed what they’re actually counting. The calculation might be flawless, but the question can still be rubbish.
I had a particularly good example of this in the mid-2000s when I was working nationally in the NHS. Somebody fairly senior within one of the IT teams had developed a dashboard and, as far as I was concerned, it had appeared from nowhere. I hadn’t been involved in its development, consulted about what information I needed or even told it was being built. Then a meeting was arranged with me at which I was presented with the finished product and asked what I would use it for.
I answered, “Nothing.”
That wasn’t, apparently, the response they were hoping for. Don't get me wrong, I wasn’t saying that the dashboard was badly constructed or that the people who’d built it hadn’t done a good job. For all I knew, they’d done exactly what they’d been asked to do. The problem was that it hadn’t been designed to answer a question that mattered to me, so I had no use for it. Somehow, this seemed to make me difficult, which I found rather odd. If somebody asks me what I’m going to use something for and I say “nothing”, I’m not sure how inventing a use would improve the situation. Perhaps I could’ve said that it looked very interesting and promised to study it carefully, and then quietly forgotten about it. That would probably have made everybody feel better. The dashboard would still have been no more useful, but at least in their eyes we’d have had a successful meeting.
To me, the sensible sequence seemed fairly obvious. Ask what questions needed answering, establish what information would help and then decide how best to present it. If a dashboard was the right tool, build one. If a report was better, produce a report. If the answer could be provided by a single number in an email, perhaps just send the number and save everyone the trouble. Starting with the dashboard and then asking what I might use it for was rather like buying a screwdriver and then wandering around the house looking for something that needed a screw. The screwdriver isn’t the problem. It’s probably a perfectly good screwdriver. The problem is that nobody particularly needs one at that moment.
Starting with the dashboard and then asking what I might use it for was rather like buying a screwdriver and then wandering around the house looking for something that needed a screw.
I encountered a much bigger version of the same issue in relation to NHS data. This time the answer wasn’t a dashboard. It was the data itself, or more precisely the desire to bring lots of different datasets together. There was understandable enthusiasm for linking information because the potential benefits were obvious. Combining information from different parts of the health system might help us understand people’s journeys, identify patterns that weren’t visible in individual datasets, improve services or answer questions that couldn’t be addressed using one source alone. All of those are perfectly reasonable reasons for linking data. What concerned me was when the ability to link datasets became a reason for doing it in the first place. I kept asking, “What are we trying to find out?”
As I remember it, I was the sole voice on the Board repeatedly asking that question in relation to care.data. I wasn’t arguing that data shouldn’t be used or that datasets should never be linked. I was saying that we shouldn’t link people’s information simply because we could. We should begin with a question that needed answering, establish what information was required and then consider whether linking the relevant data was necessary and justified. That might sound like a fairly modest proposition, particularly when dealing with people’s health information. Unfortunately, it wasn’t received that way. Once again, I was seen as difficult.
Perhaps there’s something inherently awkward about asking what something is for when everybody else is excited about what it might be capable of doing. If somebody has spent months developing the technology to link datasets, asking what question we’re trying to answer can sound as though you’re applying the brakes. The excitement is about possibility. The question is about purpose. Unfortunately, possibility is much more fun than purpose.
The excitement is about possibility. The question is about purpose. Unfortunately, possibility is much more fun than purpose.
Of course, there’s also a tendency for us to assume that more data must be better, but that's not necessarily the case. Any statistician will tell you that more data can mean more noise, more opportunities to find meaningless relationships due to multiple comparisons and more scope for people to misunderstand what they’re looking at. And of course, where personal information is involved, particularly health information, there are also questions about confidentiality, security, access, retention and legitimate use. And those aren’t just bureaucratic annoyances that get in the way of what's interesting. They’re an important, and necessary, part of the work.
This way of thinking about questions isn’t limited to data. I’ve written about versions of it before, including The Dress That Nobody Needed and The Kitchen Drawer.
In The Dress That Nobody Needed, a dress was bought by my wife for £120 rather than £300, which sounds like a saving of £180. But that's only the case if buying the £300 dress was ever a realistic option. The more useful, an usually unasked question is whether the dress was needed at all. Once you ask that, the calculation looks rather different.
The Kitchen Drawer presented a similar problem. Faced with a drawer overflowing with cables, chargers and mysterious electronic relics, the obvious question is how to find the right cable more quickly. But perhaps the better question is why we have accumulated so many cables in the first place. We can become extremely efficient at solving the wrong problem.
I think one of the less discussed effects of AI is that it makes it extraordinarily cheap to answer questions. We can produce a report in ten minutes that might once have taken three hours. We can create a presentation, analyse a dataset, generate images, summarise documents or produce recommendations almost immediately. That’s impressive, and I’m certainly not going to argue that automating repetitive work is a bad thing. I’ve spent years helping people automate data processes. If a computer can do the boring part while people concentrate on something more worthwhile, that’s usually a good outcome. But there’s a difference between making something easier to produce and making it worth producing.
there’s a difference between making something easier to produce and making it worth producing.
The three hours involved in creating a report used to provide a natural opportunity to ask whether the report was needed. If it took most of a day, you might at least have a conversation about who was going to read it and what they were going to do with it. When the same report can be produced in ten minutes, that pause can disappear. Why spend time debating whether something is worth doing when you can simply do it? That’s where I think we need to be careful.
AI is very good at producing plausible answers. That creates a different problem from the familiar issue of hallucination. An obviously ridiculous answer is relatively easy to spot. If an AI system tells me that Huddersfield is in Devon, I don’t need a PhD in artificial intelligence to question it. A polished answer is much harder to challenge. It might contain sensible language, convincing statistics and a beautifully structured argument. It might even be factually correct. It still might not answer the question we needed answering.
That’s the real parallel with the Two Ronnies sketch. The danger isn’t necessarily that the answer is wrong. It’s that the answer is perfectly reasonable but belongs to a different question. And because AI makes production so cheap, we can now do this at a scale that wasn’t previously practical. We can generate reports nobody needs, analyse data nobody intended to use, produce presentations nobody will read and automate processes that perhaps shouldn’t exist in the first place.
The technology isn’t making us ask the wrong questions. We’ve always been capable of that. It’s simply removing much of the effort and time that used to force us to think about whether the question was worth asking. That’s something I find particularly interesting because I’ve always thought that knowing how to use a tool isn’t the same as knowing when to use it. It’s reflected in the training courses I deliver. You can teach somebody how to create a graph without teaching them which graph is appropriate. You can explain how to calculate a confidence interval without teaching them how to interpret it. You can show somebody how to build a dashboard without asking what decision it’s supposed to support. You can demonstrate how to use AI without teaching them when to put the keyboard down and think. And at the moment that's probably the most important thing that we can do.
The technology isn’t making us ask the wrong questions. We’ve always been capable of that. It’s simply removing much of the effort and time that used to force us to think about whether the question was worth asking.
I’m quite fond of technology. I’ve spent enough time working with Excel, Power Query, Power Pivot, Power BI, SPSS and other statistical software to have no interest in returning to the days when everything had to be done by hand. I remember in my early career with stats tables and basic spreadsheets trying to work things out. It took hours, sometimes days. So technology can save enormous amounts of time when it’s applied to the right problem. The problem is that technology doesn’t get to decide what the right problem is; that’s still our job.
We tend to measure success by activity. Something was built, launched or delivered. A target was achieved. A report was produced. A system went live. Yet activity isn’t the same as progress. You can run faster in the wrong direction. You can automate a bad process. You can make an unnecessary report incredibly efficient. You can produce more data without learning anything. You can build a dashboard that nobody uses. You can link datasets without answering anything that needed to be answered. And now you can ask AI to help you do all of it before you’ve finished your coffee.
I don’t think the answer is to use less technology. Instead it’s to become more comfortable with asking awkward questions before everybody gets too far down the road.
I wrote about this in Ask a Silly Question many years ago, but the questions haven’t changed. What are we trying to find out? What decision will be affected? What information is genuinely required? Does the proposed solution support that decision? Does this report need to exist? Is linking the data necessary and proportionate? Was that three hour task worth doing before we celebrated reducing it to ten minutes? And perhaps the hardest question of all is whether we’re prepared to say, “I don’t know,” when we genuinely don’t know.
That isn’t always easy because organisations like certainty, or at least the appearance of it. People who seem decisive are usually rewarded more readily than people who point out that the evidence is insufficient. The person who says, “Here’s the answer,” sounds useful. The person who says, “I’m not sure we’ve asked the right question yet,” can sound as though they’re holding things up. I know this first hand. I’ve been that difficult person.
The person who says, “Here’s the answer,” sounds useful. The person who says, “I’m not sure we’ve asked the right question yet,” can sound as though they’re holding things up.
I suppose, in hindsight, that I could’ve been more accommodating. I could’ve found a use for the dashboard, said that the linked data would probably reveal something interesting or nodded along when everybody was excited about what the technology might make possible. But somebody has to ask what the thing is for. Otherwise, we’re just building things and hoping that purpose will arrive later.
That’s the connection between the dashboard, care.data, data linkage, automation and AI. They’re not really different problems. They’re all examples of an answer arriving before the question. The dashboard nobody requested, the datasets linked without a defined purpose, the report nobody reads and the automated process that saves three days of unnecessary work all begin in the same place.
AI simply gives us the ability to produce many more of these answers before we’ve worked out what we need to ask. That could be useful, but it could also be dangerous because the more capable our tools become, the less effort is involved in producing something. The less effort involved, the easier it becomes to avoid asking whether it was necessary.
the more capable our tools become, the less effort is involved in producing something. The less effort involved, the easier it becomes to avoid asking whether it was necessary.
And that’s why the apparently difficult person in the room has a useful role to play. I don’t mean the person who says no to everything, believes technology is evil or refuses to accept change because everything was apparently better when we used paper. I mean the person who continues to ask why. Why do we need this? Who needs to know? What decision will change? What data do we require? Why does it need to be linked? What happens if we don’t do it? And, perhaps most awkwardly of all, what if the answer is that we don’t need it?
I no longer work nationally and have instead made the jump to working for myself. I’m sure people I worked with breathed a sigh of relief as the difficult questions disappeared. It’s always more comfortable to surround yourself with people who agree with you. But even though I no longer work in those circles, I’ll still actively continue to be difficult. If somebody shows me a dashboard, I’ll want to know what question it answers. If someone wants to link datasets, I’ll ask why. If I’m told that AI can produce something in ten minutes that used to take three hours, I’ll still ask whether we needed it before the three hours became ten minutes.
That doesn’t mean I’m negative, or a cynic, or against dashboards, data linkage, automation or AI. It means that I’m still attached to the old-fashioned idea that we should know what we’re trying to do before we start doing it. After all, the most expensive answer in the world is still an answer to a question nobody needed to ask, and the most efficient way of producing it is still a waste of time.
the most expensive answer in the world is still an answer to a question nobody needed to ask, and the most efficient way of producing it is still a waste of time.
So, after all these years, I’ll still find myself returning to the same thing.
What was the question again?