It Doesn't Add Up
Looking back over several of the articles I've written during the past few years, I'm beginning to realise they have all been circling around the same issue from different directions.
In one of my earliest articles Ask a Silly Question, I wrote about the danger of accepting statistics at face value. The article used unemployment data as an example, showing how a measure that appears straightforward becomes much more complicated once you begin asking questions about definitions, assumptions and what is actually being counted. The central message wasn't about unemployment but that numbers rarely speak for themselves, and understanding often begins with challenging what appears obvious.
In Human vs AI, I explored a different problem. I described an experiment in which a Large Language Model struggled with a tightly constrained task of finding anagrams despite appearing highly intelligent in almost every other respect. What interested me wasn't that the AI got things wrong. Humans get things wrong all the time. What fascinated me was that it could be wrong while sounding entirely convincing. It could produce answers far more easily than it could verify them.
Then, in My Way or the AI Way, I looked at what happens when human judgement collides with algorithmic recommendations. Increasingly, professionals are finding themselves in situations where a machine suggests one course of action while experience and expertise suggest another. The question is no longer whether AI can produce answers. The question is what happens when we start trusting those answers.
The question is no longer whether AI can produce answers. The question is what happens when we start trusting those answers.
And most recently, in The Responsibility to Think, I reflected on our own obligations when using these systems. Not whether AI is intelligent, but whether we continue to exercise our own judgement and think critically when interacting with it.
When I wrote each one of these I regarded them as separate topics, especially as some were written before the advent of AI. But increasingly, I can see that they are all aspects of the same underlying challenge.
We’re becoming extraordinarily good at generating answers. Unfortunately, we’re not necessarily becoming better at understanding them. And we’re still not very good at asking the right questions
One of the common misconceptions surrounding artificial intelligence is that because it can discuss complex subjects, it must also be good at analysis. But the two are not the same thing at all. Analysis isn’t simply the production of information. It involves determining whether information means what we think it means, whether it is reliable, whether it answers the question we intended to ask, and whether alternative explanations might exist. That requires a very different set of skills.
This is where Large Language Models become particularly interesting. They speak the language of analysis fluently. They can discuss regression models, confidence intervals, epidemiology, economics and research methodology. They can explain sophisticated statistical concepts in clear language and often do so extremely well. The result is that they appear analytical.
Yet underneath that fluency sits an important limitation. A Large Language Model was not originally designed to calculate. And it wasn’t designed to reason through strict logical constraints. An LLMs primary purpose is to predict the most likely continuation of language. That distinction sounds technical, but it has enormous implications. A calculator performs arithmetic. A spreadsheet performs calculations. But Large Language Models do something fundamentally different in that they just generate plausible sequences of words.
underneath that fluency sits an important limitation. A Large Language Model was not originally designed to calculate. And it wasn’t designed to reason through strict logical constraints. An LLMs primary purpose is to predict the most likely continuation of language. That distinction sounds technical, but it has enormous implications.
Most of the time this is exactly what we want. It allows them to explain, summarise, draft and communicate at a speed that would have seemed impossible only a few years ago. The problem, however, is that generating plausible language isn’t the same thing as validating a calculation, following a set of constraints or checking whether a conclusion logically follows from its premises.
This is precisely what I encountered in Human vs AI. The exercise involved a task with clearly defined rules. But the model repeatedly drifted away from those rules, not because it lacked access to the information, but because strict rule enforcement isn't what it was designed to do. It was optimising for plausible output rather than rigorous verification.
This explains why an LLM can sometimes do things that appear contradictory. It might explain a complex statistical method while making mistake that a pocket calculator would never make. It can summarise a lengthy technical report accurately and then struggle with a constrained logic puzzle. It will write a persuasive explanation of analytical rigour while simultaneously miscounting, overlooking a rule or citing an incorrect figure.
At first glance it doesn’t add up. You see, we tend to assume that language, mathematics and reasoning belong together because they belong together in our own minds. But for Large Language Models they are largely separate capabilities. Excellence in one area doesn’t automatically imply competence in another.
we tend to assume that language, mathematics and reasoning belong together because they belong together in our own minds. But for Large Language Models they are largely separate capabilities. Excellence in one area doesn’t automatically imply competence in another.
If I needed to add up a budget, reconcile financial records, validate a dataset or perform a complex calculation, a Large Language Model would not be my first choice. I’d use software specifically designed for those purposes.
Likewise, if I needed strict logical validation, mathematical consistency checks or deterministic rule enforcement, I’d choose tools designed explicitly to carry out those tasks. As the old adage goes "You should use the best tools for the job"
The problem, of course, is that the output from an LLM often looks more intelligent than the output from a spreadsheet. The spreadsheet simply provides an answer whereas the LLM explains itself. But eloquence and accuracy are not the same thing.
Of course, many analytical failures have very little to do with arithmetic in the first place. Throughout my career I've seen analyses that were mathematically perfect and conceptually wrong. Every number was correct. Every graph was accurate. Every calculation had been performed appropriately. The problem was that the wrong thing was being measured, compared or interpreted. The answer was right but the question was wrong.
As the statistician John Tukey once observed, "Far better an approximate answer to the right question, which is often vague, than an exact answer to the wrong question, which can always be made precise." That distinction sits at the heart of analytical thinking. Precision has value, but only if we're being precise about the right thing.
This was really the lesson behind Ask a Silly Question. A measure can be technically correct while leading us towards an incorrect conclusion. A reduction in unemployment may not mean what we initially think it means. An increase in disease prevalence may not mean what we assume it means. The value comes from asking questions about context, assumptions and interpretation rather than merely accepting the headline figure.
A measure can be technically correct while leading us towards an incorrect conclusion ..... The value comes from asking questions about context, assumptions and interpretation rather than merely accepting the headline figure.
This is where experienced analysts often behave very differently from AI systems. As analysts, we're trained to think about numerators and denominators, inclusion criteria, case definitions, confounding variables, missing data, bias and uncertainty. We spend an extraordinary amount of time not calculating answers but questioning whether a calculation answers the question we were actually trying to ask. Does the question itself makes sense? Are we measuring the right thing? Are we using the correct definition? Is important information missing? Have we confused correlation with causation? What assumptions are hidden within the request?
A Large Language Model will often begin constructing an answer immediately. That's not a flaw but what it was designed to do. The difficulty is that analytical practice often requires the opposite approach. Sometimes the most valuable contribution isn’t answering the question but challenging it. As the statistician John Nelder said, "The object of a statistical test is not to answer the question, but to raise it."
This is also why the discussion in My Way or the AI Way feels increasingly relevant. As AI systems become more articulate and more persuasive, there is a temptation to trust them simply because they sound authoritative. The report is well written. The recommendation appears balanced. The explanation feels logical. Yet none of those characteristics guarantee that the underlying reasoning is sound. Good writing has never been the same thing as good analysis. Artificial intelligence simply makes that distinction easier to overlook.
Strangely, the growth of AI may actually increase the value of some of the oldest analytical skills. Understanding measures. Challenging assumptions. Recognising bias. Evaluating evidence. Asking awkward questions. The very behaviours discussed in Ask a Silly Question long before most people had heard of ChatGPT or Claude …. (insert your favourite model here).
the growth of AI may actually increase the value of some of the oldest analytical skills. Understanding measures. Challenging assumptions. Recognising bias. Evaluating evidence. Asking awkward questions
Technology changes rapidly. Human reasoning evolves much more slowly. The tools available today would have looked like science fiction when I began my career. We can now process more data, generate more reports and obtain more answers than ever before. Large Language Models can explain complex concepts, summarise vast quantities of information and help us explore ideas at remarkable speed. Used appropriately, they're genuinely transformative. But they're not calculators or statistical packages. They're also not particularly well suited to constrained logic or rigorous rule-checking. Most importantly, as I state repeatedly, they're not a substitute for understanding.
W. Edwards Deming famously said, "Without data you're just another person with an opinion." The point was never that data alone creates knowledge. It was that good decisions need evidence rather than intuition. But evidence still requires interpretation. A dataset does not explain why something happened, whether it has been measured correctly, or whether the conclusions drawn from it are justified.
This is why I keep returning to the arguments I made in Ask a Silly Question, Human vs AI, My Way or the AI Way and The Responsibility to Think. Beneath the different examples lies the same underlying lesson. Good analysis has never been about producing answers as quickly as possible. It has always been about understanding what those answers mean, how they were produced, what assumptions underpin them and whether they genuinely address the problem we are trying to solve.
An AI can generate a thousand answers in seconds. It can explain them, justify them and present them in flawless prose. The arithmetic may be correct. The methodology may be correct. The logic may even appear correct. Yet if the assumptions are wrong, the question is misguided, the data incomplete or the interpretation flawed, the conclusion can still be entirely wrong.
if the assumptions are wrong, the question is misguided, the data incomplete or the interpretation flawed, the conclusion can still be entirely wrong.
Abraham Maslow once observed, "I suppose it is tempting, if the only tool you have is a hammer, to treat everything as if it were a nail." The danger with any powerful tool is that we start adapting problems to fit the tool rather than choosing the tool that best fits the problem. With Large Language Models, there is a growing temptation to believe that every analytical challenge simply requires a better prompt. Sometimes that will help but other times it won't. A hammer can be an excellent tool, but it remains a poor choice when the job requires a screwdriver.
Large Language Models are exceptional tools for communication, exploration, synthesis and creativity. They are often useful assistants for analysis. But they are not universal analytical tools. They are not the best tools for arithmetic. They are not the best tools for constrained logical reasoning. They are not the best tools for validation and verification. When precision matters, choosing the right tool matters too.
The future almost certainly belongs to people who combine the strengths of different tools appropriately. Use calculators for calculation. Use statistical software for statistics. Use databases for retrieval. Use LLMs for exploration, explanation and communication. And most importantly, continue asking the awkward questions. Because sometimes the arithmetic is correct. The graphs are correct. The methodology is correct. The calculations are correct.
Yet something still feels wrong.
Something doesn't quite fit.
And when a beautifully written answer, based on the wrong assumptions, produced by the wrong tool, leads confidently to the wrong conclusion, there is really only one thing to say.
It doesn't add up.