Guardrails Before Gadgets

Artificial intelligence and the NHS seem destined to become increasingly intertwined. Whether we are talking about appointment booking, clinical documentation, diagnostics or patient communication, the direction of travel has been clear for some time. The latest announcement reported on by the BBC, simply moves that journey a little further forward.

Under the plans, patients using the NHS App will answer a series of questions and be directed towards what appears to be the most suitable service for their needs. That might be a GP appointment, a pharmacy, a community service, urgent care, A&E, or self-care advice. At the same time, NHS trusts are expanding the use of software that can listen to consultations and generate draft clinical notes. The intention is straightforward enough: reduce administrative burden, improve access and make better use of stretched resources.

And its difficult not to see the attraction. Most people have experienced the familiar problem of trying to contact their GP surgery at eight o'clock in the morning. We know that practices are under immense pressure, clinicians are severely overloaded, and patients often struggle to find the right route into services. If technology can help direct people more efficiently, reduce queues and free up clinical time, that has obvious appeal.

If technology can help direct people more efficiently, reduce queues and free up clinical time, that has obvious appeal.

The documentation side may be even more attractive. Healthcare professionals spend a remarkable amount of time writing notes, updating records and completing administrative tasks. If software can handle some of that burden, clinicians may be able to spend more time focused on the patient sitting in front of them.

Up to now, everything I've discussed in this article reflects either the proposals outlined in the BBC report or the potential benefits identified by NHS England and those commenting on the rollout. What follows now goes beyond the article itself and considers some of the issues that may become increasingly important as these developments progress and become embedded within routine care.

My immediate reaction to the announcement was to think back to some of the themes I've explored in previous articles, particularly When Machines Speak Medicine. In that particular article, I discussed evidence showing that even some of today's most capable language models can generate incorrect medical information with surprising confidence. The issue isn't that the systems occasionally break. Instead they are designed to generate responses that appear plausible. And most of the time those responses are accurate. But sometimes they are not. That distinction matters because healthcare isn't an environment where "mostly right" is usually good enough.

During the From Code to Consequence panel discussion on AI, ethics and governance, I made a point that I think is relevant here. Most people focus on how often a system gets things right. Healthcare tends to focus on the occasions when it gets things wrong. As I said during the discussion, "You only have to be wrong once. It's not that it's right 9 million times. It's the fact that it's wrong once."

Most people focus on how often a system gets things right. Healthcare tends to focus on the occasions when it gets things wrong. As I said during the discussion, "You only have to be wrong once. It's not that it's right 9 million times. It's the fact that it's wrong once."

That's not an argument against using AI, of course. But it is an argument for understanding the environment in which it is being deployed.

A recommendation that directs somebody with a minor ailment towards a pharmacist instead of a GP may be entirely reasonable. But a recommendation that misses the early signs of a serious condition is a different matter altogether. The challenge here is that healthcare is rarely neat and predictable. Many patients don't always describe symptoms clearly. Important details are sometimes omitted. Symptoms overlap. Conditions present atypically. Experienced clinicians spend years learning how to navigate that uncertainty.

That's why I remain cautious whenever discussions veer towards the idea that technology can replace professional judgement. In When Machines Speak Medicine I argued that AI should always be viewed as a support tool rather than a decision-maker.

The BBC article does not discuss the regulatory status of the proposed triage system. However, as AI-supported triage develops, one question that may become increasingly important is whether some components fall within existing medical device regulations. This isn't a criticism of the current proposal, nor does it imply that regulatory requirements have not been considered. However it does highlight an area that deserves attention as further details emerge.

Many people will assume that an app is simply an app. In reality though, the answer depends on what the software has been designed to do. Under current UK guidance, software that influences diagnosis, treatment decisions, monitoring or patient management can fall within the definition of software as a Medical Device. Patient-facing symptom checkers and digital tools that provide health advice are specifically identified as examples that may require regulation.

And that distinction matters. The NHS App, in its current form, operates largely as a gateway to services. Booking appointments, ordering prescriptions and accessing records would not necessarily make the application itself a regulated medical device. However, the proposed triage system that assesses symptoms and directs patients towards particular services sits in a rather different space. Depending on its intended purpose and level of influence on clinical pathways, this triage component could fall under medical device regulations and require appropriate oversight, classification and registration.

Depending on its intended purpose and level of influence on clinical pathways, this triage component could fall under medical device regulations and require appropriate oversight, classification and registration.

I suspect that most patients would assume that any software providing healthcare advice has already undergone the same type of scrutiny that we expect of other clinical tools. Whether that assumption is justified will depend on the details of implementation.

The BBC article doesn't provide answers to the following questions, and it would probably be unreasonable to expect it to do so. However, there are a number of questions that regulators, clinicians, policymakers and healthcare organisations may need to address as AI-assisted triage systems mature.


These are not all encompassing, and are not technical questions for software developers alone. They are ultimately patient safety questions.

It's also worth noting that the BBC article does not state which specific technologies underpin the triage system. So the discussion that follows in this article draws on broader research into AI used within healthcare and should therefore be viewed as a general consideration rather than a comment on the NHS system itself.

One of the themes running throughout When Machines Speak Medicine was that organisations often become captivated by impressive outputs without fully understanding the mechanisms that generated them. The Nature study I discussed previously demonstrated that even advanced systems can be manipulated through misleading inputs and can subsequently produce convincing, but incorrect, outputs. The lesson there was not that AI is useless. The conclusion was that trust should be earned through evidence rather than assumed because the technology appears sophisticated.

The lesson there was not that AI is useless. The lesson was that trust should be earned through evidence rather than assumed because the technology appears sophisticated.

If we look further ahead, there are wider human factors that may need consideration as AI becomes more common within clinical and administrative workflows. One of these issues is the tendency for people to place greater trust in recommendations generated by technology. We see it regularly. A result appears on a screen and suddenly acquires an air of authority. Yet software recommendations should only ever be starting points, not conclusions.

One of the concerns raised during general discussions around AI governance is that people may become less likely to challenge a recommendation because it comes from a machine. Ironically, maintaining a healthy degree of scepticism may become even more important as these systems become more sophisticated.

In the same way, while the BBC article highlights potential efficiency gains from AI-generated documentation, the long-term success of such systems will depend on how they perform in day-to-day clinical environments and how any inaccuracies are identified and corrected. In theory, they could release valuable time back to clinicians. In practice, they will still require careful review. A subtle transcription error, a missed symptom or an inaccurate summary may appear minor at the time, but clinical records often follow patients for years.

A subtle transcription error, a missed symptom or an inaccurate summary may appear minor at the time, but clinical records often follow patients for years.

What I find interesting is how this discussion echoes themes from another article I wrote, Rise of the Doomsayer. There, I argued that debates about AI often become trapped between two extremes. Some people believe every new development will transform healthcare for the better. Others seem convinced that catastrophe is just around the corner. Neither position is particularly helpful - the reality is more nuanced.

There are obviously genuine benefits to be gained here. Access to services may improve. Administrative burden may reduce. Clinicians may, hopefully, gain more time with patients. Services may become more efficient. Given the pressures currently facing the NHS, these are outcomes that are clearly worth pursuing.

But we need to recognise that there are also genuine risks. Some of these concerns are already referenced in the BBC article, particularly around confidentiality, patient safety and digital exclusion. Others are broader considerations that may become increasingly relevant as AI-supported services expand over the coming years. We know that systems can make mistakes. People can become over-reliant on recommendations. Digital exclusion also remains a significant concern and deserves far more attention than it often receives. While many people are comfortable managing aspects of their healthcare through apps and online services, others are not. Older adults may face barriers related to digital literacy, visual impairment, reduced dexterity, cognitive difficulties, or simply a preference for speaking directly to another person. Some may not own a suitable device or have reliable internet access. As the population continues to age, these challenges affect a substantial number of patients rather than a small minority. If digital triage becomes a primary route into NHS services, alternative pathways must remain readily available to ensure that no group is disadvantaged.

If digital triage becomes a primary route into NHS services, alternative pathways must remain readily available to ensure that no group is disadvantaged.

Alongside these questions of accessibility, there are equally important issues surrounding privacy, transparency and governance. Patients need to have confidence that their information is held securely, that recommendations are subject to appropriate oversight, and that clear accountability exists when things go wrong. Ultimately, the success of any AI-enabled healthcare system will depend not only on its technical performance but also on the safeguards that surround it. Above everything, patient safety must remain the primary consideration rather than an afterthought.

The BBC article reports calls from healthcare organisations for patient safety and professional oversight to remain central to implementation. Building on those concerns, there are several areas that may warrant continued attention as adoption increases. Continuous monitoring will be needed. Systems should be tested not only when everything is working as expected, but also in unusual situations where errors are most likely to emerge. Traditional routes into healthcare must remain available for those who cannot or do not wish to use digital tools. If the triage component is functioning as a medical device, it should be subject to the same level of scrutiny we would expect for any other technology that may influence care decisions.

The NHS announcement should not be viewed as a choice between humans and technology. It should be viewed as a test of whether the two can work together safely.

The NHS announcement should not be viewed as a choice between humans and technology. It should be viewed as a test of whether the two can work together safely.

If there is one lesson that runs through much of my recent writing on AI, it's that technology is rarely the biggest risk. The bigger risk is misunderstanding what the technology can and cannot do. Used wisely, these tools could ease pressure on an overstretched healthcare system. Used carelessly, they could create new problems while solving old ones.

The challenge now is making sure we remember the difference.

As I argued in When Machines Speak Medicine healthcare is one of those environments where a single mistake can carry consequences far beyond the immediate interaction. In that context, the question is not whether AI can help. The evidence increasingly suggests that it can. The more important question is whether we have built the regulatory, clinical and governance structures needed to ensure that help arrives safely.

The BBC article focuses primarily on the proposed benefits of the rollout and the safeguards NHS England has indicated will accompany it. The questions raised here are not intended to suggest that these systems are unsafe. Rather, they reflect the kinds of governance, regulatory and patient-safety considerations that should continue to be examined as implementation progresses.

And that, at the end of the day, is where the real discussion should be. It's not how clever the software is. It's not how quickly appointments can be routed, and it's not about how many minutes are saved. The real discussion should be about whether patients are safer because of it. Have we put in the guardrails before the gadgets?