There are three ways to handle patient calls at scale: a fully human team, a fully automated AI voice system with no human fallback, and a hybrid model where AI handles routine volume and escalates to people on defined triggers. The AI receptionist vs human debate is usually framed as a straight choice between the first two. That framing is the mistake. Judged across the six dimensions that actually matter to an NHS buyer, only one of the three models holds up under all of them, and it is not the one the loudest vendors on either side are selling.

This is the comparison to make before any contract is signed, whether you are a GP practice manager, a PCN operational lead, a trust outpatient lead, an ICB procurement lead, or an ISP head of operations. The six dimensions are cost per effective call hour, clinical safety architecture, patient experience parity across cohorts, handling of the patients who cannot use the NHS App, compliance footprint, and scalability under peak demand. Each model has a real strength and a real failure mode. Naming both honestly is the point, because a procurement committee can see through a comparison that does not.

Before the dimensions, it is worth being clear about what each model actually is, because the labels carry assumptions that do not always hold.

The three models, defined honestly

The fully human model is the status quo for most services: an in-house reception or booking team, often topped up with agency or bank staff at peak. Its strength is judgment. A person handles nuance, reads distress, manages a sensitive conversation, and holds a relationship in a way no system does. Its constraint is that human capacity is fixed and expensive, and it does not flex to demand.

The fully automated model is AI voice handling calls end to end with no human escalation path. Its strength is that it scales and it is cheap per call. Its failure mode is what happens when a call exceeds what the system can safely handle, because there is no one to hand it to. A model with no escalation route is making a bet that nothing important will ever fall outside its scope, and in a clinical setting that is not a safe bet.

The hybrid model is AI-first with human escalation on defined triggers. AI absorbs the high-volume routine calls. Anything that meets an escalation trigger, clinical sensitivity, complexity, a red flag, a patient who needs a person, routes to the human team with the context already captured. Its strength is that it matches each call to the right handler. Its failure mode is that it is only as good as its escalation design, its integration, and the human layer behind it. A badly designed hybrid is just a fully automated system with a broken handoff. A well-designed one is the only model that covers all six dimensions.

With those definitions set, here is how the three compare on each dimension that a buyer is accountable for.

Dimension 1: Cost per effective call hour

The fully human model carries the highest and least controllable cost, and the national context makes this sharper than it used to be. The NHS spent around £3 billion on agency staff in 2023/24, and providers have since been ordered to cut agency spend by 30% and to eradicate it altogether, with agency use for band 2 and 3 roles required to end by January 2026. The lever most services reached for to cover peak call demand is the lever being removed. Fixed headcount also means paying for capacity whether or not the calls arrive, and still being short when they do.

The fully automated model is cheapest per call on paper. The cost it hides is the unhandled complexity. A call the system cannot resolve does not disappear. It becomes a patient who calls back, complains, or comes to harm, and those costs land elsewhere on the balance sheet where they are harder to see.

The hybrid model is the only one where you pay for effective handled volume rather than for seat time or for calls that fail. The routine volume is absorbed by the AI layer, the escalations go to a human team that is no longer drowning in routine, and the unit cost reflects calls actually handled. As a reference point, a hybrid model priced per effective call hour can run at single-figure pounds per hour, with forwarded, abandoned, and failed-identification calls carrying no charge, which is a cost structure neither of the other models can match.

Dimension 2: Clinical safety architecture

This is the dimension where the fully automated model is weakest and where buyers should press hardest. Detecting a red flag is necessary but not sufficient. The question is what happens next. A system that identifies a safety concern but has no human to escalate to has detected a risk it cannot act on. In an NHS setting, that is the single most important gap in the voice-only model, and it is the one its vendors tend to talk about least.

The fully human model has judgment but inconsistent application under load. A reception team handling a Monday morning surge is more likely to miss a subtle red flag not because the staff lack skill but because the conditions, volume, time pressure, interruption, work against careful triage. Human judgment is the gold standard for the individual call and an unreliable safeguard at peak volume.

The hybrid model is built around the handoff. Red flags are detected in real time and escalated to a clinician or a member of staff, with the captured context attached, so the human judgment is applied to the calls that need it rather than spread thin across all of them. Built properly, to the NHS clinical risk management standard DCB0129, the safety architecture is designed in rather than assumed. The escalation route is the safety net the voice-only model does not have and the surge-pressured human model cannot reliably provide.

Dimension 3: Patient experience parity across cohorts

Parity is the dimension most comparisons skip, and it cuts against easy assumptions in both directions. The fully human model is assumed to deliver the best patient experience, and for the patient who gets through, it often does. The catch is who does not get through. A patient who hits a busy tone at 8am, gives up, and is not seen has had the worst possible experience, and that experience falls hardest on the patients who cannot easily call back: people working, people caring, people who find the system hard to navigate. A capacity ceiling is its own inequality.

The fully automated model delivers uniform experience, which sounds like parity but is not. Uniform handling fails the patients who need something other than the standard path: the distressed caller, the patient with a complex request, the patient whose first language is not English and who needs the conversation to flex. Treating every caller identically is not the same as treating every caller fairly.

The hybrid model routes by need. The routine caller gets fast, consistent handling. The caller who needs a person gets escalated to one. That is the closest any model gets to genuine parity, because it matches the handler to the patient rather than forcing every patient through the same channel. The team is still the team. What changes is that the people who most need a human are more likely to reach one, because the human team is no longer fully occupied by routine volume.

Dimension 4: Handling the patients who cannot use the NHS App

Roughly 20% of the UK population, around 11 million people, lack basic digital skills or do not use digital technology at all, and that group skews older, more deprived, and in poorer health. The much-cited 5% who cannot use the NHS App is the floor, not the ceiling, of digital exclusion. For these patients, the phone is not one channel among several. It is the channel.

This is where digital triage tools, however good, structurally exclude a cohort: if the route in requires a smartphone and the confidence to use it, the patients without either are shut out. The phone channel is the equity channel, which means the model that handles phone calls well is the model that serves the excluded cohort.

The fully human model serves these patients well in principle but is capacity-limited in practice, so the busy-tone problem hits the digitally excluded hardest, because they have the fewest alternatives. The fully automated model covers the phone channel, which is a genuine advantage over app-only triage, but a rigid voice system can struggle with exactly the patients who need flexibility. The hybrid model covers the phone channel and escalates the patients who need a person to a person. It is the model that keeps the equity channel open and staffed, which for an ICB or a practice accountable for health inequalities is not a soft consideration. It is a procurement criterion.

Dimension 5: Compliance footprint

Compliance is where the voice-only model’s missing escalation path becomes a documentation problem as well as a safety one. The NHS standards that apply, DCB0129 for clinical risk management, DTAC for digital technology assessment, and UK GDPR for data handling, all assume a clear account of how risk is managed, including what happens when something falls outside normal handling. A model with no escalation route has a weaker answer to that question by design.

The fully human model has a well-understood compliance footprint, though data handling and governance vary with how agency and bank staffing are managed. The fully automated model has the lightest operational footprint and the hardest compliance story, because the safety architecture that compliance depends on is the part it lacks. The hybrid model, built to DCB0129, DTAC, and GDPR from the start, has the most defensible footprint, because the escalation and safety architecture that compliance frameworks ask about is the same architecture the model is built on. For an IG lead or a procurement committee, the model with the clearest answer to “what happens when a call needs a human” is also the model with the clearest compliance position.

Dimension 6: Scalability under peak demand

This is the dimension the fully human model cannot win, and it is the one that drives most of the others. Patient call demand is not flat. It spikes at 8am, surges in winter, and concentrates during flu and recall campaigns. Fixed human headcount cannot flex to those peaks, which is the structural capacity mismatch at the root of the whole problem. Adding staff for the peak means paying for them through the trough, and the agency lever that used to cover the gap is being withdrawn.

The fully automated model scales without limit, which is its genuine strength, but it scales by handling everything the same way, including the calls it should not be handling alone. Infinite scale with no safety net is not the same as solving the problem.

The hybrid model scales the part that should scale, the routine absorption, and escalates the residual to a human team that is now sized for the calls that actually need people rather than for the full peak. At 8am, the AI layer absorbs the surge of routine appointment and admin calls in parallel, and the reception or booking team handles the smaller volume of escalations with the context attached. The peak is absorbed without holding fixed headcount against a spike that only happens for two hours a day. This is the dimension where the hybrid model’s advantage is largest, because it is the dimension the human model is structurally unable to address.

The pattern is consistent. The fully human model owns judgment and relationship and loses on cost and scale. The fully automated model owns cost and scale and loses on safety and parity. The hybrid model is the only one without a losing column, provided the escalation and safety architecture are real rather than cosmetic. That proviso is the whole game, which is why the question to ask any vendor is not whether they use AI but what happens to the call the AI should not handle alone.

Where Jackie sits

Jackie is the hybrid model built to NHS clinical safety standards. AI handles the routine call volume as concurrent capacity alongside the existing team, detects red flags in real time, and escalates to reception, booking, or clinical staff on defined triggers with structured context attached. It is built to DCB0129, assessed against DTAC, and run on a per-effective-call-hour basis, which is the cost structure the hybrid column above describes. The reception and booking teams stay in place. What changes is that the routine volume that was bottlenecking the queue is absorbed in parallel, and the people who most need a human are more likely to reach one.

The evidence is in the deployments. At Park Street Surgery, the hybrid model absorbed 81% of inbound calls, ran a 91.1% completion rate, handled more than 1,200 patient calls with zero outages, and recovered 51 hours of staff time in eight weeks. Headcount did not reduce. The composition of the team’s day changed. That is what the hybrid model looks like in practice rather than on a comparison table.

Make the comparison on your own service

The model you choose is the argument you will have to defend at a partnership meeting, an exec team, or a procurement committee, and the hybrid case is the one that holds across all six dimensions when the escalation and safety architecture are real. A 20-minute demo shows how the AI-first, human-escalation model runs on a live call queue, how red-flag escalation works, and how the per-effective-call-hour cost compares with your current handling. The four-week pilot runs on your existing telephony with no hardware, which is how the Park Street figures were produced in the first place.

Frequently asked questions

Is an AI receptionist better than a human receptionist?

The framing is the wrong one. A human handles judgment, sensitivity, and relationship better than any system. AI handles volume and scale better than any team. Neither alone covers all six dimensions an NHS buyer is accountable for. The hybrid model, AI for routine volume with human escalation on defined triggers, is the only one of the three without a losing column, because it matches each call to the right handler rather than forcing a single model onto every call.

What is the risk of a voice-only AI with no human escalation?

The risk is the call the system cannot safely handle. A voice-only model with no escalation path can detect a red flag but has no one to act on it, and it forces complex or distressed callers through the same path as routine ones. In a clinical setting, the absence of a human fallback is the most important gap to probe before signing a contract.

Which call handling model is best for patients who cannot use the NHS App?

The phone channel is the equity channel for the roughly 20% of people who lack basic digital skills, a group that skews older, more deprived, and in poorer health. Digital triage tools exclude them by design. A hybrid model keeps the phone channel open, handles routine calls at scale, and escalates the patients who need a person to a person, which keeps the equity channel both open and staffed.

How does a hybrid model handle the 8am peak?

The AI layer absorbs the surge of routine appointment and admin calls in parallel, so the volume hitting the human team is the smaller set of calls that need a person, with context attached. The peak is absorbed without holding fixed headcount against a spike that only lasts a couple of hours, which is the capacity mismatch the fully human model cannot solve.

What should NHS buyers ask vendors before signing?

Ask what happens to the call the AI should not handle alone. The answer reveals whether the model is genuinely hybrid or a voice-only system with a cosmetic handoff. Then ask how the safety architecture maps to DCB0129 and DTAC, and how the cost is charged: per seat, per call, or per effective call hour. Those three questions separate a defensible model from a risky one.