Walk any diagnostics trade floor this year and you will be sold artificial intelligence roughly every four metres. The connectivity box has AI. The middleware has AI. The analyser has AI, the quality module has AI, and somewhere near the coffee stand a pop-up promises an assistant that will read your control charts, write your competency records and, if you squint at the slide deck, more or less run the service while you sleep. The wave is real enough to count: by 2025 the US Food and Drug Administration had authorised more than 1,200 AI-enabled medical devices, up from around 690 at the end of 2023. But look closely at that list and roughly three quarters of it is radiology, and almost none of it is the near-patient chemistry and immunoassay work where I spend my days. It is a good time to be a vendor. It is a confusing time to be a point-of-care coordinator holding a budget and a clinical governance responsibility.
The uncomfortable part is this. Some of that noise is genuinely useful, arriving right now, and worth your attention. Some of it is a very confident piece of software laundering an unvalidated guess into what looks like a clinical decision. The two can appear identical in a demonstration. They stop looking identical the moment a result is wrong and someone has to explain why it was released.
So I want to skip the tired argument about whether AI will replace the laboratory. It will not, and asking the question that way tells you nothing useful. The question that actually earns its keep is narrower and harder: where does AI genuinely help in point-of-care testing, and where must a named human stay accountable no matter how good the software gets? My answer, which I will defend, is that the line is not drawn by capability. It is drawn by accountability and validation. Judge AI in point-of-care testing exactly as you would judge any other method: validated, monitored, and owned by a responsible person.
Where AI actually helps, today, without heroics
Strip away the marketing and a few genuinely valuable uses remain. They share a quiet family resemblance: the machine does the tedious, high-volume pattern work, and a human keeps the judgement.
The strongest near-term case is quality control. QC is repetitive, data-rich, and full of patterns that are easy to miss when you are covering four sites and a phone that will not stop ringing. Software is good at noticing that a control has drifted slowly across a fortnight, that a particular lot behaves differently from its predecessor, that one operator's runs cluster oddly, or that a shift is trending toward a boundary before it ever breaches a rule. This is not vapourware. UK laboratory scientists have been testing machine learning on patient-based real-time quality control, using the flow of ordinary patient results to catch a drifting assay hours before a scheduled control round would, and the honest reviews of this work describe it as augmenting the existing rules rather than retiring them. That is exactly the right posture. A tool that surfaces "this control is drifting, come and look" a day earlier than a busy human would have is a real gift, provided it flags for a person rather than deciding on its own.
The second honest use is documentation and structure. POCT drowns in paperwork: competency assessments, training logs, incident notes, audit narratives, the endless prose that accreditation demands. Language models are genuinely good at drafting, summarising and reshaping this kind of text. Hand a tool your rough notes and get back a tidy first draft of a competency record or an incident summary, and you have saved real hours. The key word is draft. A human reads it, corrects it, and signs it, because a signature means something.
The third is triage of workload. In a busy service the scarce resource is attention. Software that ranks what needs a human eye first, the overdue calibration, the operator whose competency lapses next week, the site that has not run a control in three days, is doing dispatch, not diagnosis. That is a safe and welcome place for automation.
The fourth is the natural-language front door. Operators forget procedures, especially the rare ones. A well-grounded assistant that answers "how do I run a QC on this device" in plain words, drawn from your own approved SOPs, can be a better on-shift help than a binder nobody opens. It supports the operator; it does not replace their training. That is why structured, assessed competency still sits at the heart of everything we teach in our POCT training, and why a good digital tool complements it rather than substituting for it.
The safe uses of AI in point-of-care testing all share one trait: the machine does the pattern-spotting, and a named human keeps the judgement.
A simple map of what fits where
It helps to sort proposed uses by risk rather than by how impressive the demo looked. Not every task carries the same consequence if the software is wrong, and your governance should follow the consequence, not the marketing.
In the green band sit the low-risk assistive uses: drafting a document a human will check, flagging a QC drift for a person to review, surfacing overdue tasks. If the software is wrong here, a human catches it in the normal course of work and nothing reaches a patient. These are the wins you can adopt with a light touch.
The amber band is assistance that must never run without a person in the loop: suggesting an interpretation of a result, prioritising which cases a clinician reviews, proposing a course of action. Useful, yes, but only as a prompt to a qualified human who remains free to disagree and is expected to. Amber uses need explicit oversight, an audit trail, and a culture where overriding the machine is normal and never punished.
The red band is short and non-negotiable. No autonomous release or interpretation of a patient result without a responsible human. No system that turns a raw number into a clinical conclusion and acts on it with nobody accountable in between. This is not conservatism for its own sake. It is where the regulatory and ethical weight of the whole field sits.
Why the red line is not optional
There is a tendency to treat the red band as an old-fashioned worry that better models will eventually retire. That misreads the situation. The line does not exist because today's software is not clever enough. It exists because of two facts that do not move when the model improves.
The first is regulatory. Software intended for a medical purpose can itself be a medical device. In the UK, the MHRA regulates software and AI as a medical device, and it is actively tightening the frame: stronger post-market surveillance duties came into force in 2025, and its AI Airlock, described as the world's first regulatory sandbox for AI as a medical device, is feeding a dedicated AI framework expected in 2026. Where a tool interprets a result or drives a clinical decision, its classification and validation are not paperwork you can skip. A clever assistant that starts quietly influencing clinical decisions has crossed a threshold, and pretending it has not does not make the regulator's view go away. This is precisely the sort of question worth taking advice on before you deploy, not after an incident.
The second is professional accountability, and it is baked into the standard many of you are accredited against. ISO 15189:2022, which now brings point-of-care testing inside the main laboratory standard and has superseded the withdrawn ISO 22870, keeps a named, competent human accountable for the quality of results. Automation does not dissolve that accountability; it just changes what the accountable person is responsible for overseeing. You cannot delegate the duty to a model, because a model cannot be held to account. When something goes wrong, an inspector, a coroner or a patient's family will ask who was responsible, and "the algorithm decided" is not an answer that protects anyone.
The red line is not about how clever the software is. It is about who answers when a result is wrong, and software cannot answer.
The specific danger: laundered confidence
The failure mode I worry about most is not a model that is obviously wrong. Obvious wrongness gets caught. The dangerous case is a model that is fluent, calm and wrong, and that presents its guess with the same clean confidence it uses for its correct answers.
We now have hard evidence that trained clinicians follow that confidence off a cliff. In a randomised trial published in NEJM AI, researchers took 44 physicians who had already completed formal AI-literacy training, gave them clinical vignettes, and let them consult a large language model. When the model's advice was sound they did well. When the model was deliberately fed wrong advice on some of the cases, mean diagnostic accuracy fell from 84.9 percent to 73.3 percent, and their single top-choice diagnosis dropped from 90.5 percent to 76.1 percent. These were not naive users being fooled by a novelty. They had been trained to interrogate the machine, and a confident wrong answer still cost them roughly fourteen percentage points of accuracy.
A human sceptic hears "the assay is probably fine" and asks how we know. A tidy dashboard that simply displays "within expected range" invites nobody to ask. That is laundered confidence: an unvalidated inference dressed in the visual language of a validated result. It is most dangerous exactly where POCT is most valuable, at the point of care, in the hands of a non-laboratorian who has been trained, quite reasonably, to trust the device in front of them.
This is why transparency is not a nice-to-have. A tool that flags a QC drift should be able to show you what it saw. An assistant that suggests an interpretation should show its working and its sources, and should be visibly comfortable saying "I am not sure, check this". A tool that cannot show you why it reached an answer is far harder to validate, to monitor and to defend, and one you cannot stand behind when challenged has no business anywhere near a patient result. Understanding what each analyte actually means, its interferences, its failure modes, its clinical weight, is what lets a competent human catch the confident error. That literacy, which we build out across our analyte guides, is not made obsolete by AI. It is what makes AI safe to use.
Judge it like a method, not like a miracle
Here is the reframe that cuts through most of the hype. You already know how to introduce a new capability responsibly, because you do it every time you bring in a new analyser or a new assay. You do not take the manufacturer's word for it. You validate against your own population, you run it in parallel, you set acceptance criteria, you monitor performance over time, and you keep a named owner. Do exactly that with AI.
The cautionary tale here comes from a hospital rather than a dispensary bench, but it teaches the point-of-care lesson perfectly. The Epic Sepsis Model is one of the most widely deployed clinical prediction tools in the world, wired into electronic records across hundreds of hospitals, firing alerts that a patient may be turning septic. Its manufacturer reported an area under the curve of 0.76 to 0.83, which reads like a solid method. Then an independent team externally validated it against 38,455 hospitalisations at their own institution. The area under the curve came out at 0.63. It caught only a third of sepsis cases, and just 12 percent of the alerts it raised were genuine sepsis, drowning clinicians in false alarms. Nothing in the software had broken. It had simply never been validated in the setting where it was being trusted.
An AI feature is a method. Treat it like one.
- Validate it before you trust it. Test the tool against known cases and edge cases from your own setting, not the vendor's demo data. Decide in advance what "good enough" means and what you will do when it is not. The sepsis model above passed its maker's numbers and failed the hospital's.
- Monitor it after you deploy it. Model performance can drift as populations, devices and reagents change, just as an assay drifts. A tool validated two years ago cannot be assumed still valid today without ongoing checks. Build in ongoing review and change control.
- Keep a human owner. Every AI feature needs a named person accountable for it, with the authority and the confidence to switch it off. If nobody owns it, nobody is watching it.
- Demand transparency. If a tool cannot show you why it flagged, suggested or scored something, you cannot validate it and you cannot defend it. Treat opacity as a red flag, not a trade secret.
- Preserve the override. The human must be able to disagree easily, and doing so must never be penalised. The moment overriding the machine feels like insubordination, your safety margin is gone.
Notice that not one of these principles is exotic. They are the ordinary disciplines of a well-run laboratory, applied to a new kind of method. If a vendor's AI cannot survive being treated like a method, that tells you something important about the vendor's AI.
The flow that keeps you safe
All of this collapses into a single picture. However clever the assistance, the shape of a safe workflow does not change. The result enters, the AI flags or assists, a responsible human decides, and the decision is recorded. Accountability never leaves the human.
Keep that shape and most AI in POCT becomes something you can adopt calmly. Break it, by letting the assist quietly become the decision, and you have not modernised your service. You have removed the one control that made it defensible.
What this means for your service
If a vendor is waving AI at you, or you are quietly wondering whether to build or buy some, here is what I would actually do.
- Start where it is boring and safe. Adopt AI first for documentation drafting, QC drift flagging and task triage. Low risk, real hours saved, human always in the loop. Bank those wins before you go anywhere near interpretation.
- Sort every proposed use into green, amber or red before you buy anything, using the map above. If a feature lives in red, it does not go in, no matter how good the demo was.
- Ask the vendor three questions. How was this validated, how will I monitor it in my setting, and what does it show me when it is unsure. Vague answers are your answer.
- Write AI into your quality system, not around it. Each tool gets a named owner, a validation record, a monitoring plan and a kill switch, documented like any other method. A little structure here saves a lot of grief later, and our POCT templates and wider resources give you a place to start.
- Invest in the humans, not just the software. The whole system depends on a competent person who can catch a confident error. Start that competence with the free POCT Fundamentals course and keep it current. If you are unsure how to govern AI in a diagnostics setting, our consultancy exists for exactly this kind of question.
The line, restated
AI is arriving in point-of-care testing, and a good deal of it is genuinely useful. It will spot drift you would have missed, draft the paperwork you dread, and point your attention where it matters most. Let it. But keep hold of the one thing no model can carry for you. When a result goes out and shapes a clinical decision, a named, competent human stands behind it. That is not a limitation to be engineered away as the technology matures. It is the point. Use the machine for the pattern-spotting. Keep the judgement, and the accountability, human.
Sources and notes
The figures in this article are drawn from peer-reviewed studies, regulator guidance and published device statistics. The automation-bias trial and the sepsis-model validation are the two load-bearing datasets; both are external, independent and clinical rather than vendor material. The sepsis example is deliberately from a hospital electronic record, not a pharmacy or POCT device, because it is the clearest published case of a confidently marketed model underperforming on independent validation. The risk-band map in Figure 1 and the workflow in Figure 4 are conceptual frameworks, labelled illustrative, not measured datasets. Device counts are marketing authorisations, not a census of every AI tool in clinical use.
- NEJM AI. Automation Bias in Large Language Model-Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy, a Randomized Clinical Trial, 2025. Preprint at medRxiv 2025.08.23.25334280. Randomised trial, 44 physicians, 6 vignettes, 3 with seeded errors.
- Wong A and colleagues. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 2021 (n=38,455 hospitalisations; AUC 0.63; sensitivity 33 percent; PPV 12 percent).
- MHRA. Software and artificial intelligence (AI) as a medical device, including the change programme, strengthened post-market surveillance and the AI Airlock sandbox.
- International Organization for Standardization. ISO 15189:2022, Medical laboratories, requirements for quality and competence, which brings point-of-care testing inside the laboratory quality standard and supersedes the withdrawn ISO 22870.
- Lorde N, Mahapatra S, Kalaria T. Machine Learning for Patient-Based Real-Time Quality Control, Analytical and Pre-analytical Error Detection in Clinical Laboratory. Diagnostics, 2024.
- US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices, the FDA's list of authorised devices.
- npj Digital Medicine. How AI is used in FDA-authorized medical devices, a taxonomy across 1,016 authorizations. 2025, source for the radiology concentration.
