A patient arrives on an acute ward feeling unwell. A finger-prick sample goes onto the ward's point-of-care analyser, and the number that comes back sits just on the wrong side of a decision threshold. The team acts: a treatment is started, an escalation is made, a call goes to the on-call registrar. A few hours later a venous sample drawn at the same time reaches the main laboratory, runs on a large automated platform, and the number that comes back sits just on the right side of the same threshold. Same patient, sampled minutes apart, same analyte, two devices, two verdicts. Somebody now has to decide which number the patient is.

Most people who work in point-of-care testing have felt this moment, even if they have never given it a name. A ward result and a lab result that "should" agree, but do not, quite. A value that trends beautifully until the day it was measured on a different machine and then jumps for no clinical reason. A reference range printed on a report that was quietly borrowed from a different method. We tend to file these under "point-of-care is less precise" and move on. That is not wrong, but it is not the real story either.

The real story is this, and it is the thesis I will defend: the same blood tested on two different point-of-care platforms can legitimately give two different numbers, and near a clinical decision threshold that difference can change what happens to the patient. Harmonisation and standardisation across methods is unfinished business. Method-dependent reference ranges and decision limits are a patient-safety issue hiding in plain sight, and until the science is finished, clinical caution near thresholds is the safety net we cannot switch off.

±15%how far ISO 15197 lets a conforming glucose meter sit from the laboratory value
1.11the factor needed just to turn a whole-blood glucose into a plasma number
13.6 to 22.2ng/L, the same troponin T assay's "abnormal" line for women versus men
8 of 27analytes where method bias blocked a shared reference interval in a harmonisation study

Why two analysers can disagree about the same blood

It is tempting to assume that measuring a glucose or a haemoglobin is a fixed physical fact, like weighing a bag of flour, and that any well-made device should land on the same answer. That intuition is the trap. Two analysers are not two scales reading the same weight. They are frequently two different chemistries, answering a slightly different question, then expressing the answer on a scale that may or may not be tied back to the same reference.

There are a handful of reasons the numbers drift apart, and each one is a real, recognised phenomenon in laboratory medicine rather than a fault to blame on a single machine.

Different measurement methods

The same analyte can be measured by genuinely different reactions. One device may use an enzymatic method, another an electrochemical sensor, another an immunoassay built around a particular antibody. Each has its own response, its own interferences and its own quirks. An immunoassay depends on which part of the target molecule its antibody recognises, and two antibodies raised against the "same" marker can see different things in the same sample. Cardiac troponin is the classic case: assays from different manufacturers are not interchangeable, and a joint working group of the professional bodies has had to publish consensus, assay-specific decision limits precisely because no single number works across them all.

Calibration and traceability

A raw signal, a colour change or a current, means nothing until it is converted into a concentration. That conversion depends on how the device was calibrated and, crucially, on what its calibration is traceable to. Where a recognised reference material and reference method exist, and where every manufacturer ties its calibration back to that same anchor, results converge. HbA1c is the success story here. After the IFCC built a reference measurement procedure and the field agreed to make every manufacturer traceable to it, results became comparable enough to report on a common scale, and yet it still takes a published master equation, NGSP percent equals 0.09148 times the IFCC number plus 2.152, to move between the two ways of expressing the very same sample. Where that chain is incomplete, two correctly working devices can sit on subtly different scales for reasons that have nothing to do with the sample in front of them. Building the scaffolding that prevents this, reference procedures, commutable reference materials and a network of reference laboratories, is exactly what bodies such as the Joint Committee for Traceability in Laboratory Medicine exist to do.

Matrix and sample type

Point-of-care testing often runs on whole blood from a fingertip; the laboratory often runs on plasma or serum from a vein. The matrix matters. Whole blood and plasma are not interchangeable, and for glucose the widely used convention is to multiply a whole-blood value by about 1.11 to reach the plasma-equivalent number, because glucose sits in the water fraction and plasma holds more water than red cells do. A meter that reports a plasma-equivalent value from a capillary drop is doing that conversion for you, on assumptions about haematocrit that are not always true. Add the ordinary variability of capillary sampling, a squeezed finger, a first drop, a cold hand, and you have another honest source of difference between the ward number and the lab number.

Two analysers are not two scales reading the same weight. They are two different chemistries answering a slightly different question, then expressing the answer on a scale that may not share the same anchor.

None of this is exotic. Different results from different methods for the same analyte is a long-standing, recognised problem in laboratory medicine, and it is precisely why international bodies run formal harmonisation and standardisation programmes to pull methods back towards a common answer. Those programmes exist because the problem is real and the work is not finished.

Why it matters most exactly at the threshold

Here is the part that turns a measurement-science curiosity into a clinical-safety question. A difference between two devices is usually clinically irrelevant. If one machine reads a value comfortably in the healthy range and another reads it a little differently but still comfortably in the healthy range, nobody is harmed and nothing changes. The difference only bites in one place: at a decision threshold, where a small numerical gap flips the clinical action from one thing to its opposite.

Clinical medicine is full of these lines. Above this value we treat; below it we reassure. Above this value we admit; below it we discharge. Above this value we anticoagulate, transfuse, escalate, repeat urgently; below it we watch and wait. The threshold is a cliff edge drawn across a continuous number, and a patient whose true value sits near that edge is in the danger zone, because now the ordinary, honest disagreement between two devices is enough to put them on different sides of the cliff.

One sampleDevice ADevice BDevice CdecisionthresholdABCsame blood, different verdict
Figure 1. One blood sample, split to three point-of-care devices, each returning a slightly different value against the same dashed decision threshold. Two devices read "below", one reads "above". Same patient, different action. Illustrative schematic of the harmonisation gap discussed in the text.

Put a real number on it. Glucose meters are held to ISO 15197:2013, which counts a meter as accurate if at least 95% of its results fall within 15% of the laboratory value once glucose is above 100 mg/dL, roughly 5.6 mmol/L. Read that carefully. Two meters can both pass, both be working exactly as designed, and both sit up to 15% either side of the laboratory comparator value, which means they can differ from each other by close to a third. Around a true plasma glucose of 6.0 mmol/L, one conforming meter can honestly report 5.5 and another 6.5. Drop an action line at 6.0 between them and the same patient is either reassured or treated depending on which meter happened to be on the bench.

4.55.05.56.06.57.07.5Plasma glucose (mmol/L)action line at 6.0ISO 15197 allows 15% either sidewhat one conforming meter may report5.56.5ReassureActTwo meters that both pass the standard, opposite verdicts
Figure 2. The tolerance a single conforming glucose meter is allowed under ISO 15197:2013, 15% either side of the laboratory value, drawn around one true reading of 6.0 mmol/L. Two meters that both pass the standard can legitimately land on opposite sides of an action line. Allowance: ISO 15197:2013 (95% of results within 15% at glucose at or above 100 mg/dL). Plotted values illustrative.

Neither meter is broken. Neither clinician is wrong given the number in front of them. The measurement spread of a single true value has straddled the line, and the patient's care has forked on which machine was in the room. Multiply that by every threshold, every ward, every clinic and every device in a health system, and the scale of the exposure becomes clear.

When the threshold itself moves

It gets one turn more subtle, because the line is not always fixed either. For some analytes the decision limit is itself method-specific and population-specific, so two services can be reading against genuinely different cut-offs while believing they share one. High-sensitivity cardiac troponin is the sharpest example. The threshold for diagnosing myocardial injury is the 99th percentile of a healthy reference population, and that value is not standardised across assays. It is not even a single number within one assay. When a US group took one widely used high-sensitivity troponin T assay and worked the same reference dataset through nine accepted statistical methods, the overall 99th percentile came out around 19.2 ng/L, but roughly 13.6 ng/L in women and up to 22.2 ng/L in men. One assay, one reference dataset, and the line that separates "injury" from "normal" moves by more than half depending on which sex-specific and statistically derived limit you apply.

99th percentile (ng/L)102013.6 ng/LWomen19.2 ng/LOverall22.2 ng/LMen+63% by sexSame assay, same population, nine calculation methods
Figure 3. The 99th percentile upper reference limit for one high-sensitivity cardiac troponin T assay, derived from the same population with nine accepted statistical methods: about 13.6 ng/L in women, 19.2 ng/L overall, and up to 22.2 ng/L in men. The decision line itself is method and population dependent. Data: Fitzgerald and colleagues, Clinica Chimica Acta, 2020 (fifth-generation cardiac troponin T).

The quiet unsafety of a borrowed reference range

If method-dependence stopped at the raw number, it would be manageable. It does not. The deeper trap is that reference intervals and clinical decision limits are themselves method-specific. A reference range is not a property of the human body alone; it is derived on a particular method, in a particular population, and it only means what it says when it is read against a result produced by that same method.

This is where services get quietly caught out. A range gets copied from the laboratory's report, or from one device's documentation, and pasted onto another device because the analyte "is the same". A decision limit lifted from a guideline gets applied to a point-of-care result without anyone checking which method that limit was established on. The number looks authoritative. The line looks official. But a range derived on one method and applied to another can be misleading, flagging normal results as abnormal or, worse, reassuring on results that the correct range would have flagged. Even when the profession sets out deliberately to build shared reference intervals, it cannot always be done: in one international harmonisation exercise only 19 of 27 analytes agreed closely enough across eight platforms to carry a common interval, which left eight where method bias made a shared range unsafe. The interface reads clean. The clinical logic underneath it can be subtly wrong.

And it compounds the moment results are trended. Trending assumes the numbers are commensurable, that a 9 last week and a 10 this week describe a real change in the patient. Trend a value across a ward device and then the lab, or across two different point-of-care platforms, as if the numbers were interchangeable, and you can manufacture a "change" that is nothing but a method difference. You can equally hide a real deterioration behind a method offset pointing the other way. The graph looks smooth and convincing, while quietly comparing numbers that were never on the same scale.

Whose job is it to fix this

It is fair to ask why the profession has not simply solved this. The honest answer is that it is being solved, slowly, and by more than one party. Harmonisation is a shared responsibility, and no single actor can deliver it alone.

Manufacturers carry the first duty: to tie calibration to recognised reference materials and methods wherever they exist, to be transparent about the method a device uses, and to state the reference intervals and limits their method actually supports rather than borrowing someone else's. External quality assessment schemes, run through bodies such as UK NEQAS and WEQAS, expose method-to-method differences by showing how the same material reads across the field, which is one of the few mechanisms that makes the problem visible at all. Standards bodies and professional colleges drive the reference-method and reference-material work that gives everyone a common anchor to aim at.

But the last mile belongs to the service that runs the devices. Standards are explicit about this. ISO 15189:2022 requires, in its clause on comparability of results (7.3.7.4), that where the same measurand is examined by more than one procedure, device or site, results are demonstrably comparable across them, and it goes further than the standard it replaced by asking laboratories to review the clinical impact of any differences on reference intervals and decision limits. That is not a nice-to-have; it is a stated expectation of an accredited service. Comparability does not happen by assumption. It has to be checked, and only the service can do the checking, because only the service knows which devices its patients are actually being measured on.

A reference range is not a property of the human body alone. It is derived on a particular method, and it only tells the truth when it is read against a result from that same method.

What this means for your service

None of this argues against point-of-care testing. Fast results in the room can change management for the better and spare patients repeat journeys, and a well-run POCT service is a genuine clinical asset. It argues for one specific discipline: treat the number as method-dependent, and be most careful exactly where it matters most, at the threshold. Here is what that looks like in practice.

  • Know your method. For every analyte on every device, know the measurement method it uses, what its calibration is traceable to, and which reference range and decision limits are correct for that method. If you cannot answer this for a device, you do not yet know what its numbers mean.
  • Use method-appropriate reference ranges, never borrowed ones. Do not copy the lab's range, or another device's range, onto a platform without confirming it applies. A range is part of the method, not a universal fact about the analyte.
  • Do not mix devices in a trend without a comparison. Before you plot results from two platforms on one line, know how those platforms compare for that analyte. If you have not established comparability, annotate the change of device on the record and treat the step with suspicion.
  • Run split-sample comparisons. Periodically measure the same sample on your point-of-care device and on the laboratory method, and look at the agreement, paying particular attention to values near clinical thresholds. This is the concrete way you meet the ISO 15189 comparability expectation rather than assuming it.
  • Escalate values near a threshold; do not trust a single device. Good anticoagulation practice already points this way. Point-of-care and laboratory INR agree well through the therapeutic range, but clinically important differences become more likely once the INR climbs above about 3.0, so many services confirm a point-of-care INR over 5.0 against the laboratory and base the warfarin change on the lab number. Borrow that habit for any borderline result: repeat, cross-check on the laboratory method, or interpret alongside the clinical picture. A borderline number from one machine is the weakest evidence in the whole system, and it is exactly where a single device should not be the final word.
  • Train the people reading the numbers, not just the people running the devices. The clinician acting on a borderline ward result needs to understand that the threshold and the method are linked. This is competency, and it belongs in your training and governance, not in folklore.

If you want help putting this on a formal footing, our consultancy designs method-comparison and split-sample programmes aligned to the ISO 15189 comparability requirement. Our training, including the free POCT Fundamentals course, covers why method and reference range travel together and how to read a result with the threshold in mind. And our analyte guides set out, analyte by analyte, why method matters for interpretation, so the people using the numbers understand what the numbers actually are.

The number is not the patient

Two devices, one patient, two numbers, and a threshold drawn between them: that is the situation, and it is not going away until harmonisation is finished, which it is not. The safest clinicians and the best-run services already behave as if this were true. They know their method, they respect the threshold, and near the line they confirm rather than commit. Until the science closes the gap, that caution is the safety net. The number on the screen is a measurement of the patient. It is not the patient. Keep the two apart, especially at the edge of a decision, and the difference between two good devices stops being a hazard and becomes just what it is: the honest, unfinished state of the art.

Sources and notes

The figures in this article are drawn from published standards, professional-body position papers and peer-reviewed sources. The glucose tolerance in Figure 2 is the accuracy criterion set by ISO 15197:2013; the 5.5 and 6.5 meter readings around a 6.0 mmol/L line are an illustrative worked example that sits inside that tolerance, not a measured dataset. The troponin values in Figure 3 are real 99th percentile results for one fifth-generation cardiac troponin T assay, and the point is precisely that they differ by sex and by statistical method. Figure 1 is a schematic. Where I describe a decision threshold, I have used clinically recognisable lines to make the argument concrete rather than to issue clinical guidance; use the reference ranges and limits validated for your own devices.

  1. International Organization for Standardization. ISO 15197:2013, in vitro diagnostic test systems: requirements for blood-glucose monitoring systems for self-testing. The criterion that 95% of results fall within 15% of the laboratory value at glucose at or above 100 mg/dL.
  2. Pardo and colleagues. The quantitative relationship between ISO 15197 accuracy criteria and mean absolute relative difference. J Diabetes Sci Technol, 2016.
  3. NGSP. IFCC and NGSP HbA1c standardisation and the master equation. NGSP percent = 0.09148 x IFCC (mmol/mol) + 2.152.
  4. Acute Care Testing. Measurement of circulating glucose: the problem of inconsistent sample and methodology. The whole-blood to plasma glucose factor of 1.11.
  5. Fitzgerald and colleagues. The 99th percentile upper reference limit for the fifth-generation cardiac troponin T assay in the United States. Clinica Chimica Acta, 2020 (overall 19.2 ng/L; women 13.5 to 13.6; men 21.4 to 22.2, across nine calculation methods).
  6. Greene and colleagues (IFCC Committee on Clinical Applications of Cardiac Bio-Markers). Establishing consensus-based, assay-specific 99th percentile upper reference limits to facilitate proper utilization of cardiac troponin measurements. Clinical Chemistry and Laboratory Medicine, 2017.
  7. Tate and colleagues. Opinion paper: deriving harmonised reference intervals, global activities. EJIFCC, 2016 (19 of 27 analytes acceptable for a common interval).
  8. Clinical Laboratory News. Point-of-care or clinical lab INR for anticoagulation monitoring: which to believe? On confirming a point-of-care INR over 5.0 against the laboratory.
  9. International Organization for Standardization. ISO 15189:2022, medical laboratories: requirements for quality and competence. Clause 7.3.7.4 on comparability of results of examinations.