A Prediction Is Only Safe When The Team Can Understand It

A Prediction Is Only Safe When The Team Can Understand It

Predictive models are most useful when clinical teams can understand the predictions, challenge them and decide whether to act on them. We asked Avi, our machine learning engineer, what has to be true before a model's output reaches a clinician.

NEX Health Intelligence

A Model That Ranks Well Can Still Mislead

A strong evaluation score does not make a predictions always safe to act on.

Models in this field are usually assessed on how well they rank patients by risk. Models have been shown to be robust predictors of an individual patient's risk of hospital-onset infections (Myall et al., 2022). Ranking metrics such as AUC-ROC tell you the order of the list is sound. However, in practice, the accuracy for an individual patient can vary significantly.

Avi puts the risk plainly:

"If we generate a predicted probability of 90%, and we just show that to the clinician, they will literally read it as a 90% chance that this person is infected. That's not really the case all the time.”

No model has access to every factor influencing a patient's risk. Some information may not be captured in routinely collected data, while small differences in exposure, biology or clinical circumstances can have large effects that are inherently difficult to model reliably.

This is why predictions should support, rather than replace, clinical judgement. The role of the model is not simply to provide a risk score, but to help clinicians understand which known factors are driving that risk and what may be contributing to it. The clinical team can then interpret that information alongside everything else they know about the patient and make the final decision.

Every Prediction Should Arrive With Its Reasons

Clinicians consistently ask why a particular patient has been flagged, and that question deserves an answer from the system rather than from the person reading it.

Our pipelines are built to pass the supporting evidence across with the score, for example that a patient has been in prolonged contact with a patient known to be infected. The clinician can then agree, disagree or look further.

This matters because NEX supports review. It does not diagnose infection, confirm transmission or decide anything on its own. A prediction without its reasoning gives an IPC team pressure without a route to act, which is the opposite of what we are for.

When the Model Surprises Us, We Stop and Look

Occasionally a prediction doesn't make sense. It's rare, and when it happens the team investigates until it can explain how the model reached that output.

Most of the time the answer is an edge case: a patient with no visible signal who was infected anyway. The cause is usually something the hospital did that our data never recorded, such as a family member's visit or transmission involving staff movement between patients.

We engineer around the cases we can catch and we are explicit about the ones we can't.

"We have to put warnings into place that sometimes the model will make errors. It's kind of a trade-off that just comes with it.”

Saying so is itself a safety measure. A team told the model is fallible reads its output with appropriate caution and keeps clinical judgement where it belongs.

Clinical Knowledge Belongs Inside the Pipeline

Some safety questions are not engineering questions.

Antibiotic exposure affects a patient's susceptibility to acquiring an infection, and Avi is clear that he is not the person to judge which drug groups carry that risk. Chang, our Chief Medical Officer, and Paul, our IPC Lead and hospital partner at Phramongkutklao Hospital in Thailand, identified the relevant groups. Once those were modelled into the outbreak pipeline, predictive performance improved.

This is an association drawn from clinical expertise rather than a causal finding of ours, and it is stronger for having come from the people qualified to state it.

Most of the Safety Work Is Maintenance

When Avi first joined NEX, roughly 90% of his time went on building and 10% on keeping existing systems running. That split has now reversed.

About 80% of his work is data cleaning. Healthcare data is messy: values are missing, some are wrong, and some are simply unexpected and have to be understood before they can be handled. Modelling, the part he enjoys most, is the remaining 20%.

Weekly team discussions still generate improvements to the pipelines, and everyone is asked what we could do better. The quiet, repetitive work is where reliability actually comes from, and reliability is what an IPC team is trusting when it acts on a prediction at seven in the morning.

More resources

Continue exploring infection intelligence

Browse the latest evidence, product thinking, and field notes from NEX.

View all resources

View all resources

Infection Intelligence for Safer Hospitals

See how NEX helps detect risks earlier, investigate outbreaks faster, and prevent avoidable infections.

Infection Intelligence for Safer Hospitals

See how NEX helps detect risks earlier, investigate outbreaks faster, and prevent avoidable infections.

Infection Intelligence for Safer Hospitals

See how NEX helps detect risks earlier, investigate outbreaks faster, and prevent avoidable infections.