Healthcare has a complicated relationship with standards.
We need them. We resist them. We create them. We ignore them. We audit them. We turn them into policies, protocols, checklists, accreditation requirements, and laminated cards stuck beside computer monitors.
And clinicians often hear the word standardization as something vaguely threatening: someone who does not understand my patient, or perhaps does not understand my work, is going to tell me exactly how to do it.
But a conversation with Dr. Steve Spear on the Leading Quality podcast gave me a different way to think about standards: as hypotheses.
As I see it, a standard can be at least three different things.
It can be a rule: Do it this way.
It can be a benchmark: This is what good looks like.
Or it can be a hypothesis: Based on what we currently know, if we do these things under these conditions, we expect these outcomes.
That third interpretation changes almost everything. A hypothesis is not asking for obedience because authority has spoken. It is making a claim that can be tested.
And perhaps that is exactly how we should think about standard work in healthcare.
A team of scientists
The hypothesis framing makes me imagine a hospital, clinic, surgical ICU, management team, quality department, or any other collection of people in healthcare as a team of scientists engaged in a continuous series of experiments.
In some ways, that is already what we are. We observe, form theories, intervene, look at what happens, and revise our understanding.
But there is an important constraint. Once we have discovered something better, our next experiments should be conducted on top of that new knowledge.
We cannot perpetually allow ourselves to do whatever we like when we already have a current best hypothesis. I emphasize ourselves deliberately because standardization is often framed as something leaders impose upon frontline workers: we decide how they should work. But if a healthcare organization is genuinely functioning as a scientific community, the standard belongs to all of us.
Someone may initially design it from the top. It may emerge from frontline experience. More likely, it should be created through some combination of both. But once we agree that this is our best current understanding of how the work should be done, everyone accepts two responsibilities.
First, we should generally work from that shared standard rather than continuously reinventing the process independently.
Second, we should continuously test the standard against reality.
The default assumption should be that the standard is never final. That preserves rigor while making the standard intellectually humble.
The opposite of standard work
I remember working at a hospital where small cards labelled “standard work” appeared beside computers. One described how patient discharge rounds were supposed to be conducted. Leadership had decided on the process, printed the cards, and placed them around the hospital.
And that was essentially the end of the experiment.
There was no meaningful feedback about whether people were actually able to conduct rounds that way. There was no test of whether the proposed process made sense in the real conditions of the units. There was no systematic mechanism for discovering why (or even knowing if) people deviated from it. There was no feedback about whether the standard itself should change. And there was little connection between adherence to the process and the outcomes it was presumably intended to produce.
It was called standard work. But in many ways it was the antithesis of what standard work should be. If we describe something as a hypothesis, printing it on a card is not the end of the work. It is the beginning.
Who is accountable when the standard fails?
The hypothesis framing changes the relationship between leadership and the frontline. Most often in healthcare today, when a leader introduces a standard as a rule, it primarily serves to increase frontline accountability. The dominant question becomes:
Why didn’t you follow the standard?
But if the standard is a hypothesis, there are at least two equally important questions:
Did we do what we said we would do?
and
When we did, did it actually produce what we expected?
That second question changes the power dynamic. It makes the standard itself accountable. A clinician who repeatedly discovers that following a standard does not produce the predicted result is not necessarily being resistant. They may be generating evidence that the organization’s theory is wrong.
Standardization should increase leadership accountability, not merely frontline accountability.
If I tell hundreds of people that this is the best way to perform an important task, I should have an unusually strong interest in discovering evidence that I am wrong. Unfortunately, organizations can sometimes behave in exactly the opposite way.
Once a process has become a policy, protocol, approved workflow, or accreditation requirement, considerable institutional energy may go into defending compliance with it.
The hypothesis framing asks us instead to actively look for disconfirming evidence. That does not mean every clinician gets to ignore a standard whenever they disagree with it. Quite the opposite. It means that we must agree to disciplined adherence to our current hypotheses until evidence shows they should be revised. By doing this, deviation, failure, and unexpected outcomes become critical information.
Not all variation means the same thing
Imagine that a patient does not improve after a standardized process is followed.
There are several possibilities.
Perhaps the standard was never actually followed.
Perhaps it was followed, but the patient was meaningfully different from the population or circumstances for which the standard works.
Perhaps the environment made reliable execution impossible.
Or perhaps the standard itself is wrong.
Those are four very different learning opportunities. Yet traditional compliance systems can flatten all of them into a single category: variance.
That wastes information.
If we genuinely believe our standards are hypotheses, then every meaningful deviation or unexpected outcome becomes an opportunity to understand what kind of failure occurred. Did the process fail? Did the environment fail to support the process? Did the prediction fail? Or did we encounter a condition our current theory does not adequately explain?
That is much closer to science.
We already know how to do this with patients
There is a striking contradiction here.
Clinicians already think this way constantly. Suppose I prescribe a medication. I do not normally think: I prescribed the evidence-based medication, therefore the job is finished.
I have made an intervention based on a prediction. After that intervention I expect the blood pressure to fall, the pain to improve, the infection to respond, or the laboratory value to change. And then I follow up.
If reality disagrees with my expectation, I rethink the diagnosis, treatment, dose, adherence, physiology, or perhaps the entire theory of what is going on. That is normal clinical reasoning.
But our approach to organizational interventions can be considerably less disciplined. Sometimes it is not even: I prescribed the evidence-based medication, therefore the job is finished. It is closer to: I prescribed a medication that seems pretty good according to my intuition, therefore the job is finished.
We introduce a new committee, change the workflow, redesign rounds, create a policy, add a form or EHR field. We train everybody and then we move on.
We routinely treat clinical interventions as hypotheses, but organizational interventions as commandments.
This connects to another point Steve made in our conversation. Clinicians already know how to examine, diagnose, treat, and follow up. The missed opportunity is that we often fail to apply that same discipline “a step or two or three away from the bedside”—to the systems that shape the care our patients ultimately receive.
What would happen if we treated the system itself with the same clinical discipline we bring to the patient?
A standard should generate evidence about itself
That leads to another important question:
How do we design standards that generate evidence about themselves?
An effective standard should ideally help us answer whether the work happened as we expected and whether it produced the result we expected. Consider the medication analogy again.
A weak standard might say:
When condition Z is present, give medication X at dose Y.
A stronger standard would implicitly contain more:
When condition Z is present, give medication X at dose Y. Confirm that it was administered correctly. Look for response A within time B. If response A does not occur, reassess.
Now the standard contains not merely an action but a test of the theory.
Healthcare already contains examples of this. Barcode medication administration can detect some mismatches at the moment work is performed rather than discovering them later through audit. A surgical count reconciles what should be present with what is actually present. Teach-back gives us an immediate test of whether a patient actually understood what we intended to communicate. Closed-loop systems for diagnostic tests can detect when an expected acknowledgement or follow-up has not occurred.
Clinical pathways can specify both an intervention and the expected response, but there are enormous opportunities to go further. Admittedly, there are also many situations where this is extremely difficult to design.
What makes this hard in healthcare is that so much of the important work is not directly observable. We can’t routinely observe, and certainly not in real time, whether the clinician recognized that the patient’s condition was changing. We often learn this only by speaking with the clinician long after the fact, when memories may have faded and their recollection may be shaped by the circumstances in which it is elicited.
Similarly, we can’t see if a handoff communicated the most important uncertainty, the receiving clinician understood the contingency plan, or the nurse knew which change should trigger escalation.
Patients discharged from hospital may not know how to take their medications or under what circumstances to return to hospital, and our records alone won’t capture this.
In these scenarios, an EHR checkbox telling us that something was “done” may be a remarkably weak test of whether the underlying work actually happened.
So what might standards that generate evidence about themselves look like?
Perhaps a discharge process does not merely require that education be documented. It includes a lightweight method for confirming what the patient or caregiver actually understood and whether the next step occurred. Drs. Amy Billett and Chris Wong (Leading Quality Episode #8) provide strong examples of such education in their work in pediatric central-line care.
Perhaps a handoff tool should do more than record that a handoff occurred. In fact, use of I-PASS as a structured handoff tool already points in this direction and has been associated with substantial reductions in medical errors and preventable adverse events.1
Perhaps an escalation pathway can identify when an expected response did not happen within the anticipated time and offer help before the delay becomes harm.
Perhaps a new rounding process contains its own measures of whether the people involved were actually able to accomplish the intended work, rather than waiting six months for a retrospective audit.
Perhaps we could design digital systems that recognize recurring workarounds. If clinicians repeatedly bypass the same step, the first organizational question should not automatically be, How do we force compliance?
It might be:
What are these clinicians discovering about our standard that we don’t yet understand?
AI could change what we can observe
Artificial intelligence makes this increasingly interesting.
Computer vision is beginning to make some previously invisible clinical processes observable. Ambient systems can increasingly understand elements of conversation and workflow.2 AI video analysis is helping to recognize when patients are at risk of falls.3 EHR data can identify sequences, omissions, delays, and recurring patterns at scales that humans could never manually review.4
In principle, these technologies could help us know whether a standard was followed, where the work departed from expectation, and whether the expected result followed.
But there is a major danger here:
A learning system and a surveillance system can use exactly the same technology.
A camera, microphone, AI model, or event log can be used to help people succeed or it can be used to catch people doing something wrong. To deploy these technologies for the benefit of healthcare workers and patients alike, leaders will need to appreciate that their people really are one of the greatest resources available to the organization.
The goal should be to give them what they need to thrive, not to build increasingly sophisticated ways of constraining them.
A good system might notice that I have forgotten an important lab test required before the antibiotic I’m prescribing and give me a nudge. It might recognize that the conditions around me are making the standard difficult to follow and offer help. It might identify that a step is repeatedly failing across hundreds of clinicians and signal that the process itself needs redesign. It might make expertise available at exactly the moment it is needed.
That is very different from creating a system whose primary purpose is to accumulate evidence against the people doing the work.
But even if we can observe more, we still have to design these systems in a way that helps clinicians rather than burdening them.
There is another constraint.
If we want people to explain meaningful deviations from standards, the mechanism cannot itself make clinical work worse. During a resuscitation, for example, an AI system might appropriately flag an amiodarone dose that appears inconsistent with the expected sequence because the discrepancy could matter immediately. What would not make sense is interrupting the team to demand that a physician document, in real time, why they departed from a protocol. Likewise, requiring contemporaneous justification during an urgent surgical procedure could increase risk rather than reduce it. The standard should create accountability for meaningful deviation without turning every deviation into an interruption.
On balance, the design of these standards would not aim to increase documentation and would be mindful of the real-world value of that documentation. Instead it would provide a way to capture meaningful deviations from expected practice in a way that supports learning, quality of care, and the well-being of clinicians and patients. The design challenge is to create visibility without creating friction.
Standards don’t prevent experimentation. They make improvement possible.
There is another reason the hypothesis framing matters. Standards are sometimes portrayed as the opposite of creativity, autonomy, or experimentation.
I think the reverse is often true.
Without a standard, we may already have enormous amounts of experimentation. But it is experimentation in all directions, at all times, conducted independently by hundreds or thousands of people.
One clinician does it this way. Another does it slightly differently. A third developed a workaround years ago. A fourth learned another process during residency.
Nobody necessarily knows that these experiments are occurring. Their results aren’t analyzed and others never get to learn from those that succeed. That is not a learning system. It is uncontrolled experimentation without observation.
This is also occurring in a setting where creative energy itself is a limited resource. Since, under the proposed hypothesis-as-standard framework, we are still asking clinicians to deploy their creativity, we owe it to them to create conditions where they can do that only when it is most useful. Creativity to deploy endless workarounds is not creativity.
A standard gives us a current shared baseline. Now, when someone finds something better, there is something against which it can be compared. If it works, the standard can change and the next round of experimentation begins from a more advanced starting point.
In that sense, the standard is not what prevents experimentation.
The standard is what allows experimentation to accumulate into improvement.
Follow the best current standard when appropriate, make meaningful deviations visible, observe the results, and investigate anomalies. Revise the standard when reality tells us our hypothesis can be improved. Repeat.
Rigorous and provisional
I increasingly think this may be a much more useful way to talk about standardization with clinicians. Professional judgment matters more than ever and should be the fuel for our improvement.
We can acknowledge that many great ideas emerge from the frontline while leaders continue to exercise their responsibility to design systems and set institutional priorities.
Leaders can set standards. Frontline clinicians can create standards. Both can challenge and improve them.
But everyone, including the people with the most organizational authority, has to accept the same bargain:
This is our best current hypothesis. We will take it seriously enough to follow it, and we will remain humble enough to try to prove it wrong.
Perhaps the most scientific healthcare organizations will be those that understand standards not as fixed truths, but as our best current hypotheses. They may be the ones that hold their standards most rigorously and most provisionally.
And perhaps that is the real opportunity.
Not fewer standards. Better hypotheses.
Listen: My full conversation with Dr. Steve Spear → Spotify | Apple Podcasts | Other Platforms
A note on Leading Quality
This is the first edition of the Leading Quality newsletter. Every other week, I’ll explore ideas about how healthcare systems improve, why meaningful change is difficult, and what leaders can do to build organizations capable of learning. These essays will draw on research, my own experience, and conversations from the Leading Quality podcast.
References
Starmer AJ, Spector ND, Srivastava R, et al. Changes in medical errors after implementation of a handoff program. N Engl J Med. 2014;371(19):1803-1812. doi:10.1056/NEJMsa1405556.
Duggan MJ, Gervase J, Schoenbaum A, et al. Clinician experiences with ambient scribe technology to assist with documentation burden and efficiency. JAMA Netw Open. 2025;8(2):e2460637. doi:10.1001/jamanetworkopen.2024.60637.
Danial M, Chow CT, Lim MH, Ayop NA, Looi I, Ch’ng ASH. AI-based patient monitoring for fall prevention in stroke patients: a pilot study at a Malaysian acute stroke unit. J Neuroeng Rehabil. 2025;22:216. doi:10.1186/s12984-025-01706-9.
Marzano L, Darwich AS, Raghothama J, et al. Diagnosing an overcrowded emergency department from its Electronic Health Records. Sci Rep. 2024;14(1):9955. doi:10.1038/s41598-024-60888-9.



