The Inference Line
A Comprehensive Framework for Evaluating AI Tools in Healthcare
*This is the executive summary of a full white paper that outlines the framework and its application. You can download it in full here.
I. EXECUTIVE SUMMARY
A little while ago, I was asked to give a presentation at a graduate program on “AI in Healthcare.” The request seemed simple enough on its face. But preparing for it forced a harder question than I expected: what do we actually mean when we say “AI in Healthcare”?
The term gets used to describe an AI scribe that transcribes a clinical note, a computer-vision system that tracks a patient’s range of motion, and a platform that autonomously adjusts a diabetic patient’s insulin dose between physician visits. The term treats all of these tools as though they were interchangeable applications of a single technology, carrying the same risk, the same regulatory burden, and the same stakes if something goes wrong. They are not. This paper is the result of taking that question seriously.
The problem.
Vendors market “AI” as a single, undifferentiated capability. AI scribes are described using the same terminology as clinical decision support tools. Media coverage treats AI the same way. We often see articles and news stories describing “AI-powered Care”, without any real distinction of the part of care delivery being “AI-powered”. Even parts of the regulatory apparatus built to govern it — state medical and professional licensing boards chief among them — are still operating without a coherent framework for telling one kind of AI tool apart from another.
That conflation isn’t simply a semantic nuisance. It has real consequences for the four groups of people actually making decisions about these tools:
entrepreneurs who mis-scope how much regulatory and validation investment their product needs before they’ve built it
healthcare leaders who can’t meaningfully evaluate a vendor’s claims because the vendor’s own category label can’t be trusted
regulators who lack a shared vocabulary for distinguishing a routine administrative tool from an autonomous clinical actor
policymakers designing reimbursement policy blind to the fact that the two categories carry fundamentally different economics and risk profiles.
In addition, each of these 4 groups have differing incentives, value-drivers, and objectives when it comes to evaluating and categorizing AI-powered AI tools and solutions. Left to their own devices, each group could end up creating its own nomenclature or categorization system that does not align well with one or more of the other groups.
The framework.
This paper introduces the Inference Line — a framework for classifying any healthcare AI tool according to a single, portable test: does the tool’s output constitute a clinical inference?
Tools that don’t — scribes, scheduling automation, agentic routing, and descriptive computer-vision tracking — fall into Category 1: Care-Enablement. These tools may enable clinicians and providers to work more efficiently and effectively, but provide no input into the clinical decision-making or recommendations associated with that patient encounter. Tools that do output what would be considered a clinical inference belong to Category 2: Care-Guiding/Directing, which itself splits into two meaningfully different tiers: 2A — Decision Support, where the tool’s output informs a clinician’s judgment but a human makes the final call every time, and 2B — Autonomous Action, where the tool executes a clinical action within clinician-defined parameters without per-instance human sign-off.
Section III develops this in full, along with the Clinical Inference Test — a short, repeatable set of questions for applying the framework to any specific tool rather than relying on how a vendor happens to market it.
Why now.
The urgency here is not abstract or theoretical. Real world situations continue to arise that bring this problem to the forefront of discussions of healthcare technology, regulation, and implementation.
For example, in the same week of June 2026, two AI-enabled care tools revealed just how unresolved this territory is. UpDoc received Food and Drug Administration (FDA) clearance for what it describes as the first Software as a Medical Device built on patient-facing large language models — a platform now operating inside Cleveland Clinic, Allegheny Health Network, and UCSF Health that autonomously adjusts insulin dosing between physician visits [22][24].
That same week, Utah’s Medical Licensing Board was fighting to suspend Doctronic, a considerably more conservative, physician-supervised prescription-renewal pilot, after being — in the board’s own words — “blindsided” by its launch [20]. Two Category 2 tools met with two entirely different regulatory postures, playing out simultaneously.
That contrast, developed fully in Section VII, is the clearest illustration available right now of why this framework is needed. It is not an isolated case: FDA clearances for AI-enabled applications have grown from five in 2017 to more than seven hundred today [12], and more than forty states introduced over two hundred and forty AI-related health bills in 2026 alone [20]. The volume of activity has outpaced the shared vocabulary needed to govern and communicate it effectively across all four involved groups.
What follows.
This paper builds out the Inference Line in full — its relationship to existing regulatory precedent already embedded in the American Medical Association’s (AMA) Current Procedural Terminology (CPT) code set and FDA guidance (Section IV); a set of real-world boundary cases, including a cardiac imaging platform that disclaims clinical judgment while functionally exercising it (Section V); the evidence base supporting the framework’s practical stakes and application (Section VI); a deep exploration of the regulatory and licensure questions the UpDoc/Doctronic contrast raises, with direct implications for physical and occupational therapy (Section VII); and closing guidance tailored separately to entrepreneurs, healthcare leaders, regulators, and policymakers (Section VIII).
The goal throughout this report is not to resolve every open question the framework surfaces — several are named explicitly as open rather than papered over — but to give every reader making decisions about healthcare AI a sharper, shared set of terms for having that conversation.
Subscribe to get the full copy of the white paper, or go to this link.


