Skip to Content
SBMI Horizontal Logo

Probabilistic Machine Learning Algorithm for Identifying Unwarranted Clinical Variation from Contextual Factors Derived from Electronic Health Records Encounter Data

Apollo McOwiti, MS (2026)

Primary advisor: Susan Fenton, PhD

Committee members: Laila R. Bekhet, PhD and Hongfang Liu, PhD

PhD thesis, McWilliams School of Biomedical Informatics at UTHealth Houston.

ABSTRACT

Unwarranted clinical variation (UCV) refers to care that is not aligned with a patient’s clinical characteristics, needs, or preferences. UCV frequently arises from local context factors (LCFs) and is associated with increased costs, unnecessary interventions, and departures from evidence-based practices. UCV represents a risk posed by healthcare organizations and clinicians, yet this risk often remains unrecognized by both providers and patients.

Detecting UCV is challenging due to the complexity of care decisions and the nuanced interpretation of variation. The state-of-the-art methods rely on centralized data aggregation and mixed-effects statistical models, which estimate relative variation across regions or sites but cannot detect absolute variation. Centralized analysis also presents two major drawbacks: first, UCV data are sensitive, and organizations may be reluctant to share proprietary clinical and operational information; second, when data access is revoked, UCV centralized reporting is disrupted. While site-level independent analyses are possible, they lack interoperability and comparability across organizations because LCFs are neither standardized nor harmonized. Moreover, existing approaches can signal UCV at geographic or site levels but fail to identify the underlying factors driving the variation. Machine learning (ML) methods leveraging LCFs for UCV detection are also lacking. This study demonstrates the feasibility of using ML to identify absolute UCV from LCFs extracted from electronic health records (EHR) across multiple sites. To enable interoperability of independent analyses, the study introduces an ontology-based standardization of LCFs for UCV analysis.

For LCFs standardization, an application ontology, UCVA, was developed using Protégé 5. UCVA incorporates semantically aligned concepts from external ontologies and introduces new concepts de novo. The study’s use case involved identifying inappropriate antibiotic prescriptions for pediatric acute viral pharyngitis. A retrospective analysis of ambulatory visits (ICD-10 J02.8) across multiple clinics of an academic health system was performed. Ensemble ML models - Random Forest, CatBoost, and Explainable Boosting Machine (EBM) - were trained on encounter-level EHR data. Model performance was assessed using nested cross-validation and AUC metrics.

UCVA comprises 78 classes, 102 instances, and 20 object properties, with a total of 944 axioms: 227 declaration axioms, 359 logical axioms, and 358 assertions. Textual definitions for classes were annotated using 'IAO:0000115' from the Information Artifact Ontology. All three ensemble models demonstrated strong performance (median AUC = 0.91). CatBoost models trained on weak labels performed comparably to those trained on gold-standard labels. Feature importance analysis identified site-level and provider-level case volumes as the most influential predictors, followed by provider credentials, experience, and encounter type. Lower provider case volumes were associated with a reduced likelihood of inappropriate treatment.

Using standardized datasets, machine learning models accurately detected absolute UCV from LCF data derived from EHR encounters. Explainable models, such as EBM, provide interpretability essential for clinical adoption. Findings support ML-based methods as scalable alternatives to traditional statistical approaches for UCV detection.