Characterization and validation of EHR computable phenotypes for Long COVID using patient-reported symptoms: Insights from the nationwide RECOVER Program
Castro, VM; Gainer, V; Wattanasin, N; et al., Journal of the American Medical Informatics Association, July 2026
View Publication on PubMedPublication Details
Abstract
Objective: Long COVID (LC) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients. In this study, we use summarized symptom reports by RECOVER-Adult cohort participants linked to electronic health record (EHR) data to characterize patients and train a computable phenotype algorithm of LC.
Materials and methods: The study included adult participants with linked Fast Health Interoperability Resource (FHIR)-sourced EHR data. We characterized EHR diagnoses, procedures, medications, lab tests, and vital sign features associated with LC. A computable phenotyping algorithm was trained and validated against patient-reported symptoms.
Main outcome and measures: We assessed model discrimination and calibration in a held-out test set. We describe important model features and evaluate model discrimination and calibration.
Results: The study included 1,501 RECOVER-Adult cohort participants with linked EHR data. 376 (25%) met criteria for highly symptomatic LC based on the RECOVER Long COVID Research Index (LCRI). EHR features associated with LC included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate. The algorithm identifying patients with highly symptomatic LC had an AUROC of 0.80 (95% confidence interval (CI) 0.74-0.85), and AUPRC of 0.58 (95% CI, 0.47-0.69).
Conclusion and relevance: These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported LC symptoms. The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.
Authors
Victor M Castro, Vivian Gainer, Nich Wattanasin, Andrew Cagan, Ana Holzbach, James Chan, Leora Horwitz, Rachel Kenney, Ivan Diaz, Hannah Mandel, Shannon Wuller, Mady Hornig, Lisa O'Brien, Andrew Wylam, James Doster, Richard A Moffitt, Emily Pfaff, Mark G Weiner, Sajjad Abedian, Michael Koropsak, Sairam Parthasarathy, Hanieh Razzaghi, Justin Manjourides, Elizabeth W Karlson, Shawn N Murphy
Keywords
Computable phenotypes; Digital health; EHR; Long COVID; Machine learning; PASC; Patient-reported symptoms