Skip to main content

RECOVER’s electronic health record study revealed early answers about COVID and Long COVID

By studying the health records of millions of people across the United States, RECOVER researchers helped uncover initial discoveries about Long COVID and contributed findings that shaped the next phases of the initiative’s research.

In 2020 and 2021, millions of people were getting sick with COVID-19—and many of those people were experiencing symptoms for months or years without feeling better. At a time when Long COVID was still an emerging condition, little was known about what these symptoms meant for these people’s long-term health.

One powerful tool was already sitting in hospitals and clinics across the country: electronic health records (EHRs). These digital medical records—which contained diagnoses, lab results, medications, and doctor's visit notes—offered a way to study COVID-19's long-term effects at a scale that would have been impossible through traditional research methods alone.

RECOVER's EHR study drew on this resource to help answer some of the most pressing questions about Long COVID. By analyzing data from over 60 million patients across 2 national health research networks, RECOVER's EHR research program has now produced more than 70 peer-reviewed scientific papers, including innovations in applying machine learning to recognize Long COVID.

As the EHR study completes its analyses and closes out in 2026, its findings leave a lasting mark on how researchers, healthcare providers, and the Long COVID community can understand the long-term effects of COVID. 

Developing a way to consistently identify Long COVID

Researchers used data from EHRs to develop computable phenotypes for Long COVID. A computable phenotype is a computer program that scans health records and identifies patients who are likely to have Long COVID based on patterns in their diagnoses and other health information. Researchers published RECOVER’s first computable phenotype paper in 2022—a critical step at a time when Long COVID was still new and there was no formal definition for the condition. 

Using this approach, researchers were able to estimate that roughly 1 in 10 adults or 1 in 20 children who had COVID-19 may have developed Long COVID. They also found that the medical diagnosis code created specifically for Long COVID (called U09.9) was not being used consistently, which meant that millions of people with the condition were not being counted by traditional tracking methods.

The computable phenotypes were also used in many other RECOVER studies to help identify groups of study participants who had Long COVID. Usage of the computable phenotypes allowed research studies to be started more efficiently and to include more patients because it did not require a doctor to manually review each patient record and confirm the presence of Long COVID.

Researchers have continued work on computable phenotypes and have published many papers on the topic, including:

Uncovering additional Long COVID symptoms

Electronic health records capture 2 kinds of data: structured and unstructured. Traditional data tools are built to look at structured data, such as diagnoses and lab results, but doctors’ observations are captured as unstructured data that these traditional tools can’t read. 

RECOVER researchers used a computational process called natural language processing (NLP) to look deeper into EHR data. NLP allowed them to examine these unstructured data—called clinical notes—in addition to the structured fields that were part of the standard medical record. RECOVER EHR researchers published NLP models to reveal additional symptoms from both adults’ records and children’s records in 2025. 

The ability to examine unstructured clinical notes was a significant step in better understanding the different ways that people experience Long COVID. Researchers were able to capture symptoms and related conditions that may not have been fully recorded in diagnosis codes, allowing them to identify people with Long COVID more quickly and accurately. This increased accuracy meant that more people with Long COVID could be included in research studies, helping researchers build a more complete understanding of the many ways Long COVID can be experienced.

Increasing knowledge about pediatric Long COVID

A major focus of RECOVER's EHR work was understanding how Long COVID affects children and teenagers—a population that had received less research attention than adults in the early stages of the COVID pandemic.

Using data from PEDSnet's nationwide network of children's hospitals, RECOVER researchers produced some of the first evidence about pediatric Long COVID. A 2022 study published in JAMA Pediatrics found that children who had COVID-19 were more likely than children who had not had COVID-19 to develop a range of conditions in the months following their illness, with myocarditis (inflammation of the heart muscle) being one of the most commonly identified.

Additional EHR studies documented the effects of COVID-19 on children's heartskidneys, and digestive systems; explored how Long COVID presented differently across racial and ethnic groups; and found that children who had COVID-19 more than once during the Omicron period faced a higher risk of developing Long COVID than children who had COVID-19 only once.

Researchers also found that vaccination helped protect children from Long COVID. Two different study methods found that the COVID-19 vaccine significantly reduced the risk of Long COVID in children and adolescents. This protection worked primarily by reducing the risk of getting sick with COVID-19 in the first place, rather than through a direct effect on Long COVID development independent of initial infection.

Searching for possible treatment targets

The EHR study also employed an innovative study type called a target trial emulation (TTE). TTEs are like traditional clinical trials in that they examine how medications may impact a person’s health. However, unlike clinical trials, which randomly assign medications to participants, TTEs analyze the EHR to study the effects of medications a person is already taking.

One recent RECOVER TTE studied whether vaccination helped protect children and adolescents who had already had COVID from becoming infected with SARS-CoV-2 (the virus that causes COVID) again. Another RECOVER TTE found that people who took the antiviral drug Paxlovid while sick with COVID may be less likely to develop certain Long COVID symptoms; however, taking Paxlovid did not protect these people from developing Long COVID

While their findings cannot replace those of clinical trials, TTEs require fewer resources and less time to design and study a treatment. This approach could offer a faster way to help identify and evaluate possible treatments to test in full clinical trials.

Awaiting more publications

Studying Long COVID using electronic health records meant building the tools to identify patients who had the condition before anyone was certain how to define it, and then enhancing those tools as research and understanding evolved.

The RECOVER EHR study demonstrates that it is possible to mobilize large-scale health data quickly in response to an emerging health condition. The infrastructure built for this study—the data tools, the analysis methods, and an understanding of which questions EHR data is well-suited to answer—is a resource that researchers can continue to draw on even after the study ends. 

The RECOVER Report, RECOVER's monthly e-newsletter, will continue to share findings from the EHR study as final findings are published. To stay up to date, explore EHR publications or subscribe to our newsletter.

This story was first announced in the RECOVER Report, RECOVER’s monthly email newsletter. Complete this form to subscribe and receive the latest updates from RECOVER.