Research Infrastructure

Data Sources

Chahal Lab integrates large-scale genomic, clinical, and administrative datasets to uncover the genetic architecture of cardiovascular disease and translate discoveries into precision medicine.

Our research draws on complementary data modalities — from whole-genome sequencing in population biobanks to longitudinal electronic health records, wearable device streams, and national administrative claims. Combining these sources enables multi-ancestry GWAS, phenome-wide association studies, and real-world outcomes research at a scale not possible from any single dataset.

Primary Resources

Core Data Sources

The foundational datasets that underpin the majority of Chahal Lab research programs.

Electronic Health Records (EHR)

Longitudinal patient records from health systems including WellSpan Health, Geisinger, and partner institutions. EHR data encompass diagnoses, procedures, medications, laboratory results, and clinical notes — standardized using the Observational Medical Outcomes Partnership (OMOP) Common Data Model to enable cross-institutional analyses.

Genomics & Biosamples

Whole-genome sequencing, genome-wide genotyping arrays, and polygenic risk score analyses drawn from large biobank cohorts. DNA extracted from blood and saliva samples undergoes short-read and long-read whole-genome sequencing to identify single nucleotide polymorphisms (SNPs), insertions/deletions, and structural variants relevant to inherited cardiovascular conditions.

Wearable Devices & Digital Health (coming soon)

Continuous biometric data from consumer wearable devices — including heart rate, physical activity, sleep architecture, and arrhythmia detection — contribute a real-world phenotyping layer that complements traditional clinical measurements. These data are particularly valuable for studying ambulatory cardiac rhythm and lifestyle determinants of cardiovascular risk.

Large Population Biobanks

Multi-ancestry population cohorts with linked genomic and phenotypic data at scale. Chahal Lab leverages several of the world's largest biobanks to perform genome-wide association studies (GWAS), Mendelian randomization analyses, and polygenic risk score validation across diverse ancestries.

Datasets

  • UK Biobank (500,000+ participants)
  • NIH All of Us Research Program (1M+ goal)
  • Our Future Health (5M+ goal, UK)

Supplementary Resources

Administrative & Imaging Data

Population-level claims databases and deep cardiac phenotyping resources that complement genomic analyses.

Administrative Claims & Hospital Data (HCUP)

The Healthcare Cost and Utilization Project (HCUP), sponsored by the Agency for Healthcare Research and Quality (AHRQ), provides the largest collection of longitudinal hospital care data in the United States. Chahal Lab uses HCUP databases to study population-level trends in cardiovascular hospitalizations, procedures, readmissions, and outcomes.

Datasets

  • National Inpatient Sample (NIS)
  • Kids' Inpatient Database (KID)
  • Nationwide Ambulatory Surgery Sample (NASS)
  • Nationwide Emergency Department Sample (NEDS)
  • Nationwide Readmissions Database (NRD)

Cardiac Imaging & Phenotyping

Echocardiography, cardiac MRI, and electrocardiographic data provide deep structural and functional cardiac phenotypes. Automated image analysis pipelines and machine learning models extract quantitative measurements — including left ventricular dimensions, ejection fraction, and wall motion — that are linked to genomic and clinical data for phenome-wide association studies.

Chahal Lab®

Advancing knowledge for a better world through cardiovascular discovery, clinical insight, and community partnership.

Institutional contact

WellSpan Health
Department of Cardiology
York, Pennsylvania

© 2026 Chahal Lab. All rights reserved.

Academic research · Clinical collaboration · Community impact