Skip to main navigation Skip to search Skip to main content

SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification

  • Zichen Zhang
  • , Ye Eun Bae
  • , Jonathan R. Bradley
  • , Lang Wu
  • , Chong Wu

Research output: Contribution to journalArticlepeer-review

Abstract

Genes with moderate to low expression heritability may explain a large proportion of complex trait etiology, but such genes cannot be sufficiently captured in conventional transcriptome-wide association studies (TWASs), partly due to the relatively small available reference datasets for developing expression genetic prediction models to capture the moderate to low genetically regulated components of gene expression. Here, we introduce a method, the Summary-level Unified Method for Modeling Integrated Transcriptome (SUMMIT), to improve the expression prediction model accuracy and the power of TWAS by using a large expression quantitative trait loci (eQTL) summary-level dataset. We apply SUMMIT to the eQTL summary-level data provided by the eQTLGen consortium. Through simulation studies and analyses of genome-wide association study summary statistics for 24 complex traits, we show that SUMMIT improves the accuracy of expression prediction in blood, successfully builds expression prediction models for genes with low expression heritability, and achieves higher statistical power than several benchmark methods. Finally, we conduct a case study of COVID-19 severity with SUMMIT and identify 11 likely causal genes associated with COVID-19 severity.

Original languageEnglish (US)
Article number6336
JournalNature communications
Volume13
Issue number1
DOIs
StatePublished - Dec 2022

ASJC Scopus subject areas

  • General Chemistry
  • General Biochemistry, Genetics and Molecular Biology
  • General
  • General Physics and Astronomy

MD Anderson CCSG core facilities

  • Biostatistics Resource Group

Fingerprint

Dive into the research topics of 'SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification'. Together they form a unique fingerprint.

Cite this