Research institution

Baylor College of Medicine

US

28 researchers0 verified9 linked papers127,583 indexed citations

Researchers

Public research profiles associated with Baylor College of Medicine.

Research from this institution

Publications linked through researcher authorship records.

2009 · Bioinformatics · 68,524 citations

The Sequence Alignment/Map format and SAMtools

SUMMARY: The Sequence Alignment/Map (SAM) format is a generic alignment format for storing read alignments against reference sequences, supporting short and long reads (up to 128 Mbp) produced by different sequencing platforms. It is flexible in style, compact in size, efficient in random access and is the format in which alignments from the 1000 Genomes Project are released. SAMtools implements various utilities for post-processing alignments in the SAM format, such as indexing, variant caller and alignment viewer, and thus provides universal tools for processing read alignments. AVAILABILITY: http://samtools.sourceforge.net.

2008 · Genome biology · 20,313 citations

Model-based Analysis of ChIP-Seq (MACS)

We present Model-based Analysis of ChIP-Seq data, MACS, which analyzes data generated by short read sequencers such as Solexa's Genome Analyzer. MACS empirically models the shift size of ChIP-Seq tags, and uses it to improve the spatial resolution of predicted binding sites. MACS also uses a dynamic Poisson distribution to effectively capture local biases in the genome, allowing for more robust predictions. MACS compares favorably to existing ChIP-Seq peak-finding algorithms, and is freely available.

2015 · Nature · 20,291 citations

A global reference for human genetic variation

The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations. Here we report completion of the project, having reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. We characterized a broad spectrum of genetic variation, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes >99% of SNP variants with a frequency of >1% for a variety of ancestries. We describe the distribution of genetic variation across the global sample, and discuss the implications for common disease studies. Results for the final phase of the 1000 Genomes Project are presented including whole-genome sequencing, targeted exome sequencing, and genotyping on high-density SNP arrays for 2,504 individuals across 26 populations, providing a global reference data set to support biomedical genetics. The 1000 Genomes Project has sought to comprehensively catalogue human genetic variation across populations, providing a valuable public genomic resource. The data obtained so far have found applications ranging from association studies and fine mapping studies to the filtering of likely neutral variants in rare-disease cohorts. The authors now report on the final phase of the project, phase 3, which covers previously uncharacterized areas of human genetic diversity in terms of the populations sampled and categories of characterized variation. The sample now includes more than 2,500 individuals from 26 global populations, with low coverage whole-genome and deep exome sequencing, as well as dense microarray genotyping. They find that while most common variants are shared across populations, rarer variants are often restricted to closely related populations. The authors also demonstrate the use of the phase 3 dataset as a reference panel for imputation to improve the resolution in genetic association studies.

2003 · Nature · 6,188 citations

The International HapMap Project

The goal of the International HapMap Project is to determine the common patterns of DNA sequence variation in the human genome and to make this information freely available in the public domain. An international consortium is developing a map of these patterns across the genome by determining the genotypes of one million or more sequence variants, their frequencies and the degree of association between them, in DNA samples from populations with ancestry from parts of Africa, Asia and Europe. The HapMap will allow the discovery of sequence variants that affect common disease, will facilitate development of diagnostic tools, and will enhance our ability to choose targets for therapeutic intervention.

2017 · Journal of the American College of Cardiology · 3,894 citations

Global, Regional, and National Burden of Cardiovascular Diseases for 10 Causes, 1990 to 2015

BACKGROUND: The burden of cardiovascular diseases (CVDs) remains unclear in many regions of the world. OBJECTIVES: The GBD (Global Burden of Disease) 2015 study integrated data on disease incidence, prevalence, and mortality to produce consistent, up-to-date estimates for cardiovascular burden. METHODS: CVD mortality was estimated from vital registration and verbal autopsy data. CVD prevalence was estimated using modeling software and data from health surveys, prospective cohorts, health system administrative data, and registries. Years lived with disability (YLD) were estimated by multiplying prevalence by disability weights. Years of life lost (YLL) were estimated by multiplying age-specific CVD deaths by a reference life expectancy. A sociodemographic index (SDI) was created for each location based on income per capita, educational attainment, and fertility. RESULTS: In 2015, there were an estimated 422.7 million cases of CVD (95% uncertainty interval: 415.53 to 427.87 million cases) and 17.92 million CVD deaths (95% uncertainty interval: 17.59 to 18.28 million CVD deaths). Declines in the age-standardized CVD death rate occurred between 1990 and 2015 in all high-income and some middle-income countries. Ischemic heart disease was the leading cause of CVD health lost globally, as well as in each world region, followed by stroke. As SDI increased beyond 0.25, the highest CVD mortality shifted from women to men. CVD mortality decreased sharply for both sexes in countries with an SDI >0.75. CONCLUSIONS: CVDs remain a major cause of health loss for all regions of the world. Sociodemographic change over the past 25 years has been associated with dramatic declines in CVD in regions with very high SDI, but only a gradual decrease or no change in most regions. Future updates of the GBD study can be used to guide policymakers who are focused on reducing the overall burden of noncommunicable disease and achieving specific global health targets for CVD.

2017 · JAMA Oncology · 2,037 citations

The Burden of Primary Liver Cancer and Underlying Etiologies From 1990 to 2015 at the Global, Regional, and National Level

Importance Liver cancer is among the leading causes of cancer deaths globally. The most common causes for liver cancer include hepatitis B virus (HBV) and hepatitis C virus (HCV) infection and alcohol use. Objective To report results of the Global Burden of Disease (GBD) 2015 study on primary liver cancer incidence, mortality, and disability-adjusted life-years (DALYs) for 195 countries or territories from 1990 to 2015, and present global, regional, and national estimates on the burden of liver cancer attributable to HBV, HCV, alcohol, and an “other” group that encompasses residual causes. Design, Settings, and Participants Mortality was estimated using vital registration and cancer registry data in an ensemble modeling approach. Single-cause mortality estimates were adjusted for all-cause mortality. Incidence was derived from mortality estimates and the mortality-to-incidence ratio. Through a systematic literature review, data on the proportions of liver cancer due to HBV, HCV, alcohol, and other causes were identified. Years of life lost were calculated by multiplying each death by a standard life expectancy. Prevalence was estimated using mortality-to-incidence ratio as surrogate for survival. Total prevalence was divided into 4 sequelae that were multiplied by disability weights to derive years lived with disability (YLDs). DALYs were the sum of years of life lost and YLDs. Main Outcomes and Measures Liver cancer mortality, incidence, YLDs, years of life lost, DALYs by etiology, age, sex, country, and year. Results There were 854 000 incident cases of liver cancer and 810 000 deaths globally in 2015, contributing to 20 578 000 DALYs. Cases of incident liver cancer increased by 75% between 1990 and 2015, of which 47% can be explained by changing population age structures, 35% by population growth, and −8% to changing age-specific incidence rates. The male-to-female ratio for age-standardized liver cancer mortality was 2.8. Globally, HBV accounted for 265 000 liver cancer deaths (33%), alcohol for 245 000 (30%), HCV for 167 000 (21%), and other causes for 133 000 (16%) deaths, with substantial variation between countries in the underlying etiologies. Conclusions and Relevance Liver cancer is among the leading causes of cancer deaths in many countries. Causes of liver cancer differ widely among populations. Our results show that most cases of liver cancer can be prevented through vaccination, antiviral treatment, safe blood transfusion and injection practices, as well as interventions to reduce excessive alcohol use. In line with the Sustainable Development Goals, the identification and elimination of risk factors for liver cancer will be required to achieve a sustained reduction in liver cancer burden. The GBD study can be used to guide these prevention efforts.

2017 · Nature · 1,607 citations

Integrated genomic and molecular characterization of cervical cancer

Cervical cancer remains one of the leading causes of cancer-related deaths worldwide. Here we report the extensive molecular characterization of 228 primary cervical cancers, one of the largest comprehensive genomic studies of cervical cancer to date. We observed notable APOBEC mutagenesis patterns and identified SHKBP1, ERBB3, CASP8, HLA-A and TGFBR2 as novel significantly mutated genes in cervical cancer. We also discovered amplifications in immune targets CD274 (also known as PD-L1) and PDCD1LG2 (also known as PD-L2), and the BCAR4 long non-coding RNA, which has been associated with response to lapatinib. Integration of human papilloma virus (HPV) was observed in all HPV18-related samples and 76% of HPV16-related samples, and was associated with structural aberrations and increased target-gene expression. We identified a unique set of endometrial-like cervical cancers, comprised predominantly of HPV-negative tumours with relatively high frequencies of KRAS, ARID1A and PTEN mutations. Integrative clustering of 178 samples identified keratin-low squamous, keratin-high squamous and adenocarcinoma-rich subgroups. These molecular analyses reveal new potential therapeutic targets for cervical cancers. This paper describes molecular subtypes of cervical cancers, including squamous cell carcinoma and adenocarcinoma clusters defined by HPV status and molecular features, and distinct molecular pathways that are activated in cervical carcinomas caused by different somatic alterations and HPV types. Cervical cancer is one of the main causes of cancer-related deaths worldwide, and 95% of cases result from human papilloma virus (HPV) infection. The Cancer Genome Atlas Research Network now reports the genomic and molecular characterization of 228 primary cervical cancers. The authors identify significantly mutated genes and pathways that differ by cervical cancer subtype, and find that keratin-low squamous, keratin-high squamous and adenocarcinoma-rich clusters are marked by different HPV types and molecular features.

2014 · Proceedings of the National Academy of Sciences · 122 citations

Evolutionary cell biology: Two origins, one objective

All aspects of biological diversification ultimately trace to evolutionary modifications at the cellular level. This central role of cells frames the basic questions as to how cells work and how cells come to be the way they are. Although these two lines of inquiry lie respectively within the traditional provenance of cell biology and evolutionary biology, a comprehensive synthesis of evolutionary and cell-biological thinking is lacking. We define evolutionary cell biology as the fusion of these two eponymous fields with the theoretical and quantitative branches of biochemistry, biophysics, and population genetics. The key goals are to develop a mechanistic understanding of general evolutionary processes, while specifically infusing cell biology with an evolutionary perspective. The full development of this interdisciplinary field has the potential to solve numerous problems in diverse areas of biology, including the degree to which selection, effectively neutral processes, historical contingencies, and/or constraints at the chemical and biophysical levels dictate patterns of variation for intracellular features. These problems can now be examined at both the within- and among-species levels, with single-cell methodologies even allowing quantification of variation within genotypes. Some results from this emerging field have already had a substantial impact on cell biology, and future findings will significantly influence applications in agriculture, medicine, environmental science, and synthetic biology.