Pablo Cingolani scite author profile

We describe a new computer program, SnpEff, for rapidly categorizing the effects of variants in genome sequences. Once a genome is sequenced, SnpEff annotates variants based on their genomic locations and predicts coding effects. Annotated genomic locations include intronic, untranslated region, upstream, downstream, splice site, or intergenic regions. Coding effects such as synonymous or non-synonymous amino acid replacement, start codon gains or losses, stop codon gains or losses, or frame shifts can be predicted. Here the use of SnpEff is illustrated by annotating ~356,660 candidate SNPs in ~117 Mb unique sequences, representing a substitution rate of ~1/305 nucleotides, between the Drosophila melanogaster w(1118); iso-2; iso-3 strain and the reference y(1); cn(1) bw(1) sp(1) strain. We show that ~15,842 SNPs are synonymous and ~4,467 SNPs are non-synonymous (N/S ~0.28). The remaining SNPs are in other categories, such as stop codon gains (38 SNPs), stop codon losses (8 SNPs), and start codon gains (297 SNPs) in the 5'UTR. We found, as expected, that the SNP frequency is proportional to the recombination frequency (i.e., highest in the middle of chromosome arms). We also found that start-gain or stop-lost SNPs in Drosophila melanogaster often result in additions of N-terminal or C-terminal amino acids that are conserved in other Drosophila species. It appears that the 5' and 3' UTRs are reservoirs for genetic variations that changes the termini of proteins during evolution of the Drosophila genus. As genome sequencing is becoming inexpensive and routine, SnpEff enables rapid analyses of whole-genome sequencing data to be performed by an individual laboratory.

show abstract

The genetic architecture of type 2 diabetes

Fuchsberger

Flannick

Teslovich

et al. 2016

Nature

936

907

View full text Add to dashboard Cite

The genetic architecture of common traits, including the number, frequency, and effect sizes of inherited variants that contribute to individual risk, has been long debated. Genome-wide association studies have identified scores of common variants associated with type 2 diabetes, but in aggregate, these explain only a fraction of heritability. To test the hypothesis that lower-frequency variants explain much of the remainder, the GoT2D and T2D-GENES consortia performed whole genome sequencing in 2,657 Europeans with and without diabetes, and exome sequencing in a total of 12,940 subjects from five ancestral groups. To increase statistical power, we expanded sample size via genotyping and imputation in a further 111,548 subjects. Variants associated with type 2 diabetes after sequencing were overwhelmingly common and most fell within regions previously identified by genome-wide association studies. Comprehensive enumeration of sequence variation is necessary to identify functional alleles that provide important clues to disease pathophysiology, but large-scale sequencing does not support a major role for lower-frequency variants in predisposition to type 2 diabetes.

show abstract

Using Drosophila melanogaster as a Model for Genotoxic Chemical Mutational Studies with a New Program, SnpSift

Cingolani¹,

Patel²,

Coon³

et al. 2012

Front. Gene.

788

659

View full text Add to dashboard Cite

This paper describes a new program SnpSift for filtering differential DNA sequence variants between two or more experimental genomes after genotoxic chemical exposure. Here, we illustrate how SnpSift can be used to identify candidate phenotype-relevant variants including single nucleotide polymorphisms, multiple nucleotide polymorphisms, insertions, and deletions (InDels) in mutant strains isolated from genome-wide chemical mutagenesis of Drosophila melanogaster. First, the genomes of two independently isolated mutant fly strains that are allelic for a novel recessive male-sterile locus generated by genotoxic chemical exposure were sequenced using the Illumina next-generation DNA sequencer to obtain 20- to 29-fold coverage of the euchromatic sequences. The sequencing reads were processed and variants were called using standard bioinformatic tools. Next, SnpEff was used to annotate all sequence variants and their potential mutational effects on associated genes. Then, SnpSift was used to filter and select differential variants that potentially disrupt a common gene in the two allelic mutant strains. The potential causative DNA lesions were partially validated by capillary sequencing of polymerase chain reaction-amplified DNA in the genetic interval as defined by meiotic mapping and deletions that remove defined regions of the chromosome. Of the five candidate genes located in the genetic interval, the Pka-like gene CG12069 was found to carry a separate pre-mature stop codon mutation in each of the two allelic mutants whereas the other four candidate genes within the interval have wild-type sequences. The Pka-like gene is therefore a strong candidate gene for the male-sterile locus. These results demonstrate that combining SnpEff and SnpSift can expedite the identification of candidate phenotype-causative mutations in chemically mutagenized Drosophila strains. This technique can also be used to characterize the variety of mutations generated by genotoxic chemicals.

show abstract

Identification and Functional Characterization of G6PC2 Coding Variants Influencing Glycemic Traits Define an Effector Transcript at the G6PC2-ABCB11 Locus

Mahajan¹,

Sim²,

Ng³

et al. 2015

PLoS Genet

100

120

View full text Add to dashboard Cite

Genome wide association studies (GWAS) for fasting glucose (FG) and insulin (FI) have identified common variant signals which explain 4.8% and 1.2% of trait variance, respectively. It is hypothesized that low-frequency and rare variants could contribute substantially to unexplained genetic variance. To test this, we analyzed exome-array data from up to 33,231 non-diabetic individuals of European ancestry. We found exome-wide significant (P<5×10-7) evidence for two loci not previously highlighted by common variant GWAS: GLP1R (p.Ala316Thr, minor allele frequency (MAF)=1.5%) influencing FG levels, and URB2 (p.Glu594Val, MAF = 0.1%) influencing FI levels. Coding variant associations can highlight potential effector genes at (non-coding) GWAS signals. At the G6PC2/ABCB11 locus, we identified multiple coding variants in G6PC2 (p.Val219Leu, p.His177Tyr, and p.Tyr207Ser) influencing FG levels, conditionally independent of each other and the non-coding GWAS signal. In vitro assays demonstrate that these associated coding alleles result in reduced protein abundance via proteasomal degradation, establishing G6PC2 as an effector gene at this locus. Reconciliation of single-variant associations and functional effects was only possible when haplotype phase was considered. In contrast to earlier reports suggesting that, paradoxically, glucose-raising alleles at this locus are protective against type 2 diabetes (T2D), the p.Val219Leu G6PC2 variant displayed a modest but directionally consistent association with T2D risk. Coding variant associations for glycemic traits in GWAS signals highlight PCSK1, RREB1, and ZHX3 as likely effector transcripts. These coding variant association signals do not have a major impact on the trait variance explained, but they do provide valuable biological insights.

show abstract

Intronic Non-CG DNA hydroxymethylation and alternative mRNA splicing in honey bees

et al. 2013

View full text Add to dashboard Cite

BackgroundPrevious whole-genome shotgun bisulfite sequencing experiments showed that DNA cytosine methylation in the honey bee (Apis mellifera) is almost exclusively at CG dinucleotides in exons. However, the most commonly used method, bisulfite sequencing, cannot distinguish 5-methylcytosine from 5-hydroxymethylcytosine, an oxidized form of 5-methylcytosine that is catalyzed by the TET family of dioxygenases. Furthermore, some analysis software programs under-represent non-CG DNA methylation and hydryoxymethylation for a variety of reasons. Therefore, we used an unbiased analysis of bisulfite sequencing data combined with molecular and bioinformatics approaches to distinguish 5-methylcytosine from 5-hydroxymethylcytosine. By doing this, we have performed the first whole genome analyses of DNA modifications at non-CG sites in honey bees and correlated the effects of these DNA modifications on gene expression and alternative mRNA splicing.ResultsWe confirmed, using unbiased analyses of whole-genome shotgun bisulfite sequencing (BS-seq) data, with both new data and published data, the previous finding that CG DNA methylation is enriched in exons in honey bees. However, we also found evidence that cytosine methylation and hydroxymethylation at non-CG sites is enriched in introns. Using antibodies against 5-hydroxmethylcytosine, we confirmed that DNA hydroxymethylation at non-CG sites is enriched in introns. Additionally, using a new technique, Pvu-seq (which employs the enzyme PvuRts1l to digest DNA at 5-hydroxymethylcytosine sites followed by next-generation DNA sequencing), we further confirmed that hydroxymethylation is enriched in introns at non-CG sites.ConclusionsCytosine hydroxymethylation at non-CG sites might have more functional significance than previously appreciated, and in honey bees these modifications might be related to the regulation of alternative mRNA splicing by defining the locations of the introns.

show abstract

Epigenetics of Early-Life Lead Exposure and Effects on Brain Development

et al. 2012

View full text Add to dashboard Cite

The epigenetic machinery plays a pivotal role in the control of many of the body's key cellular functions. It modulates an array of pliable mechanisms that are readily and durably modified by intracellular or extracellular factors. In the fast-moving field of neuroepigenetics, it is emerging that faulty epigenetic gene regulation can have dramatic consequences on the developing CNS that can last a lifetime and perhaps even affect future generations. Mounting evidence suggests that environmental factors can impact the developing brain through these epigenetic mechanisms and this report reviews and examines the epigenetic effects of one of the most common neurotoxic pollutants of our environment, which is believed to have no safe level of exposure during human development: lead.

show abstract

jFuzzyLogic: a Java Library to Design Fuzzy Logic Controllers According to the Standard for Fuzzy Control Programming

Cingolani¹,

Alcalá‐Fdez²

2013

IJCIS

164

View full text Add to dashboard Cite

Fuzzy Logic Controllers are a specific model of Fuzzy Rule Based Systems suitable for engineering applications for which classic control strategies do not achieve good results or for when it is too difficult to obtain a mathematical model. Recently, the International Electrotechnical Commission has published a standard for fuzzy control programming in part 7 of the IEC 61131 norm in order to offer a well defined common understanding of the basic means with which to integrate fuzzy control applications in control systems. In this paper, we introduce an open source Java library called jFuzzyLogic which offers a fully functional and complete implementation of a fuzzy inference system according to this standard, providing a programming interface and Eclipse plugin to easily write and test code for fuzzy control applications. A case study is given to illustrate the use of jFuzzyLogic.

show abstract

jFuzzyLogic: a robust and flexible Fuzzy-Logic inference system language implementation

Cingolani

Alcalá‐Fdez

2012

164

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Pablo Cingolani

A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff

The genetic architecture of type 2 diabetes

Using Drosophila melanogaster as a Model for Genotoxic Chemical Mutational Studies with a New Program, SnpSift

Identification and Functional Characterization of G6PC2 Coding Variants Influencing Glycemic Traits Define an Effector Transcript at the G6PC2-ABCB11 Locus

Intronic Non-CG DNA hydroxymethylation and alternative mRNA splicing in honey bees

Epigenetics of Early-Life Lead Exposure and Effects on Brain Development

jFuzzyLogic: a Java Library to Design Fuzzy Logic Controllers According to the Standard for Fuzzy Control Programming

jFuzzyLogic: a robust and flexible Fuzzy-Logic inference system language implementation

Contact Info

Product

Resources

About