Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Data-driven discovery of cancer driver genes, including tumor suppressor genes (TSGs) and oncogenes (OGs), is imperative for cancer prevention, diagnosis, and treatment. Although epigenetic alterations are important for tumor initiation and progression, most known driver genes were identified based on genetic alterations alone. Here, we developed an algorithm, DORGE (Discovery of Oncogenes and tumor suppressoR genes using Genetic and Epigenetic features), to identify TSGs and OGs by integrating comprehensive genetic and epigenetic data. DORGE identified histone modifications as strong predictors for TSGs, and it found missense mutations, super enhancers, and methylation differences as strong predictors for OGs. We extensively validated DORGE-predicted cancer driver genes using independent functional genomics data. We also found that DORGE-predicted dual-functional genes (both TSGs and OGs) are enriched at hubs in protein-protein interaction and drug-gene networks. Overall, our study has deepened the understanding of epigenetic mechanisms in tumorigenesis and revealed previously undetected cancer driver genes.
Pubmed ID: 33177077
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Database to store and display somatic mutation information and related details and contains information relating to human cancers. The mutation data and associated information is extracted from the primary literature. In order to provide a consistent view of the data a histology and tissue ontology has been created and all mutations are mapped to a single version of each gene. The data can be queried by tissue, histology or gene and displayed as a graph, as a table or exported in various formats.
Some key features of COSMIC are:
* Contains information on publications, samples and mutations. Includes samples which have been found to be negative for mutations during screening therefore enabling frequency data to be calculated for mutations in different genes in different cancer types.
* Samples entered include benign neoplasms and other benign proliferations, in situ and invasive tumours, recurrences, metastases and cancer cell lines.
Worldwide authority that approves standardized nomenclature to gene name and symbol, short form abbreviation, for each known human gene and stores all approved symbols in HGNC database. Approved human gene nomenclature. Database of gene symbols. Manually curated genes into family sets based on shared characteristics such as homology, function or phenotype. Data for protein-coding genes, pseudogenes, non-coding RNAs, phenotypes and genomic features.
View all literature mentionsSoftware platform for complex network analysis and visualization. Used for visualization of molecular interaction networks and biological pathways and integrating these networks with annotations, gene expression profiles and other state data.
View all literature mentionsSoftware package for interpreting gene expression data. Used for interpretation of a large-scale experiment by identifying pathways and processes.
View all literature mentionsEncyclopedia of DNA elements consisting of list of functional elements in human genome, including elements that act at protein and RNA levels, and regulatory elements that control cells and circumstances in which gene is active. Enables scientific and medical communities to interpret role of human genome in biology and disease. Provides identification of common cell types to facilitate integrative analysis and new experimental technologies based on high-throughput sequencing. Genome Browser containing ENCODE and Epigenomics Roadmap data. Data are available for entire human genome.
View all literature mentionsCurated protein-protein and genetic interaction repository of raw protein and genetic interactions from major model organism species, with data compiled through comprehensive curation efforts.
View all literature mentionsSoftware package for high-performance read alignment, quantification and mutation discovery.General purpose read aligner which can be used to map both genomic DNA-seq reads and RNA-seq reads. Subread aligner as fast, accurate and scalable read mapping by seed-and-vote.These programs were also implemented in Bioconductor R package Rsubread.
View all literature mentionsIntegrated database resource consisting of 16 main databases, broadly categorized into systems information, genomic information, and chemical information. In particular, gene catalogs in completely sequenced genomes are linked to higher-level systemic functions of cell, organism, and ecosystem. Analysis tools are also available. KEGG may be used as reference knowledge base for biological interpretation of large-scale datasets generated by sequencing and other high-throughput experimental technologies.
View all literature mentionsA web-based application designed with an easy-to-use interface to facilitate the high-throughput assessment and prioritization of genes and missense alterations important for cancer tumorigenesis.
View all literature mentionsA suite of tools to address common questions raised in genomic studies - mostly with regard to overlap and proximity relationships between data sets.
View all literature mentionsA read summarization program, which counts mapped reads for the genomic features such as genes and exons.
View all literature mentionsSoftware tool which predicts possible impact of amino acid substitution on structure and function of human protein using straightforward physical and comparative considerations. PolyPhen-2 is new development of PolyPhen tool for annotating coding nonsynonymous SNPs.
View all literature mentionsA unified data repository of the National Cancer Institute (NCI)'s Genomic Data Commons (GDC) that enables data sharing across cancer genomic studies in support of precision medicine. The GDC supports several cancer genome programs at the NCI Center for Cancer Genomics (CCG), including The Cancer Genome Atlas (TCGA), Therapeutically Applicable Research to Generate Effective Treatments (TARGET), and the Cancer Genome Characterization Initiative (CGCI). The GDC Data Portal provides a platform for efficiently querying and downloading high quality and complete data. The GDC also provides a GDC Data Transfer Tool and a GDC API for programmatic access.
View all literature mentionsConsortium to build comprehensive parts list of functional elements in human genome. This includes elements that act at protein and RNA levels, and regulatory elements that control cells and circumstances in which gene is active. Data from 2012-present.
View all literature mentionsProcedures for fitting the entire lasso or elastic-net regularization path for linear regression, logistic and multinomial regression models, Poisson regression and the Cox model. The algorithm uses cyclical coordinate descent in a path-wise fashion.
View all literature mentionsWeb tool to convert genome coordinates and genome annotation files between assemblies. Used to translate genomic coordinates from one assembly version into another and retrieves putative orthologous regions in other species using UCSC chained and netted alignments.
View all literature mentions