Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
https://software.broadinstitute.org/gatk/
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on July 18th,2023. Software package for genome analysis. Used for analysis of next generation genomic data in cancer.
Proper citation: IndelGenotyper (RRID:SCR_016663) Copy
http://www.cbs.dtu.dk/biotools/sequenza/
Software package for copy number estimation from tumor genome sequencing data.Tools to analyze genomic sequencing data from paired normal-tumor samples, including cellularity and ploidy estimation; mutation and copy number (allele-specific and total copy number) detection, quantification and visualization.
Proper citation: Sequenza (RRID:SCR_016662) Copy
https://github.com/TGAC/RAMPART
Software for workflow management system for de novo genome assembly of DNA sequence data.Designed to exploit high performance computing environments, such as clusters and shared memory systems.
Proper citation: Rampart (RRID:SCR_016742) Copy
https://bioconductor.org/packages/release/bioc/html/Rsubread.html
Software R package for sequence alignment and counting for R. Used for analyses of second and third generation sequencing data, for read mapping, read counting, SNP calling, short and long read alignment, quantification and mutation discovery. Includes assessment of sequence reads, read alignment, read summarization, exon-exon junction detection, fusion detection, detection of short and long indels, absolute expression calling and SNP calling. Can be used with reads generated from any of the major sequencing platforms including Illumina GA/HiSeq/MiSeq, Roche GS-FLX, ABI SOLiD and LifeTech Ion PGM/Proton sequencers.
Proper citation: Rsubread (RRID:SCR_016945) Copy
Web tool to search multiple public variant databases simultaneously and provide a unified interface to facilitate the search process. Used for integration of human and model organism genetic resources to facilitate functional annotation of the human genome. Used for analysis of human genes and variants by cross-disciplinary integration of records available in public databases to facilitate clinical diagnosis and basic research.
Proper citation: MARRVEL (RRID:SCR_016871) Copy
https://support.10xgenomics.com/de-novo-assembly/software/overview/latest/welcome
Software to generate phased, whole genome de novo assemblies from a Chromium prepared library. Used to create true diploid de novo assemblies and can separate homologous chromosomes over long distances.
Proper citation: Supernova assembler (RRID:SCR_016756) Copy
https://software.broadinstitute.org/software/discovar/blog/
Software tool for variant calling with reference and de novo assembly of genomes. The heart of DISCOVAR is a de novo genome assembler which can generate de novo assemblies for both large and small genomes.
Proper citation: Discovar assembler (RRID:SCR_016755) Copy
The Horizontal Gene Transfer DataBase (HGT-DB) is a genomic database that includes statistical parameters such as G+C content, codon and amino-acid usage, as well as information about which genes deviate in these parameters for prokaryotic complete genomes. Under the hypothesis that genes from distantly related species have different nucleotide compositions, these deviated genes may have been acquired by horizontal gene transfer.
Proper citation: Horizontal Gene Transfer-DataBase (RRID:SCR_007706) Copy
GELBANK is a government project that provides an interactive interface for the comparison of 2DE patterns in the context of proteome sequence queries. Only proteomes of species with completed genomes (bacterial genomes, some eukaryotic genomes, human proteome) are presented in the database. The image database also contains not only scanned images, but also modeled gel patterns representing a collection of images (e.g. a master pattern for a sample). 2DE gel patterns are grouped by: tissue type, sample type, staining method used, separation technique used in the first dimension (by charge), the pH-range of the media used in first dimension, technique used in the second dimension (by size). Tools pertinent to the querying of two-dimensional gel-electrophoresis are implemented and integrated into database. When searching for sequences, tools that allow allow the discovery of sequences and alignment of multiple sequences are presented. Individual 2DE gel-patterns can be displayed or a collection of patterns can be animated.
Proper citation: GELBANK (RRID:SCR_007668) Copy
http://spock.genes.nig.ac.jp/~genome/gtop.html
GTOP is a database consists of data analyses of proteins identified by various genome projects. This database mainly uses sequence homology analyses and features extensive utilization of information on three-dimensional structures. GTOP is built by the Laboratory of Gene-Product Informatics at the National Institute of Genetics. This research is supported by the Japan Science and Technology Corporation and Grants-in-Aid for Scientific Research (Genomes in category C) from the Ministry of Education, Science, Sports and Culture of Japan. We use the following methods: Prediction of 3D structure Sequence homology search of PDB, using REVERSE PSI-BLAST. Functional predictions (family classifications) Sequence homology search of Swiss-Prot, a well-annotated sequence database, with the use of BLAST. Other analytical methods We are also carrying out the following analyses: Motif Analysis(PROSITE) Family classification(Pfam) Prediction of transmembrane helix domains(SOSUI) Prediction of coiled-coil regions(Multicoil) Repetitive sequence analysis(RepAlign)
Proper citation: GTOP - Genomes To Protein structures (RRID:SCR_007698) Copy
https://omictools.com/ecgene-tool
Database of functional annotation for alternatively spliced genes. It uses a gene-modeling algorithm that combines the genome-based expressed sequence tag (EST) clustering and graph-theoretic transcript assembly procedures. It contains genome, mRNA, and EST sequence data, as well as a genome browser application. Organisms included in the database are human, dog, chicken, fruit fly, mouse, rhesus, rat, worm, and zebrafish. Annotation is provided for the whole transcriptome, not just the alternatively spliced genes. Several viewers and applications are provided that are useful for the analysis of the transcript structure and gene expression. The summary viewer shows the gene summary and the essence of other annotation programs. The genome browser and the transcript viewer are available for comparing the gene structure of splice variants. Changes in the functional domains by alternative splicing can be seen at a glance in the transcript viewer. Two unique ways of analyzing gene expression is also provided. The SAGE tags deduced from the assembled transcripts are used to delineate quantitative expression patterns from SAGE libraries available publicly. The cDNA libraries of EST sequences in each cluster are used to infer qualitative expression patterns.
Proper citation: ECgene: Gene Modeling with Alternative Splicing (RRID:SCR_007634) Copy
A database of human, chimpanzee, mouse, and rat proteases and protease inhibitors, as well as as the growing number of hereditary diseases caused by mutations in protease genes. Analysis of the human and mouse genomes has allowed us to annotate 581 human, 580 chimpanzee, 667 mouse, and 655 rat protease genes. Proteases are classified in five different classes according to their mechanism of catalysis. Proteases are a diverse and important group of enzymes representing >2% of the human, chimpanzee, mouse and rat genomes. This group of enzymes is implicated in numerous physiological processes. The importance of proteases is illustrated by the existence of 99 different hereditary diseases due to mutations in protease genes. Furthermore, proteases have been implicated in multiple human pathologies, including vascular diseases, rheumatoid arthritis, neurodegenerative processes, and cancer. During the last ten years, our laboratory has identified and characterized more than 60 human protease genes. Due to the importance of proteolytic enzymes in human physiology and pathology, we have recently introduced the concept of Degradome, as the complete repertoire of proteases expressed by a tissue or organism. Thanks to the recent completion of the human, chimpanzee, mouse, and rat genome sequencing projects, we were able to analyze and compare for the first time the complete protease repertoire in those mammalian organisms, as well as the complement of protease inhibitor genes. This webpage also contains the Supplementary Material of Human and mouse proteases: a comparative genomic approach Nat Rev Genet (2003) 4: 544-558, Genome sequence of the brown Norway rat yields insights into mammalian evolution Nature (2004) 428: 493-521, A genomic analysis of rat proteases and protease inhibitors Genome Res. (2004) 14: 609-622, and Comparative genomic analysis of human and chimpanzee proteases Genomics (2005) 86: 638-647.
Proper citation: Mammalian Degradome Database (RRID:SCR_007624) Copy
Datasets and tools for comparative analysis and annotation of all publicly available genomes from three domains of life in a uniquely integrated context. Plasmids that are not part of a specific microbial genome sequencing project and phage genomes are also included in order to increase its genomic context for comparative analysis. The user interface (see User Interface Map) allows navigating the microbial genome data space along its three key dimensions (genes, genomes, and functions), and groups together the main comparative analysis tools. Microbial genome data analysis in IMG usually starts with the definition of an analysis context in terms of selected genomes, functional annotations, and/or genes, followed by the individual or comparative analysis of genomes, functional annotations, or genes.
Proper citation: IMG (RRID:SCR_007733) Copy
https://leger2.helmholtz-hzi.de/cgi-bin/expLeger.pl
Knowledge database and visualization tool for comparative genomics of pathogenic and non-pathogenic Listeria species.Provides information on gene functions (as annotated or supposed by literature from homologous organisms) , protein expression levels under defined experimental conditions ,subcellular localization of proteins (expected and/or experimentally validated) , biological meaning of genes and proteins based on KEGG, InterPro and Gene Ontology.
Proper citation: LEGER: the post-genome Database for Listeria Research (RRID:SCR_007760) Copy
http://phylomedb.bioinfo.cipf.es
Database for phylomes, that is, complete collections of phylogenetic trees for all proteins encoded in a given genome. It aims at providing a repository of high-quality phylogenies and alignments for proteins encoded in model species. To derive a phylome, each protein encoded in a given genome is used as a seed to retrieve its homologs in other complete genomes. These sequences are aligned and processed to derive reliable phylogenies using several phylogenetic methods. Besides providing the evolutionary history of the gene families, phylomeDB includes phylogeny based predictions of orthology and paralogy relationships., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: PhylomeDB (RRID:SCR_007850) Copy
http://papilio.ab.a.u-tokyo.ac.jp/genome/index.html
Silkbase''s objective is to build a foundation for the complete genome analysis of Bombyx mori.
Proper citation: Silkworm Genome Database (RRID:SCR_008242) Copy
http://cmbi.bjmu.edu.cn/cmbidata/cgf/CGF_Database/cytokine.medic.kumamoto-u.ac.jp/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on August 26, 2016. A collection of cDNA, gene and protein records of cytokines deposited in public databases provides various information about the cytokine members of vertebrates in other databases including NCBI GenBank, Swiss-Prot, UniGene, TIGR (The Institute for Genomic Research) Gene Indices, Ensembl, Entrez Gene, Mouse Genome Informatics (MGI) and Rat Genome Database (RGD). It also provides orthologous relationship of cytokine members and includes novel members identified in the databases.
Proper citation: Cytokine Family Database (RRID:SCR_008134) Copy
http://animal.dna.affrc.go.jp/agp/index.html
Database of comparative gene mapping between species to assist the mapping of the genes related to phenotypic traits in livestock. The linkage maps, cytogenetic maps, polymerase chain reaction primers of pig, cattle, mouse and human, and their references have been included in the database, and the correspondence among species have been stipulated in the database. AGP is an animal genome database developed on a Unix workstation and maintained by a relational database management system. It is a joint project of National Institute of Agrobiological Sciences (NIAS) and Institute of the Society for Techno-innovation of Agriculture, Forestry and Fisheries (STAFF-Institute), under cooperation with other related research institutes. AGP also contains the Pig Expression Data Explorer (PEDE), a database of porcine EST collections derived from full-length cDNA libraries and full-length sequences of the cDNA clones picked from the EST collection. The EST sequences have been clustered and assembled, and their similarity to sequences in RefSeq, and UniGene determined. The PEDE database system was constructed to store sequences and similarity data of swine full-length cDNA libraries and to make them available to users. It provides interfaces for keyword and ID searches of BLAST results and enables users to obtain sequence data and names of clones of interest. Putative SNPs in EST assemblies have been classified according to breed specificity and their effect on coding amino acids, and the assemblies are equipped with an SNP search interface. The database contains porcine nucleotide sequences and cDNA clones that are ready for analyses such as expression in mammalian cells, because of their high likelihood of containing full-length CDS. PEDE will be useful for researchers who want to explore genes that may be responsible for traits such as disease susceptibility. The database also offers information regarding major and minor porcine-specific antigens, which might be investigated in regard to the use of pigs as models in various medical research applications.
Proper citation: Animal Genome Database (RRID:SCR_008165) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented on August 20,2019.The COG-database has become a powerful tool in the field of comparative genomics. The construction of this data-base is based on sequence homologies of proteins from different completely sequenced genomes. Highly homologous proteins are assigned to clusters of orthologous groups. The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies. The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. Here is a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes.
Proper citation: Phylogenetic Clusters of Orthologous Groups Ranking (RRID:SCR_008223) Copy
http://www.nisc.nih.gov/projects/comp_seq.html
Generates data for use in developing and refining computational tools for comparing genomic sequence from multiple species. The NISC Comparative Sequencing Program's goal is to establish a data resource consisting of sequences for the same set of targeted genomic regions derived from multiple animal species. The broader program includes plans for a diverse set of analytical studies using the generated sequence and the publication of a series of papers describing the results of those analysis in peer-reviewed journals in a timely fashion. Experimentally, this project involves the shotgun sequencing of mapped BAC clones. For each BAC, an assembly is first performed when a sufficient number of sequence reads have been generated to provide full shotgun coverage of the clone. At that time, the assembled sequence is submitted to the HTGS division of GenBank. Subsequent refinements of the sequence, including the generation of higher-accuracy finished sequence, results in the updating of the sequence record in GenBank. By immediately submitting our BAC-derived sequences to GenBank, it makes their data available as a public service to allow colleagues to speed up their research, consistent with the now well-established routine of sequencing centers participating in the Human Genome Project. However, at the same time, it has made considerable investment in acquiring these mapping and sequence data, including sizable efforts of graduate students, postdoctoral fellows, and other trainees. Furthermore, in most cases, large data sets involving multiple BAC sequences from multiple species must first be generated, often taking many months to accumulate, before the planned analysis can be performed and the resulting papers written and submitted for publication.
Proper citation: Comparative Vertebrate Sequencing (RRID:SCR_008213) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.