Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://www.jcvi.org/cgi-bin/tigrfams/index.cgi
Consists curated multiple sequence alignments, Hidden Markov Models (HMMs) for protein sequence classification, and associated information designed to support automated annotation of (mostly prokaryotic) proteins. Starting with release 10.0, TIGRFAMs models use HMMER3, which provides excellent search speed as well as exquisite search sensitivity. See the "TIGRFAMs Complete Listing" page to review the accession, protein name, model type, and EC number (if assigned) of all models. TIGRFAMs is a member database in InterPro. The HMM libraries and supporting files are available to download and use for free from our FTP site.
Proper citation: TIGRFAMS (RRID:SCR_005493) Copy
http://bowtie-bio.sourceforge.net/index.shtml
Software ultrafast memory efficient tool for aligning sequencing reads. Bowtie is short read aligner.
Proper citation: Bowtie (RRID:SCR_005476) Copy
An online toolbox and workflow management system for a broad range of bioinformatic and systems biology applications. The individual modules, or Bricks, are unified under a standardized interface, with a consistent look-and-feel and can flexibly be put together to comprehensive workflows. The workflow management is intuitively handled through a simple drag-and-drop system. With this system, you can edit the predefined workflows or compose your own workflows from scratch. Your own Bricks can easily be added as scripts or plug-ins and can be used in combination with pre-existing analyses. GeneXplain GmbH provides a number of state-of-the-art bricks; some of them can be obtained free of charge, while others require licensing for small fee in order to guarantee active maintenance and dynamic adaptation to the rapidly developing know-how in this field.
Proper citation: geneXplain (RRID:SCR_005573) Copy
http://mesquiteproject.org/packages/chromaseq/
A software package in Mesquite that processes chromatograms, makes contigs, base calls, etc., using in part the programs Phred and Phrap.
Proper citation: Chromaseq (RRID:SCR_005587) Copy
http://www.ebi.ac.uk/Tools/pfa/iprscan/
Software package for functional analysis of sequences by classifying them into families and predicting presence of domains and sites. Scans sequences against InterPro's signatures. Characterizes nucleotide or protein function by matching it with models from several different databases. Used in large scale analysis of whole proteomes, genomes and metagenomes. Available as Web based version and standalone Perl version and SOAP Web Service.
Proper citation: InterProScan (RRID:SCR_005829) Copy
Ratings or validation data are available for this resource
Portal to interactively visualize genomic data. Provides reference sequences and working draft assemblies for collection of genomes and access to ENCODE and Neanderthal projects. Includes collection of vertebrate and model organism assemblies and annotations, along with suite of tools for viewing, analyzing and downloading data.
Proper citation: UCSC Genome Browser (RRID:SCR_005780) Copy
Bioinformatics Resource Center for invertebrate vectors. Provides web-based resources to scientific community conducting basic and applied research on organisms considered potential agents of biowarfare or bioterrorism or causing emerging or re-emerging diseases.
Proper citation: VectorBase (RRID:SCR_005917) Copy
http://newt-omics.mpi-bn.mpg.de/index.php
Newt-omics is a database, which enables researchers to locate, retrieve and store data sets dedicated to the molecular characterization of newts. Newt-omics is a transcript-centered database, based on an Expressed Sequence Tag (EST) data set from the newt, covering ~50,000 Sanger sequenced transcripts and a set of high-density microarray data, generated from regenerating hearts. Newt-omics also contains a large set of peptides identified by mass spectrometry, which was used to validate 13,810 ESTs as true protein coding. Newt-omics is open to implement additional high-throughput data sets without changing the database structure. Via a user-friendly interface Newt-omics allows access to a huge set of molecular data without the need for prior bioinformatical expertise. The newt Notopthalmus viridescens is the master of regeneration. This organism is known for more than 200 years for its exceptional regenerative capabilities. Newts can completely replace lost appendages like limb and tail, lens and retina and parts of the central nervous system. Moreover, after cardiac injury newts can rebuild the functional myocardium with no scar formation. To date only very limited information from public databases is available. Newt-Omics aims to provide a comprehensive platform of expressed genes during tissue regeneration, including extensive annotations, expression data and experimentally verified peptide sequences with yet no homology to other publicly available gene sequences. The goal is to obtain a detailed understanding of the molecular processes underlying tissue regeneration in the newt, that may lead to the development of approaches, efficiently stimulating regenerative pathways in mammalians. * Number of contigs: 26594 * Number of est in contigs: 48537 * Number of transcripts with verified peptide: 5291 * Number of peptides: 15169
Proper citation: Newtomics (RRID:SCR_006073) Copy
http://www.nematodes.org/nembase4/
NEMBASE is a comprehensive Nematode Transcriptome Database including 63 nematode species, over 600,000 ESTs and over 250,000 proteins. Nematode parasites are of major importance in human health and agriculture, and free-living species deliver essential ecosystem services. The genomics revolution has resulted in the production of many datasets of expressed sequence tags (ESTs) from a phylogenetically wide range of nematode species, but these are not easily compared. NEMBASE4 presents a single portal into extensively functionally annotated, EST-derived transcriptomes from over 60 species of nematodes, including plant and animal parasites and free-living taxa. Using the PartiGene suite of tools, we have assembled the publicly available ESTs for each species into a high-quality set of putative transcripts. These transcripts have been translated to produce a protein sequence resource and each is annotated with functional information derived from comparison with well-studied nematode species such as Caenorhabditis elegans and other non-nematode resources. By cross-comparing the sequences within NEMBASE4, we have also generated a protein family assignment for each translation. The data are presented in an openly accessible, interactive database. An example of the utility of NEMBASE4 is that it can examine the uniqueness of the transcriptomes of major clades of parasitic nematodes, identifying lineage-restricted genes that may underpin particular parasitic phenotypes, possible viral pathogens of nematodes, and nematode-unique protein families that may be developed as drug targets.
Proper citation: NEMBASE (RRID:SCR_006070) Copy
http://hfv.lanl.gov/content/index
The Hemorrhagic Fever Viruses (HFV) sequence database collects and stores sequence data and provides a user-friendly search interface and a large number of sequence analysis tools, following the model of the highly regarded and widely used Los Alamos HIV database. The database uses an algorithm that aligns each sequence to a species-wide reference sequence. The NCBI RefSeq database is used for this; if a reference sequence is not available, a Blast search finds the best candidate. Using this method, sequences in each genus can be retrieved pre-aligned. Hemorrhagic fever viruses (HFVs) are a diverse set of over 80 viral species, found in 10 different genera comprising five different families: arena-, bunya-, flavi-, filo- and togaviridae. All these viruses are highly variable and evolve rapidly, making them elusive targets for the immune system and for vaccine and drug design. About 55,000 HFV sequences exist in the public domain today. A central website that provides annotated sequences and analysis tools will be helpful to HFV researchers worldwide.
Proper citation: HFV Database (RRID:SCR_006017) Copy
http://athina.biol.uoa.gr/DAM-Bio/
An integrated environment designed to support protein sequence and structure analysis on the web.
Proper citation: DAM-Bio (RRID:SCR_006226) Copy
http://aias.biol.uoa.gr/OMPdb/
A database of Beta-barrel outer membrane proteins from Gram-negative bacteria. The web interface of OMPdb offers the user the ability not only to view the available data, but also to submit advanced queries for text search within the database''s protein entries or run BLAST searches against the database. The most up-to-date version of the database (as well as all past versions) can be downloaded in various formats (flat text, XML format or raw FASTA sequences). For constructing OMPdb, multiple freely accessible resources were combined and a detailed literature search was performed. The classification of OMPdb''s protein entries into families is based mainly on structural and functional criteria. Information included in the database consists of sequence data, as well as annotation for structural characteristics (such as the transmembrane segments), literature references and links to other public databases, features that are unique worldwide. Along with the database, a collection of profile Hidden Markov Models that were shown to be characteristic for Beta-barrel outer membrane proteins was also compiled. This set, when used in combination with our previously developed algorithms (PRED-TMBB, MCMBB and ConBBPRED) will serve as a powerful tool in matters of discrimination and classification of novel Beta-barrel proteins and whole-genome analyses., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: OMPdb (RRID:SCR_006221) Copy
http://compbio.dfci.harvard.edu/tgi/
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 10, 2017. A pilot effort that has developed a centralized, web-based biospecimen locator that presents biospecimens collected and stored at participating Arizona hospitals and biospecimen banks, which are available for acquisition and use by researchers. Researchers may use this site to browse, search and request biospecimens to use in qualified studies. The development of the ABL was guided by the Arizona Biospecimen Consortium (ABC), a consortium of hospitals and medical centers in the Phoenix area, and is now being piloted by this Consortium under the direction of ABRC. You may browse by type (cells, fluid, molecular, tissue) or disease. Common data elements decided by the ABC Standards Committee, based on data elements on the National Cancer Institute''s (NCI''s) Common Biorepository Model (CBM), are displayed. These describe the minimum set of data elements that the NCI determined were most important for a researcher to see about a biospecimen. The ABL currently does not display information on whether or not clinical data is available to accompany the biospecimens. However, a requester has the ability to solicit clinical data in the request. Once a request is approved, the biospecimen provider will contact the requester to discuss the request (and the requester''s questions) before finalizing the invoice and shipment. The ABL is available to the public to browse. In order to request biospecimens from the ABL, the researcher will be required to submit the requested required information. Upon submission of the information, shipment of the requested biospecimen(s) will be dependent on the scientific and institutional review approval. Account required. Registration is open to everyone.. Documented on August 19,2019.The goal of The Gene Index Project is to use the available Expressed Sequence Transcript (EST) and gene sequences, along with the reference genomes wherever available, to provide an inventory of likely genes and their variants and to annotate these with information regarding the functional roles played by these genes and their products. The promise of genome projects has been a complete catalog of genes in a wide range of organisms. While genome projects have been successful in providing reference genome sequences, the problem of finding genes and their variants in genomic sequence remains an ongoing challenge. TGI has created an inventory that contains genes and their variants together with description. In addition, this resource is attempting to use these catalogs to find links between genes and pathways in different species and to provide lists of features within completed genomes that can aid in the understanding of how gene expression is regulated. DATABASES *Eukaryotic Gene Orthologues (formerly known as TOGA - TIGR Orthologous Gene Alignment): Eukaryotic Gene Orthologues (EGO) at DFGI are generated by pair-wise comparison between the Tentative Consensus (TC) sequences that comprise the Dana Farber Gene Indices from individual organisms. The reciprocal pairs of the best match were clustered into individual groups and multiple sequence alignments were displayed for each group. *GeneChip Oncology Database (GCOD):Cancer gene expression database is a collection of publicly available microarray expression data on Affymetrix GeneChip Arrays related to human cancers. Currently only datasets with available raw data (Affymetrix .CEL files) are processed. All processed datasets were subjected to extensive manual curation, uniform processing and consistent quality control. You can browse the experiments in our collection, perform statistical analysis, and download processed data; or to search gene expression profiles using Entrez gene symbol, Unigene ID, or Affymetrix probeset ID. *Gene Indices: As of July 1, 2008, there are 111 publicly available gene indices. They are separated into 4 categories for better organization and easier access. Animal: 41, Plant: 45, Protist: 15, Fungal: 10 *Genomic Maps: Human, mouse, rat, chicken, drosophila melanogaster, zebrafish, mosquito, caenorhabditis elegans, Arabidopsis thaliana, rice, yeast, fission yeast Dana-Farber Cancer Institute (DFCI) Gene Indices Software Tools: *TGI Clustering tools (TGICL): a software system for fast clustering of large EST datasets. *GICL: this package contains the scripts and all the necessary pre-compiled binaries for 32bit Linux systems. *clview: an assembly file viewer. *SeqClean:a script for automated trimming and validation of ESTs or other DNA sequences by screening for various contaminants, low quality and low-complexity sequences. *cdbfasta/cdbyank: fast indexing/retrieval of fasta records from flat file databases. *DAS/XML Genomic Viewer The Genomic viewer borrows modules from http://www.biodas.org (lstein (at) cshl.org) & http://webreference.com.
Proper citation: Gene Index Project (RRID:SCR_002148) Copy
http://www.ncbi.nlm.nih.gov/HTGS/
Database of high-throughput genome sequences from large-scale genome sequencing centers, including unfinished and finished sequences. It was created to accommodate a growing need to make unfinished genomic sequence data rapidly available to the scientific community in a coordinated effort among the International Nucleotide Sequence databases, DDBJ, EMBL, and GenBank. Sequences are prepared for submission by using NCBI's software tools Sequin or tbl2asn. Each center has an FTP directory into which new or updated sequence files are placed. Sequence data in this division are available for BLAST homology searches against either the htgs database or the month database, which includes all new submissions for the prior month. Unfinished HTG sequences containing contigs greater than 2 kb are assigned an accession number and deposited in the HTG division. A typical HTG record might consist of all the first-pass sequence data generated from a single cosmid, BAC, YAC, or P1 clone, which together make up more than 2 kb and contain one or more gaps. A single accession number is assigned to this collection of sequences, and each record includes a clear indication of the status (phase 1 or 2) plus a prominent warning that the sequence data are unfinished and may contain errors. The accession number does not change as sequence records are updated; only the most recent version of a HTG record remains in GenBank.
Proper citation: High Throughput Genomic Sequences Division (RRID:SCR_002150) Copy
http://bioafrica.mrc.ac.za/index.html
The BioAfrica HIV-1 Proteomics Resource is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. HIV Proteomics Resource contains information about each HIV-1 gene product in regard to expression, post-transcriptional / post-translational modifications, localization, functional activities, and potential interactions with viral and host macromolecules. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites.
Proper citation: BioAfrica HIV Informatics in Africa (RRID:SCR_002295) Copy
http://www.genoscope.cns.fr/spip/spip.php?lang=en
French national sequencing center with the following resources: * Sequencing ** Genoscope Projects * Environmental genomics ** Microbial diversity in wastewater ** Metabolic genomics * Bioinformatics ** Atelier for comparative genomics ** Computational Systems Biology ** Servers resources *** GGB for Generic Genome Browser: graphic interface for various databases (sequence, annotation, syntenies...) for a given organism. *** MaGe for Magnifying Microbial Genomes: annotation system for microbial genomes.
Proper citation: Genoscope (RRID:SCR_002172) Copy
Original SAMTOOLS package has been split into three separate repositories including Samtools, BCFtools and HTSlib. Samtools for manipulating next generation sequencing data used for reading, writing, editing, indexing,viewing nucleotide alignments in SAM,BAM,CRAM format. BCFtools used for reading, writing BCF2,VCF, gVCF files and calling, filtering, summarising SNP and short indel sequence variants. HTSlib used for reading, writing high throughput sequencing data.
Proper citation: SAMTOOLS (RRID:SCR_002105) Copy
https://ftp.ncbi.nlm.nih.gov/pub/mhc/mhc/Final%20Archive/
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 23, 2019 Database was open, publicly accessible platform for DNA and clinical data related to human Major Histocompatibility Complex (MHC). Data from IHWG workshops were provided as well., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: dbMHC (RRID:SCR_002302) Copy
Maintains and provides archival, retrieval and analytical resources for biological information. Central DDBJ resource consists of public, open-access nucleotide sequence databases including raw sequence reads, assembly information and functional annotation. Database content is exchanged with EBI and NCBI within the framework of the International Nucleotide Sequence Database Collaboration (INSDC). In 2011, DDBJ launched two new resources: DDBJ Omics Archive and BioProject. DOR is archival database of functional genomics data generated by microarray and highly parallel new generation sequencers. Data are exchanged between the ArrayExpress at EBI and DOR in the common MAGE-TAB format. BioProject provides organizational framework to access metadata about research projects and data from projects that are deposited into different databases.
Proper citation: DNA DataBank of Japan (DDBJ) (RRID:SCR_002359) Copy
http://www.ncbi.nlm.nih.gov/genome
Database that organizes information on genomes including sequences, maps, chromosomes, assemblies, and annotations in six major organism groups: Archaea, Bacteria, Eukaryotes, Viruses, Viroids, and Plasmids. Genomes of over 1,200 organisms can be found in this database, representing both completely sequenced organisms and those for which sequencing is in progress. Users can browse by organism, and view genome maps and protein clusters. Links to other prokaryotic and archaeal genome projects, as well as BLAST tools and access to the rest of the NCBI online resources are available.
Proper citation: NCBI Genome (RRID:SCR_002474) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.