Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
Collection of data of protein sequence and functional information. Resource for protein sequence and annotation data. Consortium for preservation of the UniProt databases: UniProt Knowledgebase (UniProtKB), UniProt Reference Clusters (UniRef), and UniProt Archive (UniParc), UniProt Proteomes. Collaboration between European Bioinformatics Institute (EMBL-EBI), SIB Swiss Institute of Bioinformatics and Protein Information Resource. Swiss-Prot is a curated subset of UniProtKB.
Proper citation: UniProt (RRID:SCR_002380) Copy
http://www.uniprot.org/taxonomy/
NEWT is the taxonomy database maintained by the UniProt group. It integrates taxonomy data compiled in the NCBI database and data specific to the UniProt Knowledgebase. Browse by hierarchy, List all, or Complete proteomes. Organisms are classified in a hierarchical tree structure. Our taxonomy database contains every node (taxon) of the tree. UniProtKB taxonomy data is manually curated: next to manually verified organism names, we provide a selection of external links, organism strains and viral host information. Species with protein sequences stored in the UniProt Knowledgebase are named according to UniProt nomenclature. We endeavour to maintain a list of manually curated species names for which protein sequence data is available. In particular, we have adopted a systematic convention for naming viral and bacterial strains and isolates. Links to external sites are chosen by the UniProt taxonomy team and show pictures and various scientific data of interest (taxonomy, biology, physiology,...).
Proper citation: NEWT (RRID:SCR_004477) Copy
http://www.uniprot.org/uniparc/
Database that contains publicly available protein sequences with stable and unique identifiers (UPI) which are never removed, changed or reassigned. UniParc tracks sequence changes in the source databases and archives the history of all changes. Information other than protein sequence must be retrieved from the UniParc source databases using the database cross-references.
Proper citation: UniParc (RRID:SCR_005818) Copy
http://www.uniprot.org/help/uniref
Databases which provide clustered sets of sequences from UniProt Knowledgebase and selected UniParc records, in order to obtain complete coverage of sequence space at several resolutions while hiding redundant sequences from view. The UniRef100 database combines identical sequences and sub-fragments with 11 or more residues (from any organism) into a single UniRef entry. The sequence of a representative protein, the accession numbers of all the merged entries, and links to the corresponding UniProtKB and UniParc records are all displayed in the entry. UniRef90 and UniRef50 are built by clustering UniRef100 sequences with 11 or more residues such that each cluster is composed of sequences that have at least 90% (UniRef90) or 50% (UniRef50) sequence identity to the longest sequence (UniRef seed sequence). All the sequences in each cluster are ranked to facilitate the selection of a representative sequence for the cluster.
Proper citation: UniRef (RRID:SCR_010646) Copy
http://www.uniprot.org/help/uniprotkb
Central repository for collection of functional information on proteins, with accurate and consistent annotation. In addition to capturing core data mandatory for each UniProtKB entry (mainly, the amino acid sequence, protein name or description, taxonomic data and citation information), as much annotation information as possible is added. This includes widely accepted biological ontologies, classifications and cross-references, and experimental and computational data. The UniProt Knowledgebase consists of two sections, UniProtKB/Swiss-Prot and UniProtKB/TrEMBL. UniProtKB/Swiss-Prot (reviewed) is a high quality manually annotated and non-redundant protein sequence database which brings together experimental results, computed features, and scientific conclusions. UniProtKB/TrEMBL (unreviewed) contains protein sequences associated with computationally generated annotation and large-scale functional characterization that await full manual annotation. Users may browse by taxonomy, keyword, gene ontology, enzyme class or pathway.
Proper citation: UniProtKB (RRID:SCR_004426) Copy
http://www.uniprot.org/program/Chordata
Data set of manually annotated chordata-specific proteins as well as those that are widely conserved. The program keeps existing human entries up-to-date and broadens the manual annotation to other vertebrate species, especially model organisms, including great apes, cow, mouse, rat, chicken, zebrafish, as well as Xenopus laevis and Xenopus tropicalis. A draft of the complete human proteome is available in UniProtKB/Swiss-Prot and one of the current priorities of the Chordata protein annotation program is to improve the quality of human sequences provided. To this aim, they are updating sequences which show discrepancies with those predicted from the genome sequence. Dubious isoforms, sequences based on experimental artifacts and protein products derived from erroneous gene model predictions are also revisited. This work is in part done in collaboration with the Hinxton Sequence Forum (HSF), which allows active exchange between UniProt, HAVANA, Ensembl and HGNC groups, as well as with RefSeq database. UniProt is a member of the Consensus CDS project and thye are in the process of reviewing their records to support convergence towards a standard set of protein annotation. They also continuously update human entries with functional annotation, including novel structural, post-translational modification, interaction and enzymatic activity data. In order to identify candidates for re-annotation, they use, among others, information extraction tools such as the STRING database. In addition, they regularly add new sequence variants and maintain disease information. Indeed, this annotation program includes the Variation Annotation Program, the goal of which is to annotate all known human genetic diseases and disease-linked protein variants, as well as neutral polymorphisms.
Proper citation: UniProt Chordata protein annotation program (RRID:SCR_007071) Copy
Non-profit academic organization for research and services in bioinformatics. Provides freely available data from life science experiments, performs basic research in computational biology, and offers user training programme, manages databases of biological data including nucleic acid, protein sequences, and macromolecular structures. Part of EMBL.
Proper citation: European Bioinformatics Institute (RRID:SCR_004727) Copy
http://www.bioextract.org/GuestLogin
An open, web-based system designed to aid researchers in the analysis of genomic data by providing a platform for the creation of bioinformatic workflows. Scientific workflows are created within the system by recording tasks performed by the user. These tasks may include querying multiple, distributed data sources, saving query results as searchable data extracts, and executing local and web-accessible analytic tools. The series of recorded tasks can then be saved as a reproducible, sharable workflow available for subsequent execution with the original or modified inputs and parameter settings. Integrated data resources include interfaces to the National Center for Biotechnology Information (NCBI) nucleotide and protein databases, the European Molecular Biology Laboratory (EMBL-Bank) non-redundant nucleotide database, the Universal Protein Resource (UniProt), and the UniProt Reference Clusters (UniRef) database. The system offers access to numerous preinstalled, curated analytic tools and also provides researchers with the option of selecting computational tools from a large list of web services including the European Molecular Biology Open Software Suite (EMBOSS), BioMoby, and the Kyoto Encyclopedia of Genes and Genomes (KEGG). The system further allows users to integrate local command line tools residing on their own computers through a client-side Java applet.
Proper citation: BioExtract (RRID:SCR_005397) Copy
Center with mission to conduct and support medical research and research training and to disseminate science-based information on diabetes and other endocrine and metabolic diseases. The NIDDK supports a wide range of medical research through grants to universities and other medical research institutions across the country.
Proper citation: NIDDK - National Institute of Diabetes and Digestive and Kidney Diseases (RRID:SCR_012895) Copy
http://xldb.fc.ul.pt/biotools/rebil/goa/
A tool for assisting the GO annotation of UniProt entries by linking the GO terms present in the uncurated annotations with evidence text automatically extracted from the documents linked to UniProt entries. Platform: Online tool
Proper citation: GoAnnotator (RRID:SCR_005792) Copy
GOTaxExplorer presents a new approach to comparative genomics that integrates functional information and families with the taxonomic classification. It integrates UniProt, Gene Ontology, NCBI Taxonomy, Pfam and SMART in one database. GOTaxExplorer provides four different query types: selection of entity sets, comparison of sets of Pfam families, semantic comparison of sets of GO terms, functional comparison of sets of gene products. This permits to select custom sets of GO terms, families or taxonomic groups. For example, it is possible to compare arbitrarily selected organisms or groups of organisms from the taxonomic tree on the basis of the functionality of their genes. Furthermore, it enables to determine the distribution of specific molecular functions or protein families in the taxonomy. The comparison of sets of GO terms allows to assess the semantic similarity of two different GO terms. The functional comparison of gene products makes it possible to identify functionally equivalent and functionally related gene products from two organisms on the basis of GO annotations and a semantic similarity measure for GO. Platform: Online tool, Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: GOTaxExplorer (RRID:SCR_005720) Copy
http://www.ebi.ac.uk/webservices/whatizit/info.jsf
A text processing system that allows you to do textmining tasks on text. It is great at identifying molecular biology terms and linking them to publicly available databases. Whatizit is also a Medline abstracts retrieval/search engine. Instead of providing the text by Copy&Paste, you can launch a Medline search. The abstracts that match your search criteria are retrieved and processed by a pipeline of your choice. Whatizit is also available as 1) a webservice and as 2) a streamed servlet. The webservice allows you to enrich content within your website in a similar way as in the wikipedia. The streamed servlet allows you to process large amounts of text.
Proper citation: Whatizit (RRID:SCR_005824) Copy
http://www.pathwaycommons.org/pc
Database of publicly available pathways from multiple organisms and multiple sources represented in a common language. Pathways include biochemical reactions, complex assembly, transport and catalysis events, and physical interactions involving proteins, DNA, RNA, small molecules and complexes. Pathways were downloaded directly from source databases. Each source pathway database has been created differently, some by manual extraction of pathway information from the literature and some by computational prediction. Pathway Commons provides a filtering mechanism to allow the user to view only chosen subsets of information, such as only the manually curated subset. The quality of Pathway Commons pathways is dependent on the quality of the pathways from source databases. Pathway Commons aims to collect and integrate all public pathway data available in standard formats. It currently contains data from nine databases with over 1,668 pathways, 442,182 interactions,414 organisms and will be continually expanded and updated. (April 2013)
Proper citation: Pathway Commons (RRID:SCR_002103) Copy
http://www.imexconsortium.org/
Interaction database from international collaboration between major public interaction data providers who share curation effort and develop set of curation rules when capturing data from both directly deposited interaction data or from publications in peer reviewed journals. Performs complete curation of all protein-protein interactions experimentally demonstrated within publication and makes them available in single search interface on common website. Provides data in standards compliant download formats. IMEx partners produce their own separate resources, which range from all encompassing molecular interaction databases, such as are maintained by IntAct, MINT and DIP, organism-centric resources such as BioGrid or MPIDB or biological domain centric, such as MatrixDB. They have committed to making records available, via PSICQUIC webservice, which have been curated to IMEx rules and are available to users as single, non-redundant set of curated publications which can be searched at the IMEx website. Data is made available in standards-compliant tab-deliminated and XML formats, enabling to visualize data using wide range of tools. Consortium is open to participation of additional partners and encourages deposition of data, prior to publication, and will supply unique accession numbers which may be referenced within final article. Submitters may send their data directly to any of member databases using variety of formats, but should conform to guidelines as to minimum information required to describe data.
Proper citation: IMEx - The International Molecular Exchange Consortium (RRID:SCR_002805) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 14,2026. Integrated database of genomic, expression and protein data for Drosophila, Anopheles, C. elegans and other organisms. You can run flexible queries, export results and analyze lists of data. FlyMine presents data in categories, with each providing information on a particular type of data (for example Gene Expression or Protein Interactions). Template queries, as well as the QueryBuilder itself, allow you to perform searches that span data from more than one category. Advanced users can use a flexible query interface to construct their own data mining queries across the multiple integrated data sources, to modify existing template queries or to create your own template queries. Access our FlyMine data via our Application Programming Interface (API). We provide client libraries in the following languages: Perl, Python, Ruby and & Java API
Proper citation: FlyMine (RRID:SCR_002694) Copy
Project that developed an open access discovery platform, called Open Pharmacological Space (OPS), via a semantic web approach, integrating pharmacological data from a variety of information resources and tools and services to question this integrated data to support pharmacological research. The project is based upon the assimilation of data already stored as triples, in the form subject-predicate-object. The software and data are available for download and local installation, under an open source and open access model. Tools and services are provided to query and visualize this data, and a sustainability plan will be in place, continuing the operation of the Open PHACTS Discovery Platform after the project funding ends. Throughout the project, a series of recommendations will be developed in conjunction with the community, building on open standards, to ensure wide applicability of the approaches used for integration of data.
Proper citation: Open PHACTS (RRID:SCR_005050) Copy
Comprehensive set of protein domain families automatically generated from UniProt Knowledge Database. Automated clustering of homologous domains generated from global comparison of all available protein sequences., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: ProDom (RRID:SCR_006969) Copy
http://loschmidt.chemi.muni.cz/predictsnp/
Consensus classifier tool that combines six of the top performing tools for the prediction of the effects of mutation on protein function. The obtained results are provided together with annotations extracted from the Protein Mutant Database and the UniProt database. A stand-alone version is also available.
Proper citation: PredictSNP (RRID:SCR_006327) Copy
Platform as a web-based interactive environment to automatically identify, explore and visualize homology and functional annotations for assembled transcripts.
Proper citation: Genotate (RRID:SCR_016659) Copy
http://plantgrn.noble.org/LegumeIP/
LegumeIP is an integrative database and bioinformatics platform for comparative genomics and transcriptomics to facilitate the study of gene function and genome evolution in legumes, and ultimately to generate molecular based breeding tools to improve quality of crop legumes. LegumeIP currently hosts large-scale genomics and transcriptomics data, including: * Genomic sequences of three model legumes, i.e. Medicago truncatula, Glycine max (soybean) and Lotus japonicus, including two reference plant species, Arabidopsis thaliana and Poplar trichocarpa, with the annotation based on UniProt TrEMBL, InterProScan, Gene Ontology and KEGG databases. LegumeIP covers a total 222,217 protein-coding gene sequences. * Large-scale gene expression data compiled from 104 array hybridizations from L. japonicas, 156 array hybridizations from M. truncatula gene atlas database, and 14 RNA-Seq-based gene expression profiles from G. max on different tissues including four common tissues: Nodule, Flower, Root and Leaf. * Systematic synteny analysis among M. truncatula, G. max, L. japonicus and A. thaliana. * Reconstruction of gene family and gene family-wide phylogenetic analysis across the five hosted species. LegumeIP features comprehensive search and visualization tools to enable the flexible query on gene annotation, gene family, synteny, relative abundance of gene expression.
Proper citation: LegumeIP (RRID:SCR_008906) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.