Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
A publicly available database of Transposed elements (TEs) which are located within protein-coding genes of 7 organisms: human, mouse, chicken, zebrafish, fruilt fly, nematode and sea squirt. Using TranspoGene the user can learn about the many aspects of the effect these TEs have on their hosting genes, such as: exonization events (including alternative splicing-related data), insertion of TEs into introns, exons, and promoters, specific location of the TE over the gene, evolutionary divergence of the TE from its consensus sequence and involvement in diseases. TranspoGene database is quickly searchable through its website, enables many kinds of searches and is available for download. TranspoGene contains information regarding specific type and family of the TEs, genomic and mRNA location, sequence, supporting transcript accession and alignment to the TE consensus sequence. The database also contains host gene specific data: gene name, genomic location, Swiss-Prot and RefSeq accessions, diseases associated with the gene and splicing pattern. The TranspoGene and microTranspoGene databases can be used by researchers interested in the effect of TE insertion on the eukaryotic transcriptome.
Proper citation: TranspoGene (RRID:SCR_005634) Copy
http://www.gene-regulation.com/pub/databases.html#transfac
Manually curated database of eukaryotic transcription factors, their genomic binding sites and DNA binding profiles. Used to predict potential transcription factor binding sites.
Proper citation: TRANSFAC (RRID:SCR_005620) Copy
http://edwardslab.bmcb.georgetown.edu/downloads/
The Peptide Sequence Database contains putative peptide sequences from human, mouse, rat, and zebrafish. Compressed to eliminate redundancy, these are about 40 fold smaller than a brute force enumeration. Current and old releases are available for download. Each species'' peptide sequence database comprises peptide sequence data from releveant species specific UniGene and IPI clusters, plus all sequences from their consituent EST, mRNA and protein sequence databases, namely RefSeq proteins and mRNAs, UniProt''s SwissProt and TrEMBL, GenBank mRNA, ESTs, and high-throughput cDNAs, HInv-DB, VEGA, EMBL, IPI protein sequences, plus the enumeration of all combinations of UniProt sequence variants, Met loss PTM, and signal peptide cleavages. The README file contains some information about the non amino-acid symbols O (digest site corresponding to a protein N- or C-terminus) and J (no digest sequence join) used in these peptide sequence databases and information about how to configure various search engines to use them. Some search engines handle (very) long sequences badly and in some cases must be patched to use these peptide sequence databases. All search engines supported by the PepArML meta-search engine can (or can be patched to) successfully search these peptide sequence databases.
Proper citation: Peptide Sequence Database (RRID:SCR_005764) Copy
http://indel.bioinfo.sdu.edu.cn/gridsphere/gridsphere
THIS RESOURCE IS NO LONGER IN SERVCE, documented September 2, 2016. Indel Flanking Region Database is an online resource for indels and the flanking regions of proteins in SCOP superfamilies, including amino acid sequences, lengths, locations, secondary structure constitutions, hydrophilicity / hydrophobicity, domain information, 3D structures and so on. It aims at providing a comprehensive dataset for analyzing the qualities of amino acid insertion/deletions(indels), substitutions and the relationship between them. The indels were obtained through the pairwise alignment of homologous structures in SCOP superfamilies. The IndelFR database contains 2,925,017 indels with flanking regions extracted from 373,402 structural alignment pairs of 12,573 non-redundant domains from 1053 superfamilies. IndelFR has already been used for molecular evolution studies and may help to promote future functional studies of indels and their flanking regions.
Proper citation: IndelFR - Indel Flanking Region Database (RRID:SCR_006050) Copy
http://www.ebi.ac.uk/thornton-srv/databases/FunTree/
FunTree provides a range of data resources to detect the evolution of enzyme function within distant structurally related clusters within domain super families as determined by CATH. To access the resource enter a specific CATH superfamily code or search for a structure / sequence / function (either via a EC code or KEGG ligand / reaction ID, PDB ID or UniProtKB ID). Or browse the resource via superfamily / function / structure / metabolites & reactions via the menu on the left panel. FunTree is a new resource that brings together sequence, structure, phylogenetic, chemical and mechanistic information for structurally defined enzyme superfamilies. Gathering together this range of data into a single resource allows the investigation of how novel enzyme functions have evolved within a structurally defined superfamily as well as providing a means to analyse trends across many superfamilies. This is done not only within the context of an enzyme''''s sequence and structure but also the relationships of their reactions. Developed in tandem with the CATH database, it currently comprises 276 superfamilies covering 1800 (70%) of sequence assigned enzyme reactions. Central to the resource are phylogenetic trees generated from structurally informed multiple sequence alignments using both domain structural alignments supplemented with domain sequences and whole sequence alignments based on commonality of multi-domain architectures. These trees are decorated with functional annotations such as metabolite similarity as well as annotations from manually curated resources such the catalytic site atlas and MACiE for enzyme mechanisms.
Proper citation: FunTree (RRID:SCR_006014) Copy
http://www.hpppi.iicb.res.in/btox/
Database of Bacterial ExoToxins for Human is a database of sequences, structures, interaction networks and analytical results for 229 exotoxins, from 26 different human pathogenic bacterial genus. All toxins are classified into 24 different Toxin classes. The aim of DBETH is to provide a comprehensive database for human pathogenic bacterial exotoxins. DBETH also provides a platform to its users to identify potential exotoxin like sequences through Homology based as well as Non-homology based methods. In homology based approach the users can identify potential exotoxin like sequences either running BLASTp against the toxin sequences or by running HMMER against toxin domains identified by DBETH from human pathogenic bacterial exotoxins. In Non-homology based part DBETH uses a machine learning approach to identify potential exotoxins (Toxin Prediction by Support Vector Machine based approach).
Proper citation: DBETH - Database for Bacterial ExoToxins for Humans (RRID:SCR_005908) Copy
This site has been developed by Kazusa DNA Research Institute for the purpose of offering the science community the analyzed sequence data produced by a multi-national Arabidopsis genome sequencing project coordinated by the Arabidopsis Genome Initiatives (AGI). The aim of this service is to enable users to browse the annotated sequence data produced by all the sequencing teams of AGI through an user-friendly graphic display system and search engines. Gene structures proposed on the annotated sequences as well as those predicted by computer programs are presented and each graphic item has a hyperlink to detailed information of the corresponding area. The nucleotide sequence data deposited in GenBank by AGI was downloaded, re-computer-analyzed at Kazusa and parsed results are displayed graphically.
Proper citation: Kazusa Arabidopsis data opening site (RRID:SCR_013511) Copy
http://viewer.shigen.info/cgi-bin/crispr/crispr.cgi
Web tool to show micro homology sequences striding over double strand break point created by CRISPR/Cas9 system. Used to search for CRISPR target site with micro-homology sequences. Used to predict deletion pattern.
Proper citation: NBRP Medaka CRISPR target site (RRID:SCR_018159) Copy
Software tool as catalog of inferred sequence binding preferences. Online library of transcription factors and their DNA binding motifs.
Proper citation: CIS-BP (RRID:SCR_017236) Copy
http://mirwalk.umm.uni-heidelberg.de/
Software tool to store the predicted and the experimentally validated microRNA (miRNA)-target interaction pairs. Predictions within the complete sequence of genes of human, mouse, and rat genomes. Integrates a comparative platform of miRNA-binding sites resulting from ten different prediction datasets.
Proper citation: miRWalk (RRID:SCR_016509) Copy
http://nucleobytes.com/index.php/4peaks
Software application for viewing and editing sequence trace files.
Proper citation: 4Peaks (RRID:SCR_000015) Copy
http://www.glycosciences.de/tools/linucs/
Service that directly converts the commonly used extended representation of complex carbohydrates into the preferred canonical description or into its inverted form. Input: A structure using the extended, non-graphic nomenclature (in ASCII writing) to describe complex carbohydrates as recommended by IUPAC. Output: A linear, unique notation. The source code (written in C), will be distributed so that software developers can easily implement their algorithm within their own application. LINUCS was chosen to fulfill to following conditions: * Input of extended, non-graphic nomenclature to describe carbohydrate structures. * Resulting linear code is closely related to notations and abbreviations recommended by IUPAC. * Number of additional rules to define the priority of the branches is low * Extended nomenclature of complex carbohydrates contains all information to define the hierarchy. * LINUCS is applicable to all types of carbohydrates (macrocyclic system are currently not implemented) . * Remaining unassigned linkage information are tolerated
Proper citation: LINUCS (RRID:SCR_001571) Copy
A curated collection of chaperonin sequence data collected from public databases or generated by a network of collaborators exploiting the cpn60 target in clinical, phylogenetic and microbial ecology studies. The database contains all available sequences for both group I and group II chaperonins. Users can search the database by Chaperonin type, group (I or II), BLAST, or other options, and can also enter and analyze FASTA sequences.
Proper citation: cpnDB: A Chaperonin Database (RRID:SCR_002263) Copy
http://ww2.sanbi.ac.za/Dbases.html
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 23, 2016. The STACKdb is knowledgebase generated by processing EST and mRNA sequences obtained from GenBank through a pipeline consisting of masking, clustering, alignment and variation analysis steps. The STACK project aims to generate a comprehensive representation of the sequence of each of the expressed genes in the human genome by extensive processing of gene fragments to make accurate alignments, highlight diversity and provide a carefully joined set of consensus sequences for each gene. The STACK project is comprised of the STACKdb human gene index, a database of virtual human transcripts, as well as stackPACK, the tools used to create the database. STACKdb is organized into 15 tissue-based categories and one disease category. STACK is a tool for detection and visualization of expressed transcript variation in the context of developmental and pathological states. The data system organizes and reconstructs human transcripts from available public data in the context of expression state. The expression state of a transcript can include developmental state, pathological association, site of expression and isoform of expressed transcript. STACK consensus transcripts are reconstructed from clusters that capture and reflect the growing evidence of transcript diversity. The comprehensive capture of transcript variants is achieved by the use of a novel clustering approach that is tolerant of sub-sequence diversity and does not rely on pairwise alignment. This is in contrast with other gene indexing projects. STACK is generated at least four times a year and represents the exhaustive processing of all publicly available human EST data extracted from GenBank. This processed information can be explored through 15 tissue-specific categories, a disease-related category and a whole-body index
Proper citation: Sequence Tag Alignment and Consensus Knowledgebase Database (RRID:SCR_002156) Copy
The Hepatitis C Virus (HCV) Database Project strives to present HCV-associated genetic and immunologic data in a user-friendly way, by providing access to the central database via web-accessible search interfaces and supplying a number of analysis tools.
Proper citation: HCV Databases (RRID:SCR_002863) Copy
Professionally curated repository for genetics, genomics and related data resources for soybean that contains the most current genetic, physical and genomic sequence maps integrated with qualitative and quantitative traits. SoyBase includes annotated Williams 82 genomic sequence and associated data mining tools. The genetic and sequence views of the soybean chromosomes and the extensive data on traits and phenotypes are extensively interlinked. This allows entry to the database using almost any kind of available information, such as genetic map symbols, soybean gene names or phenotypic traits. The repository maintains controlled vocabularies for soybean growth, development, and traits that are linked to more general plant ontologies. Contributions to SoyBase or the Breeder''s Toolbox are welcome.
Proper citation: SoyBase (RRID:SCR_005096) Copy
A web program that can locate residue periodicities in either amino acid or DNA sequences. It is based on an algorithm of Dr. A.D. McLachlan (1977). NOTE: You must use a Java compatible browser to run the application.
Proper citation: FT (RRID:SCR_006228) Copy
The Kabat Database determines the combining site of antibodies based on the available amino acid sequences. The precise delineation of complementarity determining regions (CDR) of both light and heavy chains provides the first example of how properly aligned sequences can be used to derive structural and functional information of biological macromolecules. The Kabat database now includes nucleotide sequences, sequences of T cell receptors for antigens (TCR), major histocompatibility complex (MHC) class I and II molecules, and other proteins of immunological interest. The Kabat Database searching and analysis tools package is an ASP.NET web-based portal containing lookup tools, sequence matching tools, alignment tools, length distribution tools, positional correlation tools and much more. The searching and analysis tools are custom made for the aligned data sets contained in both the SQL Server and ASCII text flat file formats. The searching and analysis tools may be run on a single PC workstation or in a distributed environment. The analysis tools are written in ASP.NET and C# and are available in Visual Studio .NET 2003/2005/2008 formats. The Kabat Database was initially started in 1970 to determine the combining site of antibodies based on the available amino acid sequences at that time. Bence Jones proteins, mostly from human, were aligned, using the now-known Kabat numbering system, and a quantitative measure, variability, was calculated for every position. Three peaks, at positions 24-34, 50-56 and 89-97, were identified and proposed to form the complementarity determining regions (CDR) of light chains. Subsequently, antibody heavy chain amino acid sequences were also aligned using a different numbering system, since the locations of their CDRs (31-35B, 50-65 and 95-102) are different from those of the light chains. CDRL1 starts right after the first invariant Cys 23 of light chains, while CDRH1 is eight amino acid residues away from the first invariant Cys 22 of heavy chains. During the past 30 years, the Kabat database has grown to include nucleotide sequences, sequences of T cell receptors for antigens (TCR), major histocompatibility complex (MHC) class I and II molecules and other proteins of immunological interest. It has been used extensively by immunologists to derive useful structural and functional information from the primary sequences of these proteins.
Proper citation: Kabat Database of Sequences of Proteins of Immunological Interest (RRID:SCR_006465) Copy
http://athina.biol.uoa.gr/bioinformatics/NON-RED/index.html
A web tool to select biological sequences from a given set, with similarity / homology less than a user-defined level. This web-based application takes as input a set of N sequences and outputs a set of sequences of user-determined redundancy. Initially, the algorithm runs an all-against-all BLAST alignment on the input data set and creates an NxN matrix of pairwise distances defined by the similarity percentages. In the next step, the algorithm removes the sequence with the largest number of neighbors, causing that sequence not to be counted as a neighbor of any other sequence during the next iterations. It then reassesses the number of neighbors of each sequence and repeats the previous step until the sequences left over have no more neighbors. The user can specify the similarity (%) threshold and the minimum coverage length of the alignments. Sequences with a similarity below the threshold or a smaller coverage than the minimum length are not considered to be neighbors.
Proper citation: NON-RED (RRID:SCR_006225) Copy
http://athina.biol.uoa.gr/bioinformatics/waveTM/
A web tool for the prediction of transmembrane segments in alpha-helical membrane proteins. A sliding window of 20 residues is used in order to calculate an average residue hydrophobicity profile, using a hydrophobicity scale. Discrete Wavelet Transform is applied on the average residue hydrophobicity signal and the different frequency coefficients produced are adaptively thresholded so that a denoised signal is reconstructed. A dynamic programming algorithm processes the denoised signal to provide the optimal model for the number, the length and the location of membrane-spanning segments. The end points of the predicted segments are extended to include flanking hydrophobic residues. Topology prediction can also be obtained in conjunction with OrienTM (Liakopoulos et al, 2001). Analysis of a non-redundant test set, provides a ~95% per segment accuracy and ~90% per residue accuracy. Now, you can: * Run waveTM on a sequence * Browse the results obtained with the algorithm * View additional material concerning the hydrophobicity scale
Proper citation: waveTM (RRID:SCR_006199) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.