Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on May 2nd, 2023. Sequence composition based classifier for metagenomic sequences. It works by capturing signatures of each sequence based on the sequence composition. Each sequence is modeled as a walk in a de Bruijn graph with underlying Markov chain properties. ClaMS captures stationary parameters of the underlying Markov chain as well as structural parameters of the underlying de Bruijn graph to form this signature. In practice, for each sequence to binned, such a signature is computed and matched to similar signatures computed for the training sets. The best match that also qualifies the normalized distance cut-off wins. In the case that the best match does not qualify this cut-off, the sequence remains un-binned.
Proper citation: Classifier for Metagenomic Sequences (RRID:SCR_004929) Copy
https://github.com/tk2/RetroSeq
A tool for discovery and genotyping of transposable element variants (TEVs) (also known as mobile element insertions) from next-gen sequencing reads aligned to a reference genome in BAM format. The goal is to call TEVs that are not present in the reference genome but present in the sample that has been sequenced. It should be noted that RetroSeq can be used to locate any class of viral insertion in any species where whole-genome sequencing data with a suitable reference genome is available. RetroSeq is a two phase process, the first being the read pair discovery phase where discorandant mate pairs are detected and assigned to a TE class (Alu, SINE, LINE, etc.) by using either the annotated TE elements in the reference and/or aligned with Exonerate to the supplied library of viral sequences.
Proper citation: RetroSeq (RRID:SCR_005133) Copy
http://seqant.genetics.emory.edu/
A free web service and open source software package that performs rapid, automated annotation of DNA sequence variants (single base mutations, insertions, deletions) discovered with any sequencing platform. Variant sites are characterized with respect to their functional type (Silent, Replacement, 5' UTR, 3' UTR, Intronic, Intergenic), whether they have been previously submitted to dbSNP, and their evolutionary conservation. Annotated variants can be viewed directly on the web browser, downloaded in a tab delimited text file, or directly uploaded in a Browser Extended Data (BED) format to the UCSC genome browser. SeqAnt further identifies all loci harboring two or more coding sequence variants that help investigators identify potential compound heterozygous loci within exome sequencing experiments. In total, SeqAnt resolves a significant bottleneck by allowing an investigator to rapidly prioritize the functional analysis of those variants of interest.
Proper citation: SeqAnt (RRID:SCR_005186) Copy
A software package for the analysis of nucleotide polymorphism from aligned DNA sequence data. DnaSP can estimate several measures of DNA sequence variation within and between populations (in noncoding, synonymous or nonsynonymous sites, or in various sorts of codon positions), as well as linkage disequilibrium, recombination, gene flow and gene conversion parameters. DnaSP can also carry out several tests of neutrality: Hudson, Kreitman and Aguad (1987), Tajima (1989), McDonald and Kreitman (1991), Fu and Li (1993), and Fu (1997) tests. Additionally, DnaSP can estimate the confidence intervals of some test-statistics by the coalescent. The results of the analyses are displayed on tabular and graphic form.
Proper citation: DnaSP (RRID:SCR_003067) Copy
Digital atlas of gene expression patterns in developing and adult mouse. Several reference atlases are also available through this site. Expression patterns are determined by non-radioactive in situ hybridization on serial tissue sections. Sections are available from several developmental ages: E10.5, E14.5 (whole embryos), E15.5, P7 and P56 (brains only). To retrieve expression patterns, search by gene name, site of expression, GenBank accession number or sequence homology. For viewing expression patterns, GenePaint.org features virtual microscope tool that enables zooming into images down to cellular resolution.
Proper citation: GenePaint (RRID:SCR_003015) Copy
https://code.google.com/p/gutentag/
An interactive, user-editable genetic sequence database tool, targeted at molecular biology research groups that can be browsed using tags. The tool is Web 2.0-flavoured, allowing users to do more than just retrieve information. Its focus on user-editability is supported by the use of tags (metadata) associated with genetic sequences. Several methods of retrieving stored data are available including tag-clouds, BLAST and keyword searches. Also, sequence tags related to HGNC gene names, conserved domains (CDD) and GO terms can be automatically generated given sequence data. The tool is constructed using the high-level Python web framework, Django, with a SQLite3 backend.
Proper citation: Gutentag (RRID:SCR_003051) Copy
http://bibiserv.techfak.uni-bielefeld.de/dialign/
Tool for multiple sequence alignment using various sources of external information that is particularly useful to detect local homologies in sequences with low overall similarity. While standard alignment methods rely on comparing single residues and imposing gap penalties, DIALIGN constructs pairwise and multiple alignments by comparing entire segments of the sequences. No gap penalty is used. This approach can be used for both global and local alignment, but it is particularly successful in situations where sequences share only local homologies. Several versions of DIALIGN are available online at GOBICS, http://dialign.gobics.de/
Proper citation: DIALIGN (RRID:SCR_003041) Copy
http://compbio.cs.sfu.ca/software-novelseq
Software pipeline to detect novel sequence insertions using high throughput paired-end whole genome sequencing data.
Proper citation: NovelSeq (RRID:SCR_003136) Copy
Database to catalog experimentally determined interactions between proteins combining information from a variety of sources to create a single, consistent set of protein-protein interactions that can be downloaded in a variety of formats. The data were curated, both, manually and also automatically using computational approaches that utilize the the knowledge about the protein-protein interaction networks extracted from the most reliable, core subset of the DIP data. Because the reliability of experimental evidence varies widely, methods of quality assessment have been developed and utilized to identify the most reliable subset of the interactions. This CORE set can be used as a reference when evaluating the reliability of high-throughput protein-protein interaction data sets, for development of prediction methods, as well as in the studies of the properties of protein interaction networks. Tools are available to analyze, visualize and integrate user's own experimental data with the information about protein-protein interactions available in the DIP database. The DIP database lists protein pairs that are known to interact with each other. By interact they mean that two amino acid chains were experimentally identified to bind to each other. The database lists such pairs to aid those studying a particular protein-protein interaction but also those investigating entire regulatory and signaling pathways as well as those studying the organization and complexity of the protein interaction network at the cellular level. Registration is required to gain access to most of the DIP features. Registration is free to the members of the academic community. Trial accounts for the commercial users are also available.
Proper citation: Database of Interacting Proteins (DIP) (RRID:SCR_003167) Copy
http://wiki.c2b2.columbia.edu/honiglab_public/index.php/Main_Page
Laboratory portal, including software, web-based tools, databases and data sets, related to their research that focuses on the development and application of biophysical and bioinformatics methods aimed at understanding the structural and energetic origins of protein-protein, protein-nucleic acid, and protein-membrane interactions. Their work includes fundamental theoretical research, the development of software tools, and applications to problems of biological importance. In this regard they maintain an active collaborative computational and experimental research program on the molecular basis of cell-cell adhesion. Other problems of current interest include protein structure prediction, the organization of protein sequence/structure space, the prediction of protein function based on protein structure, the structural origins of specificity in protein-DNA interactions, RNA function and, more generally, the electrostatic properties of biological macromolecules.
Proper citation: Honig Lab (RRID:SCR_003410) Copy
https://services.healthtech.dtu.dk/
Center for Biological Sequence Analysis of the Technical University of Denmark conducts basic research in the field of bioinformatics and systems biology and directs its research primarily towards topics related to the elucidation of the functional aspects of complex biological mechanisms. A large number of computational methods have been produced, which are offered to others via WWW servers. Several data sets are also available. The center also has experimental efforts in gene expression analysis using DNA chips and data generation in relation to the physical and structural properties of DNA. The on-line prediction services at CBS are available as interactive input forms. Most of the servers are also available as stand-alone software packages with the same functionality. In addition, for some servers, programmatic access is provided in the form of SOAP-based Web Services. The center also educates engineering students in biotechnology and systems biology and offers a wide range of courses in bioinformatics, systems biology, human health, microbiology and nutrigenomics.
Proper citation: DTU Center for Biological Sequence Analysis (RRID:SCR_003590) Copy
http://iubio.bio.indiana.edu/webapps/SeWeR/
Sequence analysis using Web Resources (SeWeR) is an integrated, Dynamic HTML (DHTML) interface to commonly used bioinformatics services available on the World Wide Web. It is highly customizable, extendable, platform neutral, completely server-independent and can be hosted as a web page as well as being used as stand-alone software running within a web browser. It doesn''t require any server to host itself. The goal of SeWeR is to turn your web-browser into a powerful sequence-analysis tool. It is written entirely in JavaScript1.2. SeWeR can be downloaded and mirrored freely. The whole package is just around 300K. You can even run it from a floppy. SeWeR is not compatible with Netscape 6. SeWeR now generates graphics. Savvy is a plasmid drawing software that generates plasmid map in the revolutionary Scalable Vector Graphics format from W3C.
Proper citation: SeWeR - SEquence analysis using WEb Resources (RRID:SCR_004167) Copy
http://www.glycosciences.de/tools/linucs/
Service that directly converts the commonly used extended representation of complex carbohydrates into the preferred canonical description or into its inverted form. Input: A structure using the extended, non-graphic nomenclature (in ASCII writing) to describe complex carbohydrates as recommended by IUPAC. Output: A linear, unique notation. The source code (written in C), will be distributed so that software developers can easily implement their algorithm within their own application. LINUCS was chosen to fulfill to following conditions: * Input of extended, non-graphic nomenclature to describe carbohydrate structures. * Resulting linear code is closely related to notations and abbreviations recommended by IUPAC. * Number of additional rules to define the priority of the branches is low * Extended nomenclature of complex carbohydrates contains all information to define the hierarchy. * LINUCS is applicable to all types of carbohydrates (macrocyclic system are currently not implemented) . * Remaining unassigned linkage information are tolerated
Proper citation: LINUCS (RRID:SCR_001571) Copy
A curated collection of chaperonin sequence data collected from public databases or generated by a network of collaborators exploiting the cpn60 target in clinical, phylogenetic and microbial ecology studies. The database contains all available sequences for both group I and group II chaperonins. Users can search the database by Chaperonin type, group (I or II), BLAST, or other options, and can also enter and analyze FASTA sequences.
Proper citation: cpnDB: A Chaperonin Database (RRID:SCR_002263) Copy
http://ww2.sanbi.ac.za/Dbases.html
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 23, 2016. The STACKdb is knowledgebase generated by processing EST and mRNA sequences obtained from GenBank through a pipeline consisting of masking, clustering, alignment and variation analysis steps. The STACK project aims to generate a comprehensive representation of the sequence of each of the expressed genes in the human genome by extensive processing of gene fragments to make accurate alignments, highlight diversity and provide a carefully joined set of consensus sequences for each gene. The STACK project is comprised of the STACKdb human gene index, a database of virtual human transcripts, as well as stackPACK, the tools used to create the database. STACKdb is organized into 15 tissue-based categories and one disease category. STACK is a tool for detection and visualization of expressed transcript variation in the context of developmental and pathological states. The data system organizes and reconstructs human transcripts from available public data in the context of expression state. The expression state of a transcript can include developmental state, pathological association, site of expression and isoform of expressed transcript. STACK consensus transcripts are reconstructed from clusters that capture and reflect the growing evidence of transcript diversity. The comprehensive capture of transcript variants is achieved by the use of a novel clustering approach that is tolerant of sub-sequence diversity and does not rely on pairwise alignment. This is in contrast with other gene indexing projects. STACK is generated at least four times a year and represents the exhaustive processing of all publicly available human EST data extracted from GenBank. This processed information can be explored through 15 tissue-specific categories, a disease-related category and a whole-body index
Proper citation: Sequence Tag Alignment and Consensus Knowledgebase Database (RRID:SCR_002156) Copy
The Hepatitis C Virus (HCV) Database Project strives to present HCV-associated genetic and immunologic data in a user-friendly way, by providing access to the central database via web-accessible search interfaces and supplying a number of analysis tools.
Proper citation: HCV Databases (RRID:SCR_002863) Copy
Professionally curated repository for genetics, genomics and related data resources for soybean that contains the most current genetic, physical and genomic sequence maps integrated with qualitative and quantitative traits. SoyBase includes annotated Williams 82 genomic sequence and associated data mining tools. The genetic and sequence views of the soybean chromosomes and the extensive data on traits and phenotypes are extensively interlinked. This allows entry to the database using almost any kind of available information, such as genetic map symbols, soybean gene names or phenotypic traits. The repository maintains controlled vocabularies for soybean growth, development, and traits that are linked to more general plant ontologies. Contributions to SoyBase or the Breeder''s Toolbox are welcome.
Proper citation: SoyBase (RRID:SCR_005096) Copy
Database that provides the genome sequence assembly of the International Rice Genome Sequencing Project (IRGSP), manually curated annotation of the sequence, and other genomics information that could be useful for comprehensive understanding of the rice biology. RAP-DB contains clone positions, structures and functions of genes validated by cDNAs, RNA genes detected by massively parallel signature sequencing (MPSS) technology and sequence similarity, flanking sequences of mutant lines, transposable elements, etc. Other annotation data such as Gnomon can be displayed along with those of RAP for comparison.
Proper citation: RAP-DB (RRID:SCR_006610) Copy
http://www.ncbi.nlm.nih.gov/CCDS/
Database (anonymous FTP) resulting from a collaborative effort to identify a core set of human and mouse protein coding regions that are consistently annotated and of high quality. The long term goal is to support convergence towards a standard set of gene annotations. Collaborators are EBI, NCBI, UCSC, WTSI and the initial results are also available from the participants'''' genome browser Web sites. In addition, CCDS identifiers are indicated on the relevant NCBI RefSeq and Entrez Gene records and in Map Viewer displays of RNA (RefSeq) and Gene annotations on the reference assembly.
Proper citation: Consensus CDS (RRID:SCR_006729) Copy
Database devoted to protein domains. It is also a collection of tools for the investigation of the relationships between protein sequences and motifs described on them.
Proper citation: MyHits (RRID:SCR_006757) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.