Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://sv.gersteinlab.org/pemer/
Software package as computational framework with simulation-based error models for inferring genomic structural variants from massive paired-end sequencing data. Package is composed of three modules, PEMer workflow, SV-Simulation and BreakDB. PEMer workflow is a sensitive software for detecting SVs from paired-end sequence reads. SV-Simulation randomly introduces SVs into a given genome and generates simulated paired-end reads from novel genome.
Proper citation: PEMer (RRID:SCR_005263) Copy
http://en.wikipedia.org/wiki/Gene_Wiki
The Gene Wiki is a project that facilitates transferring information on human genes to Wikipedia article stubs with the goal of promoting collaboration and expansion of the articles. Number of gene articles The human genome contains an estimated 20,00025,000 protein-coding genes. The goal of the Gene Wiki project is to create seed articles for every notable human gene, that is, every gene whose function has been assigned in the peer-reviewed scientific literature. Approximately half of human genes have assigned function, therefore the total number of articles seeded by the Gene Wiki project would be expected to be in the range of 10,000 - 15,000. To date, approximately 10,271 articles have been created or augmented to include Gene Wiki project content. Expansion Once seed articles have been established, the hope and expectation is that these will be annotated and expanded by editors ranging in experience from the lay audience to students to professionals and academics. Proteins encoded by genes The majority of genes encode proteins hence understanding the function of a gene generally requires understanding of the function of the corresponding protein. In addition to including basic information about the gene, the project therefore also includes information about the protein encoded by the gene. Stubs for the Gene Wiki project are created by a bot and contain links to the following primary gene/protein databases * HUGO Gene Nomenclature Committee official gene name * Entrez Gene database * OMIM (Mendelian Inheritance in Man) database that catalogues all the known diseases with a genetic component * Amigo Gene Ontology * HomoloGene gene homologs in other species * SymAtlasRNA gene expression pattern in tissues * Protein Data Bank 3D structure of protein encoded by the gene * Uniprot (universal protein resource) a central repository of protein data
Proper citation: Gene Wiki (RRID:SCR_005317) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on May 3rd,2023. A software program designed to accurately map sequence data obtained from next-generation sequencing machines (specifically that of Solexa/Illumina) back to a genome of any size. By using the posterior probability of mapping a given read to a specific genomic loation, we are able to account for repetitive reads by distributing them across several regions in the genome. In addition, the output of the program is created in such a way that it can be easily viewed through other free and readily- available programs. Several benchmark data sets were created with spiked-in duplicate regions, and GNUMAP was able to more accurately account for these duplicate regions.
Proper citation: GNUMAP (RRID:SCR_005482) Copy
http://sourceforge.net/projects/cushaw2/files/CUSHAW2-GPU/
Software program (based on CUSHAW2) designed and optimized for Kepler-based GPUs, but still workable on earlier-generation Fermi-based ones.
Proper citation: CUSHAW2-GPU (RRID:SCR_005480) Copy
http://mbgd.genome.ad.jp/CGAT/
A comparative genome analysis tool for detailed comparison of closely related bacterial-sized genomes. It visualizes precomputed pairwise genome alignments on both dotplot and alignment viewers. Users can add information on this alignment, such as existence of tandem repeats or interspersed repetitive sequences and changes in codon usage bias, to facilitate interpretation of the observed genomic changes. Besides visualization functionalities, it also provides a general framework to process genome-scale alignments using various existing alignment programs. CGAT employs a client-server architecture, which consists of AlignmentViewer (client; a Java application) and DataServer (a set of Perl scripts). The DataServer package contains data construction scripts and CGI scripts and the AlignmentViewer program visualizes the alignment data obtained from the server thorough the HTTP protocol.
Proper citation: CGAT (RRID:SCR_005550) Copy
http://cushaw2.sourceforge.net/homepage.htm#latest
Software package for next-generation sequencing read alignment that is fast and parallel gapped read alignment to large genomes, such as the human genome.
Proper citation: CUSHAW (RRID:SCR_005479) Copy
http://manatee.sourceforge.net/
Manatee is a web-based gene evaluation and genome annotation tool; Manatee can store and view annotation for prokaryotic and eukaryotic genomes. The Manatee interface allows biologists to quickly identify genes and make high quality functional assignments, such as GO classifications, using search data, paralogous families, and annotation suggestions generated from automated analysis. Manatee can be downloaded and installed to run under the CGI area of a web server, such as Apache. Platform: Online tool, Linux compatible, Solaris
Proper citation: Manatee (RRID:SCR_005685) Copy
A web-based software package for comparative genomics.
Proper citation: Sybil (RRID:SCR_005593) Copy
http://www.sanger.ac.uk/resources/software/lookseq/
A web-based application for alignment visualization, browsing and analysis of genome sequence data.
Proper citation: LookSeq (RRID:SCR_005625) Copy
https://www.hgsc.bcm.edu/software/mercury
An automated, flexible, and extensible analysis workflow that provides accurate and reproducible genomic results at scales ranging from individuals to large cohorts. The analysis pipeline is deployed in local hardware and the Amazon Web Services cloud via the DNAnexus platform.
Proper citation: Mercury (RRID:SCR_004231) Copy
A promoter database of Saccharomyces cerevisiae. Users can explore the promoter regions of ~6000 genes and ORFs in yeast genome, annotate putative regulatory sites of all genes and ORFs, locate intergenic regions, and retrieve sequence of the promoter region. In regards to regulatory elements and transcription factors, users can provide information on transcriptionally related genes, browse matrix and consensus sequences, view the correlation between elements, observe binding affinity and expression, and look at genomewise distribution. SCPD also provides some simple but useful tools for promoter sequence analysis. Gene, consensus and matrix records may be submitted.
Proper citation: SCPD - Saccharomyces cerevisiae promoter database (RRID:SCR_004412) Copy
A collaborative ontology for the definition of sequence features used in biological sequence annotation. SO was initially developed by the Gene Ontology Consortium. Contributors to SO include the GMOD community, model organism database groups such as WormBase, FlyBase, Mouse Genome Informatics group, and institutes such as the Sanger Institute and the EBI. Input to SO is welcomed from the sequence annotation community. The OBO revision is available here: http://sourceforge.net/p/song/svn/HEAD/tree/ SO includes different kinds of features which can be located on the sequence. Biological features are those which are defined by their disposition to be involved in a biological process. Biomaterial features are those which are intended for use in an experiment such as aptamer and PCR_product. There are also experimental features which are the result of an experiment. SO also provides a rich set of attributes to describe these features such as polycistronic and maternally imprinted. The Sequence Ontologies use the OBO flat file format specification version 1.2, developed by the Gene Ontology Consortium. The ontology is also available in OWL from Open Biomedical Ontologies. This is updated nightly and may be slightly out of sync with the current obo file. An OWL version of the ontology is also available. The resolvable URI for the current version of SO is http://purl.obolibrary.org/obo/so.owl.
Proper citation: SO (RRID:SCR_004374) Copy
http://bix.ucsd.edu/projects/singlecell/
Software package for short read data from single cells that improves assembly through use of progressively increasing coverage cutoff. Used for single cell Illumina sequences, allows variable coverage datasets to be utilized with assembly of E. coli and S. aureus single cell reads. Assembles single cell genome of uncultivated SAR324 clade of Deltaproteobacteria.
Proper citation: Velvet-SC (RRID:SCR_004377) Copy
http://bioinfo.unl.edu/raiphy.php
A semi-supervised metagenomic fragment classification software program that utilizes the genome signatures to characterize the DNA sequences and taxonomic classification is based on an information theoretic measure referred as Relative Abundance Index (RAI). A DNA sequence of unknown source is classified and taxonomically labeled based on the phylogenetic profiles of the previously sequenced genomes. The profiles are iteratively updated using the unknown DNA sequences and the classification results. After a few cycles, the metagenome is classified into operational taxonomic units.
Proper citation: RAIphy (RRID:SCR_004720) Copy
Database of genetic and molecular biology data for the model higher plant Arabidopsis thaliana. Data available includes the complete genome sequence along with gene structure, gene product information, metabolism, gene expression, DNA and seed stocks, genome maps, genetic and physical markers, publications, and information about the Arabidopsis research community. Gene product function data is updated every two weeks from the latest published research literature and community data submissions. Gene structures are updated 1-2 times per year using computational and manual methods as well as community submissions of new and updated genes. TAIR also provides extensive linkouts from data pages to other Arabidopsis resources. The data can be searched, viewed and analyzed. Datasets can also be downloaded. Pages on news, job postings, conference announcements, Arabidopsis lab protocols, and useful links are provided.
Proper citation: TAIR (RRID:SCR_004618) Copy
http://www.genedb.org/Homepage/Lmajor
Database of the most recent sequence updates and annotations for the L. major genome. New annotations are constantly being added to keep up with published manuscripts and feedback from the Trypanosomatid research community. You may search by Protein Length, Molecular Mass, Gene Type, Date, Location, Protein Targeting, Transmembrane Helices, Product, GO, EC, Pfam ID, Curation and Comments, and Dbxrefs. BLAST and other tools are available. Leishmania species cause a spectrum of human diseases in tropical and subtropical regions of the world. We have sequenced the 36 chromosomes of the 32.8-megabase haploid genome of Leishmania major (Friedlin strain) and predict 911 RNA genes, 39 pseudogenes, and 8272 protein-coding genes, of which 36% can be ascribed a putative function. These include genes involved in host-pathogen interactions, such as proteolytic enzymes, and extensive machinery for synthesis of complex surface glycoconjugates. The Pathogen Genomics group at the Wellcome Trust Sanger Institute played a major role in sequencing the genome of Leishmania major (see Ivens et al.) Details of the centres involved and which chromosomes they sequenced, are given. The sequence data were obtained by adopting several parallel approaches, including complete cosmid sequencing, whole chromosome shotguns and/or BAC sequencing/skimming. The Leishmania parasite is an intracellular pathogen of the immune system targeting macrophages and dendritic cells. The disease Leishmaniasis affects the populations of 88 counties worldwide with symptoms ranging from disfiguring cutaneous and muco-cutaneous lesions that can cause widespread destruction of mucous membranes to visceral disease affecting the haemopoetic organs. In collaboration with GeneDB, the EuPathDB genomic sequence data and annotations are regularly deposited on TriTrypDB where they can be integrated with other datasets and queried using customized queries.
Proper citation: GeneDB Lmajor (RRID:SCR_004613) Copy
http://www.ncbi.nlm.nih.gov/nucest
Nucleotide database as collection of sequences from several sources, including GenBank, RefSeq, TPA and PDB. Genome, gene and transcript sequence data provide the foundation for biomedical research and discovery.
Proper citation: Nucleotide database (RRID:SCR_004630) Copy
https://computation-rnd.llnl.gov/lmat/
Open-source software tool to assign taxonomic labels to as many reads as possible in very large metagenomic datasets and report the taxonomic profile of the input sample. The quick "single pass" analysis of every read allows read binning to support additional more computationally expensive analysis such as metagenomic assembly or sensitive database searches on targeted subsets of reads.
Proper citation: LMAT (RRID:SCR_004646) Copy
Service providing functional analysis of proteins by classifying them into families and predicting domains and important sites. They combine protein signatures from a number of member databases into a single searchable resource, capitalizing on their individual strengths to produce a powerful integrated database and diagnostic tool. This integrated database of predictive protein signatures is used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures. You can access the data programmatically, via Web Services. The member databases use a number of approaches: # ProDom: provider of sequence-clusters built from UniProtKB using PSI-BLAST. # PROSITE patterns: provider of simple regular expressions. # PROSITE and HAMAP profiles: provide sequence matrices. # PRINTS provider of fingerprints, which are groups of aligned, un-weighted Position Specific Sequence Matrices (PSSMs). # PANTHER, PIRSF, Pfam, SMART, TIGRFAMs, Gene3D and SUPERFAMILY: are providers of hidden Markov models (HMMs). Your contributions are welcome. You are encouraged to use the ''''Add your annotation'''' button on InterPro entry pages to suggest updated or improved annotation for individual InterPro entries.
Proper citation: InterPro (RRID:SCR_006695) Copy
Database of peer-reviewed, continually updated annotation for the Pseudomonas aeruginosa PAO1 reference strain genome expanded to include all Pseudomonas species to facilitate cross-strain and cross-species genome comparisons with high quality comparative genomics. The database contains robust assessment of orthologs, a novel ortholog clustering method, and incorporates five views of the data at the sequence and annotation levels (Gbrowse, Mauve and custom views) to facilitate genome comparisons. Other features include more accurate protein subcellular localization predictions and a user-friendly, Boolean searchable log file of updates for the reference strain PAO1. The current annotation is updated using recent research literature and peer-reviewed submissions by a worldwide community of PseudoCAP (Pseudomonas aeruginosa Community Annotation Project) participating researchers. If you are interested in participating, you are invited to get involved. Many annotations, DNA sequences, Orthologs, Intergenic DNA, and Protein sequences are available for download.
Proper citation: Pseudomonas Genome Database (RRID:SCR_006590) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.