Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://www.sci.unisannio.it/docenti/rampone/
Data set of Homo Sapiens Exons, Introns and Splice regions extracted from GenBank Rel.123 with an aim of giving standardized material to train and to assess the prediction accuracy of computational approaches for gene identification and characterization. From the complete GenBank (Primate Sequences Division) Rel.123 (162,557 entries), entries of Human Nuclear DNA including Complete CDS and more than one Exon have been selected, and 4523 exons and 3802 introns have been extracted from these entries. Details about extracted exons and introns are reported (Locus, number, Start and End position in the entry, sequence, length, G+C content, presence of not AGCT data (nucleotide scan check)). Statistics are also reported (overall nucleotides, average G+C content, nucleotide scan check results, number of not GT starting / AG ending introns, minimum / maximum / average length, length standard deviation). 3799+3799 donor and acceptor sites, as windows of 140 nucleotides around each splice site have been extracted. After discarding sequences not including canonical GTAG junctions (65+74), including insufficient data (not enough material for a 140 nucleotide window) (686+589), including not AGCT bases (29+30), and redundant (218+226) there are 2796+ 2880 windows. Finally, there are 271,937 + 332,296 windows of false splice sites, selected by searching canonical GTAG pairs in not splicing positions. The false sites in a range of +/- 60 from a true splice site are marked as proximal.
Proper citation: HS3D - Homo Sapiens Splice Sites Dataset (RRID:SCR_002939) Copy
http://rp-www.cs.usyd.edu.au/~yangpy/software/MFGE.html
A hybrid software system for feature selection and sample classification of high-dimensional datasets. It is designed for microarray but can be applied to any other high-dimensional datasets. It uses multiple filters to produce a normalized score for each feature. The score is an indication of the usefulness of each feature. It is then translated into a frequency map with more useful features receive a higher frequency in the map.
Proper citation: MF-GE (RRID:SCR_003509) Copy
Curated lists of genes associated to speech / language phenotypes and structural or functional abnormalities observed in patient populations. Entrez ID gene information, as well as gene expression profiles from the Allen Brain Atlas are available. You can also download expression data for a given gene in JSON or XML format.
Proper citation: Speech Language Disorders Database (RRID:SCR_003655) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 11, 2023. Archiving services, insertional site analysis, pharmacology and toxicology resources, and reagent repository for academic investigators and others conducting gene therapy research. Databases and educational resources are open to everyone. Other services are limited to gene therapy investigators working in academic or other non-profit organizations. Stores reserve or back-up clinical grade vector and master cell banks. Maintains samples from any gene therapy related Pharmacology or Toxicology study that has been submitted to FDA by U.S. academic investigator that require storage under Good Laboratory Practices. For certain gene therapy clinical trials, FDA has required post-trial monitoring of patients, evaluating clinical samples for evidence of clonal expansion of cells. To help academic investigators comply with this FDA recommendation, the NGVB offers assistance with clonal analysis using LAM-PCR and LM-PCR technology.
Proper citation: National Gene Vector Biorepository (RRID:SCR_004760) Copy
http://www.linked-neuron-data.org/
Neuroscience data and knowledge from multiple scales and multiple data sources that has been extracted, linked, and organized to support comprehensive understanding of the brain. The core is the CAS Brain Knowledge base, a very large scale brain knowledge base based on automatic knowledge extraction and integration from various data and knowledge sources. The LND platform provides services for neuron data and knowledge extraction, representation, integration, visualization, semantic search and reasoning over the linked neuron data. Currently, LND extracts and integrates semantic data and knowledge from the following resources: PubMed, INCF-CUMBO, Allen Reference Atlas, NIF, NeuroLex, MeSH, DBPedia/Wikipedia, etc.
Proper citation: Linked Neuron Data (RRID:SCR_003658) Copy
http://bc02.iis.sinica.edu.tw/gobu/manual/index.html
Gene Ontology Browsing Utility (GOBU) (GOBU) is a Java-based software program for integrating biological annotation catalogs under an extendable software architecture. Users may interact with the Gene Ontology and user-defined hierarchy data of genes, and then use its plugins to (and not limited to) (1) browse the GO hierarchy with user defined data, (2) browse GO-oriented expression levels in the user data, (3) compute GO enrichment, and/or (4) customize data reporting. A set of classes and utility functions has been established so that a customized program can be made as a plugin or a command-line tool that programmically manipulate the Gene Ontology and specified user data. See the source code repository for examples. Reference Lin WD, Chen YC, Ho JM, Hsiao CD. GOBU: Toward an Integration Interface for Biological Objects. Journal of Information Science and Engineering. 2006 22(1):19-29. Platform: Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: Gene Ontology Browsing Utility (GOBU) (RRID:SCR_005662) Copy
http://www.sgn.cornell.edu/bulk/input.pl?modeunigene
Allows users to download Unigene or BAC information using a list of identifiers or complete datasets with FTP., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Sol Genomics Network - Bulk download (RRID:SCR_007161) Copy
http://linux1.softberry.com/spldb/SpliceDB.html
Database of canonical and non-canonical mammalian splice sites. The information about verified splice site sequences for canonical and non-canonical sites is presented with the supporting evidence. Weight matrices were built for the major splice groups, which can be incorporated into gene prediction programs.
Proper citation: SpliceDB (RRID:SCR_006262) Copy
http://aws.amazon.com/1000genomes/
A dataset containing the full genomic sequence of 1,700 individuals, freely available for research use. The 1000 Genomes Project is an international research effort coordinated by a consortium of 75 companies and organizations to establish the most detailed catalogue of human genetic variation. The project has grown to 200 terabytes of genomic data including DNA sequenced from more than 1,700 individuals that researchers can now access on AWS for use in disease research free of charge. The dataset containing the full genomic sequence of 1,700 individuals is now available to all via Amazon S3. The data can be found at: http://s3.amazonaws.com/1000genomes The 1000 Genomes Project aims to include the genomes of more than 2,662 individuals from 26 populations around the world, and the NIH will continue to add the remaining genome samples to the data collection this year. Public Data Sets on AWS provide a centralized repository of public data hosted on Amazon Simple Storage Service (Amazon S3). The data can be seamlessly accessed from AWS services such Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic MapReduce (Amazon EMR), which provide organizations with the highly scalable compute resources needed to take advantage of these large data collections. AWS is storing the public data sets at no charge to the community. Researchers pay only for the additional AWS resources they need for further processing or analysis of the data. All 200 TB of the latest 1000 Genomes Project data is available in a publicly available Amazon S3 bucket. You can access the data via simple HTTP requests, or take advantage of the AWS SDKs in languages such as Ruby, Java, Python, .NET and PHP. Researchers can use the Amazon EC2 utility computing service to dive into this data without the usual capital investment required to work with data at this scale. AWS also provides a number of orchestration and automation services to help teams make their research available to others to remix and reuse. Making the data available via a bucket in Amazon S3 also means that customers can crunch the information using Hadoop via Amazon Elastic MapReduce, and take advantage of the growing collection of tools for running bioinformatics job flows, such as CloudBurst and Crossbow.
Proper citation: 1000 Genomes Project and AWS (RRID:SCR_008801) Copy
https://cran.r-project.org/web/packages/babelgene/index.html
Software R package to convert between human and non-human gene orthologs/homologs. Integrates orthology assertion predictions sourced from multiple databases as compiled by the HGNC Comparison of Orthology Predictions (HCOP) (Wright et al. 2005 , Eyre et al. 2007 , Seal et al. 2011 ).
Proper citation: babelgene (RRID:SCR_027117) Copy
http://www-personal.umich.edu/~jianghui/rseqdiff/
An R package that can detect differential gene and isoform expressions from RNA-seq data of multiple biological conditions. The approach considers three cases for each gene: 1) no differential expression, 2) differential expression without differential splicing and 3) differential splicing.
Proper citation: rSeqDiff (RRID:SCR_001683) Copy
https://cran.r-project.org/src/contrib/Archive/QuasiSeq/
Software package to apply the QL, QLShrink and QLSpline methods to quasi-Poisson or quasi-negative binomial models for identifying differentially expressed genes in RNA-seq data.
Proper citation: QuasiSeq (RRID:SCR_001715) Copy
http://datahub.io/dataset/kupkb
A collection of omics datasets (mRNA, proteins and miRNA) that have been extracted from PubMed and other related renal databases, all related to kidney physiology and pathology giving KUP biologists the means to ask queries across many resources in order to aggregate knowledge that is necessary for answering biological questions. Some microarray raw datasets have also been downloaded from the Gene Expression Omnibus and analyzed by the open-source software GeneArmada. The Semantic Web technologies, together with the background knowledge from the domain's ontologies, allows both rapid conversion and integration of this knowledge base. SPARQL endpoint http://sparql.kupkb.org/sparql The KUPKB Network Explorer will help you visualize the relationships among molecules stored in the KUPKB. A simple spreadsheet template is available for users to submit data to the KUPKB. It aims to capture a minimal amount of information about the experiment and the observations made.
Proper citation: Kidney and Urinary Pathway Knowledge Base (RRID:SCR_001746) Copy
http://gmod.org/wiki/Main_Page
A collection of open source software tools for creating and managing genome-scale biological databases. GMOD is made up databases, applications, and adaptor software that connects these components together. You can use it to create a small laboratory database of genome annotations, or a large web-accessible community database. At first GMOD just featured model organisms but now any organism with any kind of sequence associated with it is a good candidate as a subject for a GMOD database. There are GMOD databases with just protein sequence in them, with EST sequence only, those that are concerned primarily with gene expression, and even those dedicated to collections of RNA sequence. They have also heard of GMOD databases for oligonucleotides and plasmids.
Proper citation: Generic Model Organism Database Project (RRID:SCR_001731) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 23,2022. Time-series data sets spanning twelve time-points between E12-P9 for exploring cerebellar development of the mouse in time and space. The database contains a number of mutant / wildtype microarray datasets including two complete wildtype microarray time-series (C57BL/6 and DBA/2J). The dataset also includes in situ hybridization and bioinformatic analyses. Exploration of this dataset will allow the investigator to assess differential gene expression profiles from a developing mutant cerebella, to assess the temporal changes in gene expression in the wildtype, and to verify the cellular expression of these genes in images from our in situ hybridization library. Using the database, the investigator can explore the developmental expression or differential expression patterns of a particular gene, or create lists of similarly expression genes by building simple search algorithms. These lists can then be mined across all the datasets in both space and time. Cb GRiTS's current datasets represent gene expression analyses from multiple cerebellar mutant and wildtype single time-point and developmental series.
Proper citation: Cerebellar Gene Regulation in Time and Space Database (RRID:SCR_001699) Copy
The Physiome Project is a worldwide public domain effort to provide a computational framework for understanding human and other eukaryotic physiology. It aims to develop integrative models at all levels of biological organization, from genes to the whole organism via gene regulatory networks, protein pathways, integrative cell function, and tissue and whole organ structure/function relations. Additionally, an important goal of the project is to develop applications for teaching physiology. Current projects include the development of: - ontologies to organize biological knowledge and access to databases - markup languages to encode models of biological structure and function in a standard format for sharing between different application programs and for re-use as components of more comprehensive models - databases of structure at the cell, tissue and organ levels - software to render computational models of cell function such as ion channel electrophysiology, cell signaling and metabolic pathways, transport, motility, the cell cycle, etc. in 2 & 3D graphical form - software for displaying and interacting with the organ models which will allow the user to move across all spatial scales Sponsors: This project is supported by the International Union of Physiological Sciences (IUPS), the IEEE Engineering. in Medicine and Biology (EMBS), and the International Federation for Medical and Biological Engineering (IFMBE)
Proper citation: International Union of Physiological Sciences: Physiome Project (RRID:SCR_001760) Copy
http://incf.org/about/programs/modeling/blue-gene-access
Through this site, INCF provides he neuroinformatics community with access to an IBM Blue Gene/L supercomputer. INCF owns a share of a BlueGene/L (BG/L) supercomputer located at the Parallel Computer Center (PDC) at The Royal Institute of Technology (KTH) in Stockholm. Allocations are now available through the INCF Secretariat. During an initial evaluation phase, a limited numbers of large-scale computing projects will be selected, based on the suitability of the project for supercomputing. Research groups with limited access to supercomputers at their home institutions are given priority. Approved projects are regularly re-evaluated. New projects are approved based on availability and usage load of the BG/L. The Blue Gene/L supercomputer project is aimed at expanding the horizon of high-performance computing to unprecedented levels of scale and performance. Blue Gene/L is the first supercomputer in the Blue Gene family. The full Blue Gene/L consists of 64 racks containing 65,536 high-performance compute nodes. Each node (nodes and chips are the same in the Blue Gene system) contains two embedded 32-bit PowerPC processors. Furthermore, the same chip that is used for compute nodes is also used for the 1,024 I/O nodes. A three-dimensional torus network and a collective network are used to interconnect all nodes. The full system contains 33 terabytes of main memory; it is designed to achieve 183.5 teraflops peak performance using one of the processors of each node for computation and the other processor for communication, and 367 teraflops using both processors for computation. Another key architectural feature of this supercomputer is the link chip component and five Blue Gene/L networks, the PowerPC 440 core and floating-point enhancements, the on-chip and off-chip distributed memory system, the node- and system-level design for high reliability, and the comprehensive approach to fault isolation. One of the key objectives in Blue Gene/L design is to achieve cost/performance comparable to the COTS (Commodity Off The Shelf) approach, while at the same time incorporating a processor and network combination so powerful that it revolutionizes the performance of supercomputer systems. Sponsors: This resource is supported by the INCF.
Proper citation: International Neuroinformatics Coordinating Facility: Blue Gene/L Access (RRID:SCR_001755) Copy
http://csg.sph.umich.edu//abecasis/MACH/index.html
A Markov Chain based software tool for haplotyping, genotype imputation and disease association analysis that can resolve long haplotypes or infer missing genotypes in samples of unrelated individuals.
Proper citation: MACH 1.0 (RRID:SCR_001759) Copy
http://www.ncbi.nlm.nih.gov/projects/homology/maps/
This page provides quick access to the Comparative mapping functions available in the Map Viewer. Currently, comparative maps are calculated using HomoloGene orthology predictions. Once the gene pairs have been established, blocks of conserved syteny can be established using the positions of each gene object in their respective builds. Sponsors: This resource is supported by NCBI.
Proper citation: Homology Maps Page (RRID:SCR_001666) Copy
Data analysis service that searches PubMed literature database (abstracts) about specific relationships between proteins, genes, or keywords using a NLP-based text-mining approach. The results are returned as a graph. The synonym database used in Chilibot is available, without fee, for academic use only. Several different search methods are supported including: * searching for relationship between two genes, proteins or keywords * searching for relationships between many genes, proteins, or keywords * searching for relationships between two lists of genes, proteins, or keywords Advanced options include: * Automated hypothesis generation (graph) * Restricting context using keywords * Providing your own synonyms * Modifying synonyms provided by Chilibot * Color coding nodes with gene expression values * Special search: modulation
Proper citation: Chilibot: Gene and Protein relationships from MEDLINE (RRID:SCR_001705) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.