Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://rgd.mcw.edu/rgdCuration/?module=portal&func=show&name=nuro
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on May 12,2023. Portal that provides researchers with easy access to data on rat genes, QTLs, strain models, biological processes and pathways related to neurological diseases. This resource also includes dynamic data analysis tools.
Proper citation: Rat Genome Database: Neurological Disease Portal (RRID:SCR_008685) Copy
A robust, secure, medical-grade, web application that lives in the cloud and has the ability to analyze and annotate entire human genomes in a rapid and cost-effective way.
Proper citation: Tute Genomics (RRID:SCR_008672) Copy
http://roadmapepigenomics.org/
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on July 11, 2022. Project for human epigenomic data from experimental pipelines built around next-generation sequencing technologies to map DNA methylation, histone modifications, chromatin accessibility and small RNA transcripts in stem cells and primary ex vivo tissues selected to represent normal counterparts of tissues and organ systems frequently involved in human disease. Consortium expects to deliver collection of normal epigenomes that will provide framework or reference for comparison and integration within broad array of future studies. Consortium is also committed to development, standardization and dissemination of protocols, reagents and analytical tools to enable research community to utilize, integrate and expand upon this body of data.
Proper citation: Roadmap Epigenomics Project (RRID:SCR_008924) Copy
http://research-pub.gene.com/gmap/
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 29, 2016. A software program for mapping and aligning cDNA sequences to a genome. The program maps and aligns a single sequence with minimal startup time and memory requirements, and provides fast batch processing of large sequence sets. The program generates accurate gene structures, even in the presence of substantial polymorphisms and sequence errors, without using probabilistic splice site models. Methodology underlying the program includes a minimal sampling strategy for genomic mapping, oligomer chaining for approximate alignment, sandwich DP for splice site detection, and microexon identification with statistical significance testing.
Proper citation: GMAP (RRID:SCR_008992) Copy
https://www.mtocdb.org/?next=/browse/results/
A database of over 300 Electron Microscopy (EM) images of centrioles and centriole related structures from almost 60 species, described by a controlled vocabulary allowing detailed description of the observed structures. This knowledge is supplemented by a manually curated list of proteins known to be involved in centriole assembly, their (putative) orthologs, and localization information. mtocDB aims to characterize the naturally occurring morphological variation observed in centrioles and centriole associated structure alongside molecular information on the proteins involved in their assembly. Examining these in an evolutionary context will allow the cell biology community to infer meaningful relationships between cellular assembly mechanisms and the structures they form. This community resource for cell biologists interested in the the evolution of centrioles and centriole related structures aims to bridge the gap between structural morphology and molecular function by examining naturally occurring structural variation in a phylogenomic context. Centrioles are cylindrical microtubule arrays required for stability and duplication of the centrosome in animal cells, and for the assembly of cilia and flagella in many eukaryotes. The presence of centrioles throughout most eukaryotic branches suggests that this structure was present in the last eukaryotic common ancestor. Although centrioles show a typically well conserved structure, they can perform several functions and display a diversity of accessory structures. However, this diversity is not properly classified beyond model organisms, and the information contained in decades of electronic microscopy of other organisms remains untapped.
Proper citation: mtocDB (RRID:SCR_008933) Copy
http://gwas.biosciencedbc.jp/cgi-bin/hvdb/hv_top.cgi
A repository database to achieve continuous and intensive management of GWAS data and variation data identified by next generation sequencing (NGS) and data-sharing among researchers. In this database, variations including short/long insertions / deletions and structural variations related to disease susceptibility, virus resistance, and drug response are registered along with statistical genetic results and simple clinical characteristics to clarify the locus specific characteristics. Currently this database contains information extracted from scientific papers and next generation sequencing results and other small scale experimental results of several research laboratories. Mutation data submission is greatly appreciated.
Proper citation: Human Variation DB (RRID:SCR_009014) Copy
http://www.evocontology.org/site/Main/EvocOntologyDotOrg
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 10, 2017. A pilot effort that has developed a centralized, web-based biospecimen locator that presents biospecimens collected and stored at participating Arizona hospitals and biospecimen banks, which are available for acquisition and use by researchers. Researchers may use this site to browse, search and request biospecimens to use in qualified studies. The development of the ABL was guided by the Arizona Biospecimen Consortium (ABC), a consortium of hospitals and medical centers in the Phoenix area, and is now being piloted by this Consortium under the direction of ABRC. You may browse by type (cells, fluid, molecular, tissue) or disease. Common data elements decided by the ABC Standards Committee, based on data elements on the National Cancer Institute''s (NCI''s) Common Biorepository Model (CBM), are displayed. These describe the minimum set of data elements that the NCI determined were most important for a researcher to see about a biospecimen. The ABL currently does not display information on whether or not clinical data is available to accompany the biospecimens. However, a requester has the ability to solicit clinical data in the request. Once a request is approved, the biospecimen provider will contact the requester to discuss the request (and the requester''s questions) before finalizing the invoice and shipment. The ABL is available to the public to browse. In order to request biospecimens from the ABL, the researcher will be required to submit the requested required information. Upon submission of the information, shipment of the requested biospecimen(s) will be dependent on the scientific and institutional review approval. Account required. Registration is open to everyone., documented September 6, 2016. Set of orthogonal controlled vocabularies that unifies gene expression data by facilitating a link between the genome sequence and expression phenotype information. The system associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. The four core eVOC ontologies provide an appropriate set of detailed human terms that describe the sample source of human experimental material such as cDNA and SAGE libraries. These expression terms are linked to libraries and transcripts allowing the assessment of tissue expression profiles, differential gene expression levels and the physical distribution of expression across the genome. Analysis is currently possible using EST and SAGE data, with microarray data being incorporated. The eVOC data is increasingly being accepted as a standard for describing gene expression and eVOC ontologies are integrated with the Ensembl EnsMart database, the Alternate Transcript Diversity Project and the UniProt Knowledgebase. Several groups are currently working to provide shared development of this resource such that it is of maximum use in unifying transcript expression information.
Proper citation: eVOC (RRID:SCR_010704) Copy
http://www.broad.mit.edu/node/305
The Connectivity Map aims to generate a detailed map that links gene patterns associated with disease to corresponding patterns produced by drug candidates and a variety of genetic manipulations. The Connectivity Map is the most comprehensive effort yet for using genomics in a drug-discovery framework. It allows researchers to screen compounds against genome-wide disease signatures, rather than a pre-selected set of target genes. Drugs are paired with diseases using sophisticated pattern-matching methods with a high level of resolution and specificity. To build a Connectivity Map, the Broad Institute brings together molecular biologists, genomics specialists, computational scientists, pharmacologists, chemists and chemical biologists, as well as expertise from across the breadth and depth of medicine.Connectivity map is a large public database of signatures of drugs and genes, and pattern-matching tools to detect similarities among these signatures.The parent site for the Broad Institute at MIT has a software library of software applications developed for use in genetic analysis.
Proper citation: National Institute of Mental Health (NIMH) Human Genetics Initiative (RRID:SCR_007436) Copy
http://www.geisha.arizona.edu/geisha/
Online repository for chicken in situ hybridization information. This site presents whole mount in situ hybridization images and corresponding probe and genomic information for genes expressed in chicken embryos in Hamburger Hamilton stages 1-25 (0.5-5 days). The GEISHA project began in 1998 to investigate using high throughput whole mount in situ hybridization to identify novel, differentially expressed genes in chicken embryos. An initial expression screen of approximately 900 genes demonstrated feasibility of the approach, and also highlighted the need for a centralized repository of in situ hybridization expression data. Objectives: The goals of the GEISHA project are to obtain whole mount in situ hybridization expression information for all differentially expressed genes in the chicken embryo between HH stages 1-25, to integrate expression data with the chicken genome browsers, and to offer this information through a user-friendly graphical user interface. In situ hybridization images are obtained from three sources: 1. In house high throughput in situ hybridization screening: cDNAs obtained from several embryonic cDNA libraries or from EST repositories are screened for expression using high throughput in situ hybridization approaches. 2. Literature curation: Agreements with journals permit posting of published in situ hybridization images and related information on the GEISHA site. 3. Unpublished in situ hybridization information from other laboratories: laboratories generally publish only a small fraction of their in situ hybridization data. High quality images for which probe identity can be verified are welcome additions to GEISHA.
Proper citation: GEISHA - Gallus Expression in Situ Hybridization Analysis: A Chicken Embryo Gene Expression Database (RRID:SCR_007440) Copy
Comprehensive catalogue of animal genome size data. Haploid DNA contents (C-values, in picograms) are available for 4972 species (3231 vertebrates and 1741 non-vertebrates) based on 6518 records from 669 published sources. Data may be submitted directly to the database or reprints and notifications of new papers may be sent to database curation staff.
Proper citation: Animal Genome Size Database (RRID:SCR_007551) Copy
http://gene3d.biochem.ucl.ac.uk/Gene3D/
A large database of CATH protein domain assignments for ENSEMBL genomes and Uniprot sequences. Gene3D is a resource of form studying proteins and the component domains. Gene3D takes CATH domains from Protein Databank (PDB) structures and assigns them to the millions of protein sequences with no PDB structures using Hidden Markov models. Assigning a CATH superfamily to a region of a protein sequence gives information on the gross 3D structure of that region of the protein. CATH superfamilies have a limited set of functions and so the domain assignment provides some functional insights. Furthermore most proteins have several different domains in a specific order, so looking for proteins with a similar domain organization provides further functional insights. Strict confidence cut-offs are used to ensure the reliability of the domain assignments. Gene3D imports functional information from sources such as UNIPROT, and KEGG. They also import experimental datasets on request to help researchers integrate there data with the corpus of the literature. The website allows users to view descriptions for both single proteins and genes and large protein sets, such as superfamilies or genomes. Subsets can then be selected for detailed investigation or associated functions and interactions can be used to expand explorations to new proteins. The Gene3D web services provide programmatic access to the CATH-Gene3D annotation resources and in-house software tools. These services include Gene3DScan for identifying structural domains within protein sequences, access to pre-calculated annotations for the major sequence databases, and linked functional annotation from UniProt, GO and KEGG., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Gene3D (RRID:SCR_007672) Copy
http://genomics.senescence.info/
Collection of databases and tools designed to help researchers study the genetics of human ageing using modern approaches such as functional genomics, network analyses, systems biology and evolutionary analyses. A major resource in HAGR is GenAge, which includes a curated database of genes related to human aging and a database of ageing- and longevity-associated genes in model organisms. Another major database in HAGR is AnAge. Featuring over 4,000 species, AnAge provides a compilation of data on aging, longevity, and life history that is ideal for the comparative biology of aging. GenDR is a database of genes associated with dietary restriction based on genetic manipulation experiments and gene expression profiling. Other projects include evolutionary studies, genome sequencing, cancer genomics, and gene expression analyses. The latter allowed them to identify a set of genes commonly altered during mammalian aging which represents a conserved molecular signature of aging. Software, namely in the form of scripts for Perl and SPSS, is made available for users to perform a variety of bioinformatic analyses potentially relevant for studying aging. The Perl toolkit, entitled the Ageing Research Computational Tools (ARCT), provides modules for parsing files, data-mining, searching and downloading data from the Internet, etc. Also available is an SPSS script that can be used to determine the demographic rate of aging for a given population. An extensive list of links regarding computational biology, genomics, gerontology, and comparative biology is also available.
Proper citation: Human Ageing Genomic Resources (RRID:SCR_007700) Copy
The Centre d''Etude du Polymorphisme Humain (CEPH) is a research laboratory, the main activities of which are the setting up, storage, processing and distribution of DNA collections for the identification of genetic factors conferring susceptibility to complex disorders. These collections are established in partnership and full collaboration with external French or international research groups. The Foundation currently hosts the CEPH reference panel, the HGDP panel (Human genome Diversity Cell Line Panel) and several collections amounting mid-2008 to more than 250 000 samples. The goal of CEPH is to understand complex multifactorial disorders necessitates the establishment of structures facilitating access to large and integrated collection of individuals, characterized by a large number of variables emanating from different technologies and platforms. To achieve this goal, CEPH facilitates the setting up of integrated analyses combining clinical, genetic and environmental data, for the identification of susceptibility factors to complex multifactorial disorders Additionally, CEHP allows the reception, storage, processing and distribution of biological sample collections. At the same time, it promotes and participates in the design and setting up of genetic studies: - in partnership and full collaboration with external research groups - giving access to a large number of variables - in a sufficient number of subjects - allowing large scale integrated analyses
Proper citation: Centre dEtude du Polymorphisme Humain (RRID:SCR_008026) Copy
http://www.dnaform.jp/products/cage_e.html
Expression profiling and promoter identification software tool for transcriptional network analysis and transcriptome characterization. DeepCAGE, the combination of next-generation sequencing with next generation expression profiling provides unsurpassed solutions for expression profiling and genome annotation. CAGE will be the experimental approach at need to link gene expression and control regions in the genome. With the availability of next-generation sequencing methods, DNAFORM now offers DeepCAGE services. DeepCAGE libraries are prepared for direct analysis by an Illumina/Solexa Sequencer. One sequencing run using one channel on an Illumina/Solexa Sequencer can yield in over 4,000,000 reads per sample. CAGE is based on our full-length cDNA library technology, where an adaptor is ligated to the 5''''-end of full-length cDNAs, which introduces a recognition site for a Class IIs restriction endonuclease adjacent to the 5''''-end of the cDNA. The Class IIs restriction endonuclease, here MmeI, allows for the cloning of short tags as derived from the 5''''-end of transcripts into concatemers for high-throughput sequencing. CAGE tags are further characterized by mapping to genomic sequences, which enables the identification of transcriptional start sites. As such CAGE can contribute to projects in Gene Discovery, Gene Expression, and Promoter Identification. After the genome sequencing projects have provided us with the genetic blueprints for many organisms, new questions have to be answered on how to correlate the observed genotypes with related phenotypes, and how to understand the regulation of genetic information in time and space. The dynamics of living systems and the functional behavior of cells in multicellular organisms has thus become the subject of the emerging field of system biology. Integration of experimental approaches and computer aided theories on a system level will be the fundamental principle to drive systems biology in order to understand the principles behind complex regulatory networks, which will be an ambitious goal requiring new approaches in life sciences. For ordering and additional information, please contact us under contact_at_dnaform.jp
Proper citation: CAGE (RRID:SCR_007574) Copy
Database of information about restriction enzymes and related proteins containing published and unpublished references, recognition and cleavage sites, isoschizomers, commercial availability, methylation sensitivity, crystal, genome, and sequence data. DNA methyltransferases, homing endonucleases, nicking enzymes, specificity subunits and control proteins are also included. Several tools are available including REBsites, BLAST against REBASE, NEBcutter and REBpredictor. Putative DNA methyltransferases and restriction enzymes, as predicted from analysis of genomic sequences, are also listed. REBASE is updated daily and is constantly expanding. Users may submit new enzyme and/or sequence information, recommend references, or send them corrections to existing data. The contents of REBASE may be browsed from the web and selected compilations can be downloaded by ftp (ftp.neb.com). Additionally, monthly updates can be requested via email.,
Proper citation: REBASE (RRID:SCR_007886) Copy
http://www.projects.roslin.ac.uk/cdiv/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 16, 2013. The objective of the project is the standardization of micro-satellite markers used within participating laboratories, use of DNA markers to define genetic diversity and to enable monitoring of breeds to promote conservation programs where required, and the determination of diversity present in rare and local breeds across Europe. The blood typing laboratories are now beginning to use micro-satellite markers as an alternative to serology for parentage verification, and are selecting a common set to be used from the several hundred micro-satellite markers available that cover the bovine genome, produced as part of the Bovine genome mapping project (See BovMaP). Work with micro-satellite markers has shown that they are valuable tools for examining genetic diversity and phylogeny in many species. However, for work carried out in different laboratories to be comparable, it is essential that the same markers are used. To maintain the compatibility of data generated by the various typing labs, it is essential that all laboratories adopt the same markers and typing protocols. It is therefore of paramount importance that the blood typing laboratories and research labs that are examining the genetic structure of the cattle populations adopt a common panel of the best micro-satellite markers available. Some pilot comparative work has been undertaken through the International Society for Animal Genetics, but so far this has only involved the blood typing laboratories. One objective of this project is to facilitate the comparison of the micro-satellite markers currently in use in the different types of laboratory and determine the efficiency of the markers available in revealing genetic differences within and among breeds. It will also be important to compare the use of markers in different laboratories to determine how robust they are and how easily results can be compared. From comparison of the markers, those that are most suitable will be selected to form a panel which will be recommended for pedigree validation and genetic surveys. Cattle are an important source of food in Europe, and intense selection has resulted in the development of specialized breeds. Selection for high-producing dairy cattle has been successful, but one associated drawback is that the cattle population, both in Europe and North America, has been skewed dramatically towards one breed, the Holstein/Friesian. So there has been a decline in the number of individuals of other breeds, and hence a general erosion of the genetic base of the cattle population. The progressive move towards the North American-type Holstein animals has also resulted in the requirement for high input/high output farming and intensive management schemes. The impact of this on the environment has been significant, e.g. pollution problems arising from the need for high nitrogen fertilizers to produce sufficient high quality fodder, and disposal problems associated with slurry waste. Poorer areas of the community have been unable to compete with such farming systems, and are more suited to low input/low output farming using traditional stock. It is however the future perspective that is of greatest concern. It is impossible to predict requirements for cattle production - quality, production type, management systems, etc. The ability to switch rapidly to alternative production will be dependent on the genetic base of the population available to selection programs. It is therefore essential to maintain the greatest genetic diversity possible in the cattle population. Whilst current farming practices are perceived to be both efficient and acceptable, the breeds less favored by commercial farmers will dwindle. It is therefore important that on an European scale efficient management of these breeds maintains the widest genetic base possible. This project aims to carry out a survey of the current genetic base of the European cattle population and to provide the tools to assist breeding programs to maintain a broad base. The blood typing laboratories are now beginning to use micro-satellite markers as an alternative to serology for parentage verification, and are selecting a common set to be used from the several hundred micro-satellite markers available that cover the bovine genome, produced as part of the Bovine genome mapping project. Early work to measure genetic diversity used blood groups to show differences between breeds and the diversity present. Unfortunately, the number of loci available are limited, with only the B system being sufficiently polymorphic to be really useful. However, since there is a wealth of information available from such typing, this information can be used to estimate changes in the genetic structure of cattle populations across Europe over the past twenty years. More recently mini-satellite probes have been used to generate ''genetic fingerprints'' which have been used to show differences between individuals. Such fingerprints have been used to estimate genetic diversity - the greater the number of bands revealed by the fingerprint being equated with greater diversity. This is valid within limits. The main disadvantage of the fingerprint approach is that the chromosomal location and number of loci being sampled, and so the proportion of the genome examined, is unknown. The allelic bands on the gel cannot be easily identified, so allele inheritance cannot be addressed making it impossible to trace ancestry. Through the EC funded BovMaP project, large numbers of highly polymorphic micro-satellite markers have become available, which are being mapped on the bovine genome. These markers are particularly suited to measuring genetic diversity, and markers can be selected to cover the entire genome.
Proper citation: CaDBase: Genetic diversity in cattle (RRID:SCR_008146) Copy
http://locus.jouy.inra.fr/cgi-bin/bovmap/intro.pl
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. Database containing information on the cattle genome comprising loci list, phenes list, homology query, cattle maps, gene list, and chromosome homology. The objective of BovMap is to develop a set of anchored loci for the cattle genome map. In total, 58 clones were hybridized with chromosomes and identified loci on 22 of the 31 different bovine chromosomes. Three clones contained satellite DNA. Two or more markers were placed on 12 chromosomes. Sequencing of the microsatellites and flanking regions was performed directly from 43 cosmids, as previously reported. Primers were developed for 39 markers and used to describe the polymorphism associated with the corresponding loci. Users are also allowed to summit their own data for Bovmap. An integrated cytogenetic and meiotic map of the bovine genome has also been developed around the Bovmap database. One objective that Bovmap uses as the mapping strategy for the bovine genome uses large insert clones as a tool for physical mapping and as a source of highly polymorphic microsatellites for genetic typing.
Proper citation: BovMap Database (RRID:SCR_008145) Copy
The E. coli Genome Project has the goal of completely sequencing the E. coli and human genomes. They began isolation of an overlapping lambda clonebank of E. coli K-12 strain MG1655. Those clones served as the starting material in our initial efforts to sequence the whole genome. Improvements in sequencing technology have since reached the point where whole-genome sequencing of microbial genomes is routine, and the human genome has in fact been completed. They initiated additional sequencing efforts, concentrating on pathogenic members of the family Enterobacteriaceae -- to which E. coli belongs. They also began a systematic functional characterization of E. coli K-12 genes and their regulation, using the whole genome sequence to address how the over 4000 genes of this organism act together to enable its survival in a wide range of environments.
Proper citation: E. coli Genome project (RRID:SCR_008139) Copy
MitoRes, is a comprehensive and reliable resource for massive extraction of sequences and sub-sequences of nuclear genes and encoded products targeting mitochondria in metazoa. It has been developed for supporting high-throughput in-silico analyses aimed to studies of functional genomics related to mitochondrial biogenesis, metabolism and to their pathological dysfunctions. It integrates information from the most accredited world-wide databases to bring together gene, transcript and encoded protein sequences associated to annotations on species name and taxonomic classification, gene name, functional product, organelle localization, protein tissue specificity, Enzyme Classification (EC), Gene Ontology (GO) classification and links to other related public databases. The section Cluster, has been dedicated to the collection of data on protein clustering of the entire catalogue of MitoRes protein sequences based on all versus all global pair-wise alignments for assessing putative intra- and inter-species functional relationships. The current version of MitoRes is based on the UniProt release 4 and contains 64 different metazoan species. The incredible explosion of knowledge production in Biology in the past two decades has created a critical need for bioinformatic instruments able to manage data and facilitate their retrieval and analysis. Hundreds of biological databases have been produced and the integration of biological data from these different resources is very important when we want to focus our efforts towards the study of a particular layer of biological knowledge. MitoRes is a completely rebuilt edition of MitoNuc database, which has been extensively modified to deal successfully with the challenges of the post genomic era. Its goal is to represent a comprehensive and reliable resource supporting high-quality in-silico analyses aimed to the functional characterization of gene, transcript and amino acid sequences, encoded by the nuclear genome and involved in mitochondrial biogenesis, metabolism and pathological dysfunctions in metazoa. The central features of MitoRes are: # an integrated catalogue of protein, transcript and gene sequences and sub-sequences # a Web-based application composed of a wide spectrum of search/retrieval facilities # a sequence export manager allowing massive extraction of bio-sequences (genes, introns, exons, gene flanking regions, transcripts, UTRs, CDS, proteins and signal peptides) in FASTA, EMBL and GenBank formats. It is an interconnected knowledge management system based on a MySQL relational database, which ensures data consistency and integrity, and on a Web Graphical User Interface (GUI), built in Seagull PHP Framework, offering a wide range of search and sequence extraction facilities. The database is compiled extracting and integrating information from public resources and data generated by the MitoRes team. The MitoRes database consists of comprehensive sequence entries whose core data are protein, transcript and gene sequences and taxonomic information describing the biological source of the protein. Additional information include: bio-sequences structure and location, biological function of protein product and dynamic links to both, external public databases used as data resources and public databases reporting complementary information. The core entity of the MitoRes database is represented by the protein so that each MitoRes entry is generated for each protein reported in the UniProt database as a nuclear encoded protein involved in mitochondrial biogenesis and function. Sponsors: MitoRes has been supported by Ministero Universit e Ricerca Scientifica, Italy (PRIN, Programma Biotecnologie legge 95/95-MURST 5, Proiect MURST Cluster C03/2000, CEGBA). Currently it is supported by operating grants from the Ministero dellIstruzione, dellUniversit e della Ricerca (MIUR), Italy (PNR 2001-2003 (FIRB art.8) D.M. 199, Strategic Program: Post-genome, grant 31-063933 and Project n.2, Cluster C03 L. 488/929).
Proper citation: MitoRes (RRID:SCR_008208) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 29, 2016. An algorithm that finds articles most relevant to a genetic sequence. In the genomic era, researchers often want to know more information about a biological sequence by retrieving its related articles. However, there is no available tool yet to achieve conveniently this goal. Here, a new literature-mining tool MedBlast is developed, which uses natural language processing techniques, to retrieve the related articles of a given sequence. An online server of this program is also provided. The genome sequencing projects generate such a large amount of data every day that many molecular biologists often encounter some sequences that they know nothing about. Literature is usually the principal resource of such information. It is relatively easy to mine the articles cited by the sequence annotation; however, it is a difficult task to retrieve those relevant articles without direct citation relationship. The related articles are those described in the given sequence (gene/protein), or its redundant sequences, or the close homologs in various species. They can be divided into two classes: direct references, which include those either cited by the sequence annotation or citing the sequence in its text; indirect references, those which contain gene symbols of the given sequence. A few additional issues make the task even more complicated: (1) symbols may have aliases; and (2) one sequence may have a couple of relatives that we want to take into account too, which include redundant (e.g. protein and gene sequences) and close homologs. Here the issues are addressed by the development of the software MedBlast, which can retrieve the related articles of the given sequence automatically. MedBlast uses BLAST to extend homology relationships, precompiled species-specific thesauruses, a useful semantics technique in natural language processing (NLP), to extend alias relationship, and EUtilities toolset to search and retrieve corresponding articles of each sequence from PubMed. MedBlast take a sequence in FASTA format as input. The program first uses BLAST to search the GenBank nucleic acid and protein non-redundant (nr) databases, to extend to those homologous and corresponding nucleic acid and protein sequences. Users can input the BLAST results directly, but it is recommended to input the result of both protein and nucleic acid nr databases. The hits with low e-values are chosen as the relatives because the low similarity hits often do not contain specific information. Very long sequences, e.g. 100k, which are usually genomic sequences, are discarded too, for they do not contain specific direct references. User can adjust these parameters to meet their own needs.
Proper citation: MedBlast (RRID:SCR_008202) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.