Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://www.ebi.ac.uk/webservices/whatizit/info.jsf
A text processing system that allows you to do textmining tasks on text. It is great at identifying molecular biology terms and linking them to publicly available databases. Whatizit is also a Medline abstracts retrieval/search engine. Instead of providing the text by Copy&Paste, you can launch a Medline search. The abstracts that match your search criteria are retrieved and processed by a pipeline of your choice. Whatizit is also available as 1) a webservice and as 2) a streamed servlet. The webservice allows you to enrich content within your website in a similar way as in the wikipedia. The streamed servlet allows you to process large amounts of text.
Proper citation: Whatizit (RRID:SCR_005824) Copy
A semantically annotated corpus of 240 MEDLINE abstracts (167 on the subject of E. coli species and 73 on the subject of the Human species) intended for training information extraction (IE) systems and/or resources which are used to extract events from biomedical literature. The corpus has been manually annotated with events relating to gene regulation by biologists. Each event is centered on either a verb (e.g. transcribe) or nominalized verb (e.g. transcription) and annotation consists of identifying, as exhaustively as possible, the structurally-related arguments of the verb or nominalized verb within the same sentence. Each event argument is then assigned the following information: * A semantic role from a fixed set of 13 roles which are tailored to the biomedical domain. * A biomedical concept type (where appropriate). The corpus in available for download in 2 formats: * A standoff format, based on the BioNLP'09 Shared Task format * An XML format, based on the GENIA event annotation format
Proper citation: GREC Corpus (RRID:SCR_006719) Copy
http://www.nactem.ac.uk/genia/
Resources and tools from a project to automatically extract useful information from texts written by scientists to help overcome the problems caused by information overload. The primary annotated resource created is the GENIA corpus, a collection of biomedical literature which consists of multiple layers of annotation, encompassing both syntactic and semantic annotation. The project also created or coordinated the annotation of multiple other corpus resources. Additionally, a rich set of automatic tools are available for various annotation tasks, most trained on various parts of the GENIA corpus annotations. The GENIA corpus was developed to provide a reference material for the development of bio-TM systems. The corpus currently contains 1,999 Medline abstracts which were collected using the three MeSH terms, human, blood cells, and transcription factors. The corpus has been annotated with various levels of linguistic and semantic information. The GENIA corpus includes the following: * POS annotation * Treebank * Coreference Annotation * Term annotation * Event annotation * Relation annotation * Cellular localization * Disease-Gene association * Pathway corpus The GENIA Project initiated the BioNLP Shared Task series and has organized a number of tasks in three different shared task events, many using resources based on GENIA Corpus annotations. Tools include: * XConc suite: a collection of XML-based tools which are integrated to support the corpus development and annotation.
Proper citation: GENIA Project: Mining literature for knowledge in molecular biology (RRID:SCR_007990) Copy
http://www.ebi.ac.uk/Rebholz-srv/ebimed/
A web application that combines Information Retrieval and Extraction from Medline. EBIMed finds Medline abstracts in the same way PubMed does. Then it goes a step beyond and analyses them to offer a complete overview on associations between UniProt protein/gene names, GO annotations, Drugs and Species. The results are shown in a table that displays all the associations and links to the sentences that support them and to the original abstracts. By selecting relevant sentences and highlighting the biomedical terminology EBIMed enhances your ability to acquire knowledge, relate facts, discover implications and, overall, have a good overview economizing the effort in reading.
Proper citation: EBIMed (RRID:SCR_005314) Copy
http://www.chibi.ubc.ca/WhiteText/
Freely available corpus of manually annotated brain region mentions created to facilitate text mining of neuroscience literature. The corpus contains 1,377 abstracts with 18,242 brain region annotations. Interannotator agreement was evaluated for a subset of the documents, and was 90.7% and 96.7% for strict and lenient matching respectively. We observed a large vocabulary of over 6,000 unique brain region terms and 17,000 words. For automatic extraction of brain region mentions we evaluated simple dictionary methods and complex natural language processing techniques. The dictionary methods based on neuroanatomical lexicons recalled 36% of the mentions with 57% precision. The best performance was achieved using a conditional random field (CRF) with a rich feature set. Features were based on morphological, lexical, syntactic and contextual information. The CRF recalled 76% of mentions at 81% precision, by counting partial matches recall and precision increase to 86% and 92% respectively. We suspect a large amount of error is due to coordinating conjunctions, previously unseen words and brain regions of less commonly studied organisms. We found context windows, lemmatization and abbreviation expansion to be the most informative techniques. We encourage you to test new methods and applications of the dataset. Please contact us if you do, we would like to hear about and link to your work. The abstracts are from PubMed/Medline, specifically The Journal of Comparative Neurology.
Proper citation: Automated recognition of brain region mentions in neuroscience literature. (RRID:SCR_002731) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within NIF that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.