Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Cell-specific patterns of gene expression are determined by combinatorial actions of sequence-specific transcription factors at cis-regulatory elements. Studies indicate that relatively simple combinations of lineage-determining transcription factors (LDTFs) play dominant roles in the selection of enhancers that establish cell identities and functions. LDTFs require collaborative interactions with additional transcription factors to mediate enhancer function, but the identities of these factors are often unknown. We have shown that natural genetic variation between individuals has great utility for discovering collaborative transcription factors. Here, we introduce MMARGE (Motif Mutation Analysis of Regulatory Genomic Elements), the first publicly available suite of software tools that integrates genome-wide genetic variation with epigenetic data to identify collaborative transcription factor pairs. MMARGE is optimized to work with chromatin accessibility assays (such as ATAC-seq or DNase I hypersensitivity), as well as transcription factor binding data collected by ChIP-seq. Herein, we provide investigators with rationale for each step in the MMARGE pipeline and key differences for analysis of datasets with different experimental designs. We demonstrate the utility of MMARGE using mouse peritoneal macrophages, liver cells, and human lymphoblastoid cells. MMARGE provides a powerful tool to identify combinations of cell type-specific transcription factors while simultaneously interpreting functional effects of non-coding genetic variation.
Pubmed ID: 29893919
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
International functional genomics data collection generated from microarray or next-generation sequencing (NGS) platforms. Repository of functional genomics data supporting publications. Provides genes expression data for reuse to the research community where they can be queried and downloaded. Integrated with the Gene Expression Atlas and the sequence databases at the European Bioinformatics Institute. Contains a subset of curated and re-annotated Archive data which can be queried for individual gene expression under different biological conditions across experiments. Data collected to MIAME and MINSEQE standards. Data are submitted by users or are imported directly from the NCBI Gene Expression Omnibus.
View all literature mentionsA dataset containing the full genomic sequence of 1,700 individuals, freely available for research use. The 1000 Genomes Project is an international research effort coordinated by a consortium of 75 companies and organizations to establish the most detailed catalogue of human genetic variation. The project has grown to 200 terabytes of genomic data including DNA sequenced from more than 1,700 individuals that researchers can now access on AWS for use in disease research free of charge. The dataset containing the full genomic sequence of 1,700 individuals is now available to all via Amazon S3. The data can be found at: http://s3.amazonaws.com/1000genomes The 1000 Genomes Project aims to include the genomes of more than 2,662 individuals from 26 populations around the world, and the NIH will continue to add the remaining genome samples to the data collection this year. Public Data Sets on AWS provide a centralized repository of public data hosted on Amazon Simple Storage Service (Amazon S3). The data can be seamlessly accessed from AWS services such Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic MapReduce (Amazon EMR), which provide organizations with the highly scalable compute resources needed to take advantage of these large data collections. AWS is storing the public data sets at no charge to the community. Researchers pay only for the additional AWS resources they need for further processing or analysis of the data. All 200 TB of the latest 1000 Genomes Project data is available in a publicly available Amazon S3 bucket. You can access the data via simple HTTP requests, or take advantage of the AWS SDKs in languages such as Ruby, Java, Python, .NET and PHP. Researchers can use the Amazon EC2 utility computing service to dive into this data without the usual capital investment required to work with data at this scale. AWS also provides a number of orchestration and automation services to help teams make their research available to others to remix and reuse. Making the data available via a bucket in Amazon S3 also means that customers can crunch the information using Hadoop via Amazon Elastic MapReduce, and take advantage of the growing collection of tools for running bioinformatics job flows, such as CloudBurst and Crossbow.
View all literature mentionsSoftware package for analysis of large-scale genetic data sets with hundreds of thousands of markers genotyped on thousands of samples. BEAGLE can * phase genotype data (i.e. infer haplotypes) for unrelated individuals, parent-offspring pairs, and parent-offspring trios. * infer sporadic missing genotype data. * impute ungenotyped markers that have been genotyped in a reference panel. * perform single marker and haplotypic association analysis. * detect genetic regions that are homozygous-by-descent in an individual or identical-by-descent in pairs of individuals. Beagle can also be used in conjunction with PRESTO, a program for fast and flexible permutation testing. PRESTO can compute empirical distributions of order statistics, analyze stratified data, and determine significance levels for one-stage and two-stage genetic association studies. BEAGLE is written in Java and runs on any computing platform with a Java version 1.6 interpreter (e.g. Windows, Unix, Linux, Solaris, Mac).
View all literature mentionsCollection of curated, non-redundant genomic DNA, transcript RNA, and protein sequences produced by NCBI. Provides a reference for genome annotation, gene identification and characterization, mutation and polymorphism analysis, expression studies, and comparative analyses. Accessed through the Nucleotide and Protein databases.
View all literature mentionsA next-generation web-based application that aims to provide an integrated solution for both visualization and analysis of deep-sequencing data, along with simple access to public datasets.
View all literature mentionsSoftware application that can be used for converting Eland, Maq (.map), BED or other files into WIG files and identifying areas of enrichment (ChIP-Seq analysis).
View all literature mentionsSoftware tools for Motif Discovery and next-gen sequencing analysis. Used for analyzing ChIP-Seq, GRO-Seq, RNA-Seq, DNase-Seq, Hi-C and numerous other types of functional genomics sequencing data sets. Collection of command line programs for unix style operating systems written in Perl and C++.
View all literature mentionsMus musculus with name C57BL/6J from IMSR.
View all literature mentionsMus musculus with name BALB/cJ from IMSR.
View all literature mentionsMus musculus with name NOD/ShiLtJ from IMSR.
View all literature mentionsMus musculus with name SPRET/EiJ from IMSR.
View all literature mentions