Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
This service exclusively searches for literature that cites resources. Please be aware that the total number of searchable documents is limited to those containing RRIDs and does not include all open-access literature.
We have previously shown that 5' halves from tRNAGlyGCC and tRNAGluCUC are the most enriched small RNAs in the extracellular space of human cell lines, and especially in the non-vesicular fraction. Extracellular RNAs are believed to require protection by either encapsulation in vesicles or ribonucleoprotein complex formation. However, deproteinization of non-vesicular tRNA halves does not affect their retention in size-exclusion chromatography. Thus, we considered alternative explanations for their extracellular stability. In-silico analysis of the sequence of these tRNA-derived fragments showed that tRNAGly 5' halves can form homodimers or heterodimers with tRNAGlu 5' halves. This capacity is virtually unique to glycine tRNAs. By analyzing synthetic oligonucleotides by size exclusion chromatography, we provide evidence that dimerization is possible in vitro. tRNA halves with single point substitutions preventing dimerization are degraded faster both in controlled nuclease digestion assays and after transfection in cells, showing that dimerization can stabilize tRNA halves against the action of cellular nucleases. Finally, we give evidence supporting dimerization of endogenous tRNAGlyGCC 5' halves inside cells. Considering recent reports have shown that 5' tRNA halves from Ala and Cys can form tetramers, our results highlight RNA intermolecular structures as a new layer of complexity in the biology of tRNA-derived fragments.
DNA methylation is an epigenetic mechanism known to affect gene expression and aberrant DNA methylation patterns have been described in cancer. However, only a small fraction of differential methylation events target genes with a defined role in cancer, raising the question of how aberrant DNA methylation contributes to carcinogenesis. As recently a link has been suggested between methylation patterns arising in ageing and those arising in cancer, we asked which aberrations are unique to cancer and which are the product of normal ageing processes. We therefore compared the methylation patterns between ageing and cancer in multiple tissues. We observed that hypermethylation preferentially occurs in regulatory elements, while hypomethylation is associated with structural features of the chromatin. Specifically, we observed consistent hypomethylation of late-replicating, lamina-associated domains. The extent of hypomethylation was stronger in cancer, but in both ageing and cancer it was proportional to the replication timing of the region and the cell division rate of the tissue. Moreover, cancer patients who displayed more hypomethylation in late-replicating, lamina-associated domains had higher expression of cell division genes. These findings suggest that different cell division rates contribute to tissue- and cancer type-specific DNA methylation profiles.
Cell-specific patterns of gene expression are determined by combinatorial actions of sequence-specific transcription factors at cis-regulatory elements. Studies indicate that relatively simple combinations of lineage-determining transcription factors (LDTFs) play dominant roles in the selection of enhancers that establish cell identities and functions. LDTFs require collaborative interactions with additional transcription factors to mediate enhancer function, but the identities of these factors are often unknown. We have shown that natural genetic variation between individuals has great utility for discovering collaborative transcription factors. Here, we introduce MMARGE (Motif Mutation Analysis of Regulatory Genomic Elements), the first publicly available suite of software tools that integrates genome-wide genetic variation with epigenetic data to identify collaborative transcription factor pairs. MMARGE is optimized to work with chromatin accessibility assays (such as ATAC-seq or DNase I hypersensitivity), as well as transcription factor binding data collected by ChIP-seq. Herein, we provide investigators with rationale for each step in the MMARGE pipeline and key differences for analysis of datasets with different experimental designs. We demonstrate the utility of MMARGE using mouse peritoneal macrophages, liver cells, and human lymphoblastoid cells. MMARGE provides a powerful tool to identify combinations of cell type-specific transcription factors while simultaneously interpreting functional effects of non-coding genetic variation.
The recent identification and development of RNA-guided enzymes for programmable cleavage of target nucleic acids offers exciting possibilities for both therapeutic and biotechnological applications. However, critical challenges such as expensive guide RNAs and inability to predict the efficiency of target recognition, especially for highly-structured RNAs, remain to be addressed. Here, we introduce a programmable RNA restriction enzyme, based on a budding yeast Argonaute (AGO), programmed with cost-effective 23-nucleotide (nt) single-stranded DNAs as guides. DNA guides offer the advantage that diverse sequences can be easily designed and purchased, enabling high-throughput screening to identify optimal recognition sites in the target RNA. Using this DNA-induced slicing complex (DISC) programmed with 11 different guide DNAs designed to span the sequence, sites of cleavage were identified in the 352-nt human immunodeficiency virus type 1 5'-untranslated region. This assay, coupled with primer extension and capillary electrophoresis, allows detection and relative quantification of all DISC-cleavage sites simultaneously in a single reaction. Comparison between DISC cleavage and RNase H cleavage reveals that DISC not only cleaves solvent-exposed sites, but also sites that become more accessible upon DISC binding. This study demonstrates the advantages of the DISC system for programmable cleavage of highly-structured, functional RNAs.
MicroRNAs often occur in families whose members share an identical 5' terminal 'seed' sequence. The seed is a major determinant of miRNA activity, and family members are thought to act redundantly on target mRNAs with perfect seed matches, i.e. sequences complementary to the seed. However, recently sequences outside the seed were reported to promote silencing by individual miRNA family members. Here, we examine this concept and the importance of miRNA specificity for the robustness of developmental gene control. Using the let-7 miRNA family in Caenorhabditis elegans, we find that seed match imperfections can increase specificity by requiring extensive pairing outside the miRNA seed region for efficient silencing and that such specificity is needed for faithful worm development. In addition, for some target site architectures, elevated miRNA levels can compensate for a lack of complementarity outside the seed. Thus, some target sites require higher miRNA concentration for silencing than others, contrasting with a traditional binary distinction between functional and non-functional sites. We conclude that changing miRNA concentrations can alter cellular miRNA target repertoires. This diversifies possible biological outcomes of miRNA-mediated gene regulation and stresses the importance of target validation under physiological conditions to understand miRNA functions in vivo.
The HSF and FOXO families of transcription factors play evolutionarily conserved roles in stress resistance and lifespan. In humans, the rs2802292 G-allele at FOXO3 locus has been associated with longevity in all human populations tested; moreover, its copy number correlated with reduced frequency of age-related diseases in centenarians. At the molecular level, the intronic rs2802292 G-allele correlated with increased expression of FOXO3, suggesting that FOXO3 intron 2 may represent a regulatory region. Here we show that the 90-bp sequence around the intronic single nucleotide polymorphism rs2802292 has enhancer functions, and that the rs2802292 G-allele creates a novel HSE binding site for HSF1, which induces FOXO3 expression in response to diverse stress stimuli. At the molecular level, HSF1 mediates the occurrence of a promoter-enhancer interaction at FOXO3 locus involving the 5'UTR and the rs2802292 region. These data were confirmed in various cellular models including human HAP1 isogenic cell lines (G/T). Our functional studies highlighted the importance of the HSF1-FOXO3-SOD2/CAT/GADD45A cascade in cellular stress response and survival by promoting ROS detoxification, redox balance and DNA repair. Our findings suggest the existence of an HSF1-FOXO3 axis in human cells that could be involved in stress response pathways functionally regulating lifespan and disease susceptibility.
Accumulating evidence indicates that transcription factor (TF) binding sites, or cis-regulatory elements (CREs), and their clusters termed cis-regulatory modules (CRMs) play a more important role than do gene-coding sequences in specifying complex traits in humans, including the susceptibility to common complex diseases. To fully characterize their roles in deriving the complex traits/diseases, it is necessary to annotate all CREs and CRMs encoded in the human genome. However, the current annotations of CREs and CRMs in the human genome are still very limited and mostly coarse-grained, as they often lack the detailed information of CREs in CRMs. Here, we integrated 620 TF ChIP-seq datasets produced by the ENCODE project for 168 TFs in 79 different cell/tissue types and predicted an unprecedentedly completely map of CREs in CRMs in the human genome at single nucleotide resolution. The map includes 305 912 CRMs containing a total of 1 178 913 CREs belonging to 736 unique TF binding motifs. The predicted CREs and CRMs tend to be subject to either purifying selection or positive selection, thus are likely to be functional. Based on the results, we also examined the status of available ChIP-seq datasets for predicting the entire regulatory genome of humans.
The unprecedented growth of high-throughput sequencing has led to an ever-widening annotation gap in protein databases. While computational prediction methods are available to make up the shortfall, a majority of public web servers are hindered by practical limitations and poor performance. Here, we introduce PANNZER2 (Protein ANNotation with Z-scoRE), a fast functional annotation web server that provides both Gene Ontology (GO) annotations and free text description predictions. PANNZER2 uses SANSparallel to perform high-performance homology searches, making bulk annotation based on sequence similarity practical. PANNZER2 can output GO annotations from multiple scoring functions, enabling users to see which predictions are robust across predictors. Finally, PANNZER2 predictions scored within the top 10 methods for molecular function and biological process in the CAFA2 NK-full benchmark. The PANNZER2 web server is updated on a monthly schedule and is accessible at http://ekhidna2.biocenter.helsinki.fi/sanspanz/. The source code is available under the GNU Public Licence v3.
Small-molecule compounds that target mismatched base pairs in DNA offer a novel prospective for cancer diagnosis and therapy. The potent anticancer antibiotic echinomycin functions by intercalating into DNA at CpG sites. Surprisingly, we found that the drug strongly prefers to bind to consecutive CpG steps separated by a single T:T mismatch. The preference appears to result from enhanced cooperativity associated with the binding of the second echinomycin molecule. Crystallographic studies reveal that this preference originates from the staggered quinoxaline rings of the two neighboring antibiotic molecules that surround the T:T mismatch forming continuous stacking interactions within the duplex. These and other associated changes in DNA conformation allow the formation of a minor groove pocket for tight binding of the second echinomycin molecule. We also show that echinomycin displays enhanced cytotoxicity against mismatch repair-deficient cell lines, raising the possibility of repurposing the drug for detection and treatment of mismatch repair-deficient cancers.
Aberrant chromatin transformation dysregulates gene expression and may be an important driver of tumorigenesis. However, the functional role of chromosomal dynamics in tumorigenesis remains to be elucidated. Here, using in vitro and in vivo experiments, we reveal a novel long noncoding (lncing) driver at chr12p13.3, in which a novel lncRNA GALNT8 Antisense Upstream 1 (GAU1) is initially activated by an open chromatin status, triggering recruitment of the transcription elongation factor TCEA1 at the oncogene GALNT8 promoter and cis-activates the expression of GALNT8. Analysis of The Cancer Genome Atlas (TCGA) clinical database revealed that the GAU1/GALNT8 driver serves as an important indicative biomarker, and targeted silencing of GAU1 via the HKP-encapsulated method exhibited therapeutic efficacy in orthotopic xenografts. Our study presents a novel oncogenetic mechanism in which aberrant tuning of the chromatin state at specific chromosomal loci exposes factor-binding sites, leading to recruitment of trans-factor and activation of oncogenetic driver, thereby provide a novel alternative concept of chromatin dynamics in tumorigenesis.
Calling variants from next-generation sequencing (NGS) data or discovering discordant sequences between two NGS data sets is challenging. We developed a computer algorithm, ADIScan1, to call variants by comparing the fractions of allelic reads in a tester to the universal reference genome. We then created ADIScan2 by modifying the algorithm to directly compare two sets of NGS data and predict discordant sequences between two testers. ADIScan1 detected >99.7% of variants called by GATK with an additional 724 393 SNVs. ADIScan2 identified ∼500 candidates of discordant sequences in each of two pairs of the monozygotic twins. About 200 of these candidates were included in the ∼2800 predicted by VarScan2. We verified 66 true discordant sequences among the candidates that ADIScan2 and VarScan2 exclusively predicted. ADIScan2 detected many discordant sequences overlooked by VarScan2 and Mutect, which specialize in detecting low frequency mutations in genetically heterogeneous cancerous tissues. Numbers of verified sequences alone were >5 times more than expected based on recently estimated mutation rates from whole genome sequences. Estimated post-zygotic mutation rates were 1.68 × 10-7 in this study. ADIScan1 and 2 would complement existing tools in screening causative mutations of diverse genetic diseases and comparing two sets of genome sequences, respectively.
With the emergence of Next Generation Sequencing (NGS) technologies, a large volume of sequence data in particular de novo sequencing was rapidly produced at relatively low costs. In this context, computational tools are increasingly important to assist in the identification of relevant information to understand the functioning of organisms. This work introduces BASiNET, an alignment-free tool for classifying biological sequences based on the feature extraction from complex network measurements. The method initially transform the sequences and represents them as complex networks. Then it extracts topological measures and constructs a feature vector that is used to classify the sequences. The method was evaluated in the classification of coding and non-coding RNAs of 13 species and compared to the CNCI, PLEK and CPC2 methods. BASiNET outperformed all compared methods in all adopted organisms and datasets. BASiNET have classified sequences in all organisms with high accuracy and low standard deviation, showing that the method is robust and non-biased by the organism. The proposed methodology is implemented in open source in R language and freely available for download at https://cran.r-project.org/package=BASiNET.
Small RNAs (sRNAs) are short, non-coding RNAs that play critical roles in many important biological pathways. They suppress the translation of messenger RNAs (mRNAs) by directing the RNA-induced silencing complex to their sequence-specific mRNA target(s). In plants, this typically results in mRNA cleavage and subsequent degradation of the mRNA. The resulting mRNA fragments, or degradome, provide evidence for these interactions, and thus degradome analysis has become an important tool for sRNA target prediction. Even so, with the continuing advances in sequencing technologies, not only are larger and more complex genomes being sequenced, but also degradome and associated datasets are growing both in number and read count. As a result, existing degradome analysis tools are unable to process the volume of data being produced without imposing huge resource and time requirements. Moreover, these tools use stringent, non-configurable targeting rules, which reduces their flexibility. Here, we present a new and user configurable software tool for degradome analysis, which employs a novel search algorithm and sequence encoding technique to reduce the search space during analysis. The tool significantly reduces the time and resources required to perform degradome analysis, in some cases providing more than two orders of magnitude speed-up over current methods.
Derivatives of 5-hydroxyuridine (ho5U), such as 5-methoxyuridine (mo5U) and 5-oxyacetyluridine (cmo5U), are ubiquitous modifications of the wobble position of bacterial tRNA that are believed to enhance translational fidelity by the ribosome. In gram-negative bacteria, the last step in the biosynthesis of cmo5U from ho5U involves the unique metabolite carboxy S-adenosylmethionine (Cx-SAM) and the carboxymethyl transferase CmoB. However, the equivalent position in the tRNA of Gram-positive bacteria is instead mo5U, where the methyl group is derived from SAM and installed by an unknown methyltransferase. By utilizing a cmoB-deficient strain of Escherichia coli as a host and assaying for the formation of mo5U in total RNA isolates with methyltransferases of unknown function from Bacillus subtilis, we found that this modification is installed by the enzyme TrmR (formerly known as YrrM). Furthermore, X-ray crystal structures of TrmR with and without the anticodon stemloop of tRNAAla have been determined, which provide insight into both sequence and structure specificity in the interactions of TrmR with tRNA.
Astrocytes play crucial roles in the central nervous system, and defects in astrocyte function are closely related to many neurological disorders. Studying the mechanism of gliogenesis has important implications for understanding and treating brain diseases. Epigenetic regulations have essential roles during mammalian brain development. Here, we demonstrate that histone H2A.Z.1 is necessary for the specification of multiple neural precursor cells (NPCs) and has specialized functions that regulate gliogenesis. Depletion of H2A.Z.1 suppresses gliogenesis and results in reduced astrocyte differentiation. Additionally, H2A.Z.1 regulates the acetylation of H3K56 (H3K56ac) by cooperating with the chaperone of ASF1a. Furthermore, RNA-seq data indicate that folate receptor 1 (FOLR1) participates in gliogenesis through the JAK-STAT signaling pathway. Taken together, our results demonstrate that H2A.Z.1 is a key regulator of gliogenesis because it interacts with ASF1a to regulate H3K56ac and then directly affects the expression of FOLR1, which acts as a signal-transducing component of the JAK-STAT signaling pathway.
Long intergenic non-coding RNAs (lincRNAs) are non-coding transcripts >200 nucleotides long that do not overlap protein-coding sequences. Importantly, such elements are known to be tissue-specifically expressed and to play a widespread role in gene regulation across thousands of genomic loci. However, very little is known of the mechanisms for the evolutionary biogenesis of these RNA elements, especially given their poor conservation across species. It has been proposed that lincRNAs might arise from pseudogenes. To test this systematically, we developed a novel method that searches for remnants of protein-coding sequences within lincRNA transcripts; the hypothesis is that we can trace back their biogenesis from protein-coding genes or posterior transposon/retrotransposon insertions. Applying this method, we found 203 human lincRNA genes with regions significantly similar to protein-coding sequences. Our method provides a visualization tool to trace the evolutionary biogenesis of lincRNAs with respect to protein-coding genes by sequence divergence. Subsequently, we show the expression correlation between lincRNAs and their identified parental protein-coding genes using public RNA-seq repositories, hinting at novel gene regulatory relationships. In summary, we developed a novel computational methodology to study non-coding gene sequences, which can be applied to identify the evolutionary biogenesis and function of lincRNAs.
Genome-wide association studies (GWAS), relying on hundreds of thousands of individuals, have revealed >200 genomic loci linked to metabolic disease (MD). Loss of insulin sensitivity (IS) is a key component of MD and we hypothesized that discovery of a robust IS transcriptome would help reveal the underlying genomic structure of MD. Using 1,012 human skeletal muscle samples, detailed physiology and a tissue-optimized approach for the quantification of coding (>18,000) and non-coding (>15,000) RNA (ncRNA), we identified 332 fasting IS-related genes (CORE-IS). Over 200 had a proven role in the biochemistry of insulin and/or metabolism or were located at GWAS MD loci. Over 50% of the CORE-IS genes responded to clinical treatment; 16 quantitatively tracking changes in IS across four independent studies (P = 0.0000053: negatively: AGL, G0S2, KPNA2, PGM2, RND3 and TSPAN9 and positively: ALDH6A1, DHTKD1, ECHDC3, MCCC1, OARD1, PCYT2, PRRX1, SGCG, SLC43A1 and SMIM8). A network of ncRNA positively related to IS and interacted with RNA coding for viral response proteins (P < 1 × 10-48), while reduced amino acid catabolic gene expression occurred without a change in expression of oxidative-phosphorylation genes. We illustrate that combining in-depth physiological phenotyping with robust RNA profiling methods, identifies molecular networks which are highly consistent with the genetics and biochemistry of human metabolic disease.
In contrast to GNRA tetraloop receptors that are common in RNA, receptors for the more thermostable UNCG loops have remained elusive for almost three decades. An analysis of all RNA structures with resolution ≤3.0 Å from the PDB allowed us to identify three previously unnoticed receptors for UNCG and GNRA tetraloops that adopt a common UNCG fold, named 'Z-turn' in agreement with our previously published nomenclature. These receptors recognize the solvent accessible second Z-turn nucleotide in different but specific ways. Two receptors participating in a complex network of tertiary interactions are associated with the rRNA UUCG and GAAA Z-turns capping helices H62 and H35a in rRNA large subunits. Structural comparison of fully assembled ribosomes and comparative sequence analysis of >6500 rRNA sequences helped us recognize that these motifs are almost universally conserved in rRNA, where they may contribute to organize the large subunit around the subdomain-IV four-way junction. The third UCCG receptor was identified in a rRNA/protein construct crystallized at acidic pH. These three non-redundant Z-turn receptors are relevant for our understanding of the assembly of rRNA and other long-non-coding RNAs, as well as for the design of novel folding motifs for synthetic biology.
Riboswitches are structured mRNA sequences that regulate gene expression by directly binding intracellular metabolites. Generating the appropriate regulatory response requires the RNA rapidly and stably acquire higher-order structure to form the binding pocket, bind the appropriate effector molecule and undergo a structural transition to inform the expression machinery. These requirements place riboswitches under strong kinetic constraints, likely restricting the sequence space accessible by recurrent structural modules such as the kink turn and the T-loop. Class-II cobalamin riboswitches contain two T-loop modules: one directing global folding of the RNA and another buttressing the ligand binding pocket. While the T-loop module directing folding is highly conserved, the T-loop associated with binding is substantially less so, with no clear consensus sequence. To further understand the functional role of the binding-associated module, a functional genetic screen of a library of riboswitches with the T-loop and its interacting nucleotides was used to build an experimental phylogeny comprised of sequences that possess a wide range of cobalamin-dependent regulatory activity. Our results reveal conservation patterns of the T-loop and its interaction with the binding core that allow for rapid tertiary structure formation and demonstrate its importance for generating strong ligand-dependent repression of mRNA expression.
Acquisition of foreign DNA by natural transformation is an important mechanism of adaptation and evolution in diverse microbial species. Here, we characterize the mechanism of ComM, a broadly conserved AAA+ protein previously implicated in homologous recombination of transforming DNA (tDNA) in naturally competent Gram-negative bacterial species. In vivo, we found that ComM was required for efficient comigration of linked genetic markers in Vibrio cholerae and Acinetobacter baylyi, which is consistent with a role in branch migration. Also, ComM was particularly important for integration of tDNA with increased sequence heterology, suggesting that its activity promotes the acquisition of novel DNA sequences. In vitro, we showed that purified ComM binds ssDNA, oligomerizes into a hexameric ring, and has bidirectional helicase and branch migration activity. Based on these data, we propose a model for tDNA integration during natural transformation. This study provides mechanistic insight into the enigmatic steps involved in tDNA integration and uncovers the function of a protein required for this conserved mechanism of horizontal gene transfer.
Welcome to the NIF Resources search. From here you can search through a compilation of resources used by NIF and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that NIF has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on NIF then you can log in from here to get additional features in NIF such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into NIF you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the facets that you can filter your papers by.
From here we'll present any options for the literature, such as exporting your current results.
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.
Year:
Count: