Review
Nature Reviews Genetics 11, 345-355 (May 2010) | doi:10.1038/nrg2776
Alternative splicing and evolution: diversification, exon definition and function
Hadas Keren1, Galit Lev-Maor1 & Gil Ast1 About the authors
Top of page
Abstract
Over the past decade, it has been shown that alternative splicing (AS) is a major mechanism for the enhancement of transcriptome and proteome diversity, particularly in mammals. Splicing can be found in species from bacteria to humans, but its prevalence and characteristics vary considerably. Evolutionary studies are helping to address questions that are fundamental to understanding this important process: how and when did AS evolve? Which AS events are functional? What are the evolutionary forces that shaped, and continue to shape, AS? And what determines whether an exon is spliced in a constitutive or alternative manner? In this Review, we summarize the current knowledge of AS and evolution and provide insights into some of these unresolved questions.
Transcription regulation, Epigenetic and Next-generation sequencing related research and advances.
Friday, April 16, 2010
Friday, April 9, 2010
Molecular basis of S100 proteins interacting with the p53 homologs p63 and p73
Oncogene (2010) 29, 2024–2035; doi:10.1038/onc.2009.490; published online 8 February 2010
Molecular basis of S100 proteins interacting with the p53 homologs p63 and p73
J van Dieck1, T Brandt1, D P Teufel1, D B Veprintsev1, A C Joerger1 and A R Fersht1
1MRC Centre for Protein Engineering, Hills Road, Cambridge, UK
Correspondence: Professor AR Fersht, MRC Centre for Protein Engineering, Cambridge University, Hills Road, Cambridge, Cambs CB2 0QH, UK. E-mail: arf25@cam.ac.uk
Received 6 August 2009; Revised 16 October 2009; Accepted 27 October 2009; Published online 8 February 2010.
Top of page
Abstract
S100 proteins modulate p53 activity by interacting with its tetramerization (p53TET, residues 325–355) and transactivation (residues 1–57) domains. In this study, we characterized biophysically the binding of S100A1, S100A2, S100A4, S100A6 and S100B to homologous domains of p63 and p73 in vitro by fluorescence anisotropy, analytical ultracentrifugation and analytical gel filtration. We found that S100A1, S100A2, S100A4, S100A6 and S100B proteins bound different p63 and p73 tetramerization domain variants and naturally occurring isoforms with varying affinities in a calcium-dependent manner. Additional interactions were observed with peptides derived from the p63 and p73 N-terminal transactivation domains. Importantly, S100 proteins bound p63 and p73 with different affinities in their different oligomeric states, similarly to the differential modes of binding to p53. On the basis of our data, we hypothesize that S100 proteins regulate the oligomerization state of all three p53 family members and their isoforms, with a potential physiological relevance in developmental and disease-related processes. The regulation of the p53 family by S100 is complicated and depends on the target preference of each individual S100 protein, the concentration of the proteins and calcium, as well as the splicing variation of p63 or p73. Our results outlining the complexity of the interaction should be considered when studying the functional effects of S100 proteins in their biological context.
Keywords:
S100; p63; p73; tumor suppressor; protein–protein interaction
Molecular basis of S100 proteins interacting with the p53 homologs p63 and p73
J van Dieck1, T Brandt1, D P Teufel1, D B Veprintsev1, A C Joerger1 and A R Fersht1
1MRC Centre for Protein Engineering, Hills Road, Cambridge, UK
Correspondence: Professor AR Fersht, MRC Centre for Protein Engineering, Cambridge University, Hills Road, Cambridge, Cambs CB2 0QH, UK. E-mail: arf25@cam.ac.uk
Received 6 August 2009; Revised 16 October 2009; Accepted 27 October 2009; Published online 8 February 2010.
Top of page
Abstract
S100 proteins modulate p53 activity by interacting with its tetramerization (p53TET, residues 325–355) and transactivation (residues 1–57) domains. In this study, we characterized biophysically the binding of S100A1, S100A2, S100A4, S100A6 and S100B to homologous domains of p63 and p73 in vitro by fluorescence anisotropy, analytical ultracentrifugation and analytical gel filtration. We found that S100A1, S100A2, S100A4, S100A6 and S100B proteins bound different p63 and p73 tetramerization domain variants and naturally occurring isoforms with varying affinities in a calcium-dependent manner. Additional interactions were observed with peptides derived from the p63 and p73 N-terminal transactivation domains. Importantly, S100 proteins bound p63 and p73 with different affinities in their different oligomeric states, similarly to the differential modes of binding to p53. On the basis of our data, we hypothesize that S100 proteins regulate the oligomerization state of all three p53 family members and their isoforms, with a potential physiological relevance in developmental and disease-related processes. The regulation of the p53 family by S100 is complicated and depends on the target preference of each individual S100 protein, the concentration of the proteins and calcium, as well as the splicing variation of p63 or p73. Our results outlining the complexity of the interaction should be considered when studying the functional effects of S100 proteins in their biological context.
Keywords:
S100; p63; p73; tumor suppressor; protein–protein interaction
De novo motif identification improves the accuracy of predicting transcription factor binding sites in ChIP-Seq data analysis
De novo motif identification improves the accuracy of predicting transcription factor binding sites in ChIP-Seq data analysis
Valentina Boeva1,2,3,4, Didier Surdez1,2, Noƫlle Guillon1,2, Franck Tirode1,2, Anthony P. Fejes5, Olivier Delattre1,2 and Emmanuel Barillot1,3,4,*
1Institut Curie, 26 rue d’Ulm, 2INSERM, U830, Genetics and Biology of Cancer, 3INSERM, U900, Bioinformatics, Biostatistics, Epidemiology and Computational Systems Biology of Cancer, Paris, F-75248, 4Mines ParisTech, Fontainebleau, F-77300, France and 5Genome Sciences Centre, BC Cancer Agency, Vancouver, British Columbia, V5Z 4S6, Canada
*To whom correspondence should be addressed. Tel: ; Fax: +33 1 56 24 69 11; Email: micsa@curie.fr
Received November 10, 2009. Revised February 23, 2010. Accepted March 15, 2010.
Dramatic progress in the development of next-generation sequencing technologies has enabled accurate genome-wide characterization of the binding sites of DNA-associated proteins. This technique, baptized as ChIP-Seq, uses a combination of chromatin immunoprecipitation and massively parallel DNA sequencing. Other published tools that predict binding sites from ChIP-Seq data use only positional information of mapped reads. In contrast, our algorithm MICSA (Motif Identification for ChIP-Seq Analysis) combines this source of positional information with information on motif occurrences to better predict binding sites of transcription factors (TFs). We proved the greater accuracy of MICSA with respect to several other tools by running them on datasets for the TFs NRSF, GABP, STAT1 and CTCF. We also applied MICSA on a dataset for the oncogenic TF EWS-FLI1. We discovered >2000 binding sites and two functionally different binding motifs. We observed that EWS-FLI1 can activate gene transcription when (i) its binding site is located in close proximity to the gene transcription start site (up to ~150 kb), and (ii) it contains a microsatellite sequence. Furthermore, we observed that sites without microsatellites can also induce regulation of gene expression—positively as often as negatively—and at much larger distances (up to ~1 Mb).
Valentina Boeva1,2,3,4, Didier Surdez1,2, Noƫlle Guillon1,2, Franck Tirode1,2, Anthony P. Fejes5, Olivier Delattre1,2 and Emmanuel Barillot1,3,4,*
1Institut Curie, 26 rue d’Ulm, 2INSERM, U830, Genetics and Biology of Cancer, 3INSERM, U900, Bioinformatics, Biostatistics, Epidemiology and Computational Systems Biology of Cancer, Paris, F-75248, 4Mines ParisTech, Fontainebleau, F-77300, France and 5Genome Sciences Centre, BC Cancer Agency, Vancouver, British Columbia, V5Z 4S6, Canada
*To whom correspondence should be addressed. Tel: ; Fax: +33 1 56 24 69 11; Email: micsa@curie.fr
Received November 10, 2009. Revised February 23, 2010. Accepted March 15, 2010.
Dramatic progress in the development of next-generation sequencing technologies has enabled accurate genome-wide characterization of the binding sites of DNA-associated proteins. This technique, baptized as ChIP-Seq, uses a combination of chromatin immunoprecipitation and massively parallel DNA sequencing. Other published tools that predict binding sites from ChIP-Seq data use only positional information of mapped reads. In contrast, our algorithm MICSA (Motif Identification for ChIP-Seq Analysis) combines this source of positional information with information on motif occurrences to better predict binding sites of transcription factors (TFs). We proved the greater accuracy of MICSA with respect to several other tools by running them on datasets for the TFs NRSF, GABP, STAT1 and CTCF. We also applied MICSA on a dataset for the oncogenic TF EWS-FLI1. We discovered >2000 binding sites and two functionally different binding motifs. We observed that EWS-FLI1 can activate gene transcription when (i) its binding site is located in close proximity to the gene transcription start site (up to ~150 kb), and (ii) it contains a microsatellite sequence. Furthermore, we observed that sites without microsatellites can also induce regulation of gene expression—positively as often as negatively—and at much larger distances (up to ~1 Mb).
Detection of splice junctions from paired-end RNA-seq data by SpliceMap
Detection of splice junctions from paired-end RNA-seq data by SpliceMap
Kin Fai Au1, Hui Jiang1,2, Lan Lin3, Yi Xing3 and Wing Hung Wong1,*
1Department of Statistics, Stanford University, Stanford, CA 94305, 2Stanford Genome Technology Center, 855 California Ave, Palo Alto, CA 94304 and 3Department of Internal Medicine and Department of Biomedical Engineering, University of Iowa, Iowa City, IA, 52242, USA
*To whom correspondence should be addressed. Tel: ; Fax: +1 650 725 8977; Email: whwong@stanford.edu
Received December 7, 2009. Revised March 10, 2010. Accepted March 12, 2010.
Alternative splicing is a prevalent post-transcriptional process, which is not only important to normal cellular function but is also involved in human diseases. The newly developed second generation sequencing technique provides high-throughput data (RNA-seq data) to study alternative splicing events in different types of cells. Here, we present a computational method, SpliceMap, to detect splice junctions from RNA-seq data. This method does not depend on any existing annotation of gene structures and is capable of finding novel splice junctions with high sensitivity and specificity. It can handle long reads (50–100 nt) and can exploit paired-read information to improve mapping accuracy. Several parameters are included in the output to indicate the reliability of the predicted junction and help filter out false predictions. We applied SpliceMap to analyze 23 million paired 50-nt reads from human brain tissue. The results show at this depth of sequencing, RNA-seq can support reliable detection of splice junctions except for those that are present at very low level. Compared to current methods, SpliceMap can achieve 12% higher sensitivity without sacrificing specificity.
Kin Fai Au1, Hui Jiang1,2, Lan Lin3, Yi Xing3 and Wing Hung Wong1,*
1Department of Statistics, Stanford University, Stanford, CA 94305, 2Stanford Genome Technology Center, 855 California Ave, Palo Alto, CA 94304 and 3Department of Internal Medicine and Department of Biomedical Engineering, University of Iowa, Iowa City, IA, 52242, USA
*To whom correspondence should be addressed. Tel: ; Fax: +1 650 725 8977; Email: whwong@stanford.edu
Received December 7, 2009. Revised March 10, 2010. Accepted March 12, 2010.
Alternative splicing is a prevalent post-transcriptional process, which is not only important to normal cellular function but is also involved in human diseases. The newly developed second generation sequencing technique provides high-throughput data (RNA-seq data) to study alternative splicing events in different types of cells. Here, we present a computational method, SpliceMap, to detect splice junctions from RNA-seq data. This method does not depend on any existing annotation of gene structures and is capable of finding novel splice junctions with high sensitivity and specificity. It can handle long reads (50–100 nt) and can exploit paired-read information to improve mapping accuracy. Several parameters are included in the output to indicate the reliability of the predicted junction and help filter out false predictions. We applied SpliceMap to analyze 23 million paired 50-nt reads from human brain tissue. The results show at this depth of sequencing, RNA-seq can support reliable detection of splice junctions except for those that are present at very low level. Compared to current methods, SpliceMap can achieve 12% higher sensitivity without sacrificing specificity.
A Signal-Noise Model for Significance Analysis of ChIP-seq with Negative Control
A Signal-Noise Model for Significance Analysis of ChIP-seq with Negative Control
Han Xu 1,3, Lusy Handoko 2, Xueliang Wei 4, Chaopeng Ye 2, Jianpeng Sheng 5, Chia-Lin Wei 2, Feng Lin 3,* and Wing-Kin Sung 1,4,*
1Computational & Mathematical Biology Group, Genome Institute of Singapore, 138672, Singapore; 2Genome Technology & Biology Group, Genome Institute of Singapore, 138672, Singapore; 3School of Computer Engineering, Nanyang Technological University, 637553, Singapore; 4School of Computing, National University of Singapore, 117543, Singapore; 5School of Biological Science, Nanyang Techno-logical University, 637551, Singapore
*To whom correspondence should be addressed. Feng Lin, Wing-Kin Sung, E-mail: sungk@gis.a-star.edu.sg, asflin@ntu.edu.sg
Abstract
Motivation: ChIP-seq is becoming the main approach to the genome-wide study of protein-DNA interactions and histone modifications. Existing informatics tools perform well to extract strong ChIP-enriched sites. However, two questions remain to be answered: (a) to which extent is a ChIP-seq experiment able to reveal the weak ChIP-enriched sites? (b) are the weak sites biologically meaningful? To answer these questions, it is necessary to identify the weak ChIP signals from background noise.
Results: We propose a linear signal-noise model, in which a noise rate was introduced to represent the fraction of noise in a ChIP library. We developed an iterative algorithm to estimate the noise rate using a control library, and derived a library-swapping strategy for the FDR estimation. These approaches were integrated in a general-purpose framework, named CCAT (Control based ChIP-seq Analysis Tool), for the significance analysis of ChIP-seq. Applications to H3K4me3 and H3K36me3 datasets showed CCAT predicted significantly more ChIP-enriched sites than previous methods did. With the high sensitivity of CCAT prediction, we revealed distinct chromatin features associated to the strong and weak H3K4me3 sites.
Availability: http://cmb.gis.a-star.edu.sg/ChIPSeq/tools.htm
Han Xu 1,3, Lusy Handoko 2, Xueliang Wei 4, Chaopeng Ye 2, Jianpeng Sheng 5, Chia-Lin Wei 2, Feng Lin 3,* and Wing-Kin Sung 1,4,*
1Computational & Mathematical Biology Group, Genome Institute of Singapore, 138672, Singapore; 2Genome Technology & Biology Group, Genome Institute of Singapore, 138672, Singapore; 3School of Computer Engineering, Nanyang Technological University, 637553, Singapore; 4School of Computing, National University of Singapore, 117543, Singapore; 5School of Biological Science, Nanyang Techno-logical University, 637551, Singapore
*To whom correspondence should be addressed. Feng Lin, Wing-Kin Sung, E-mail: sungk@gis.a-star.edu.sg, asflin@ntu.edu.sg
Abstract
Motivation: ChIP-seq is becoming the main approach to the genome-wide study of protein-DNA interactions and histone modifications. Existing informatics tools perform well to extract strong ChIP-enriched sites. However, two questions remain to be answered: (a) to which extent is a ChIP-seq experiment able to reveal the weak ChIP-enriched sites? (b) are the weak sites biologically meaningful? To answer these questions, it is necessary to identify the weak ChIP signals from background noise.
Results: We propose a linear signal-noise model, in which a noise rate was introduced to represent the fraction of noise in a ChIP library. We developed an iterative algorithm to estimate the noise rate using a control library, and derived a library-swapping strategy for the FDR estimation. These approaches were integrated in a general-purpose framework, named CCAT (Control based ChIP-seq Analysis Tool), for the significance analysis of ChIP-seq. Applications to H3K4me3 and H3K36me3 datasets showed CCAT predicted significantly more ChIP-enriched sites than previous methods did. With the high sensitivity of CCAT prediction, we revealed distinct chromatin features associated to the strong and weak H3K4me3 sites.
Availability: http://cmb.gis.a-star.edu.sg/ChIPSeq/tools.htm
Thursday, April 8, 2010
Global methylation profiling of lymphoblastoid cell lines reveals epigenetic contributions to autism spectrum disorders and a novel autism candidate
Published online before print April 7, 2010 as doi: 10.1096/fj.10-154484.
Global methylation profiling of lymphoblastoid cell lines reveals epigenetic contributions to autism spectrum disorders and a novel autism candidate gene, RORA, whose protein product is reduced in autistic brain
AnhThu Nguyen, Tibor A. Rauch, Gerd P. Pfeifer, and Valerie W. Hu
E-mail contact: bcmvwh@gwumc.edu
Autism is currently considered a multigene disorder with epigenetic influences. To investigate the contribution of DNA methylation to autism spectrum disorders, we have recently completed large-scale methylation profiling by CpG island microarray analysis of lymphoblastoid cell lines derived from monozygotic twins discordant for diagnosis of autism and their nonautistic siblings. Methylation profiling revealed many candidate genes differentially methylated between discordant MZ twins as well as between both twins and nonautistic siblings. Bioinformatics analysis of the differentially methylated genes demonstrated enrichment for high-level functions including gene transcription, nervous system development, cell death/survival, and other biological processes implicated in autism. The methylation status of 2 of these candidate genes, BCL-2 and retinoic acid-related orphan receptor alpha (RORA), was further confirmed by bisulfite sequencing and methylation-specific PCR, respectively. Immunohistochemical analyses of tissue arrays containing slices of the cerebellum and frontal cortex of autistic and age- and sex-matched control subjects revealed decreased expression of RORA and BCL-2 proteins in the autistic brain. Our data thus confirm the role of epigenetic regulation of gene expression via differential DNA methylation in idiopathic autism, and furthermore link molecular changes in a peripheral cell model with brain pathobiology in autism.—Nguyen, A., Rauch, T. A., Pfeifer, G. P., Hu, V. W. Global methylation profiling of lymphoblastoid cell lines reveals epigenetic contributions to autism spectrum disorders and a novel autism candidate gene, RORA, whose protein product is reduced in autistic brain.
Global methylation profiling of lymphoblastoid cell lines reveals epigenetic contributions to autism spectrum disorders and a novel autism candidate gene, RORA, whose protein product is reduced in autistic brain
AnhThu Nguyen, Tibor A. Rauch, Gerd P. Pfeifer, and Valerie W. Hu
E-mail contact: bcmvwh@gwumc.edu
Autism is currently considered a multigene disorder with epigenetic influences. To investigate the contribution of DNA methylation to autism spectrum disorders, we have recently completed large-scale methylation profiling by CpG island microarray analysis of lymphoblastoid cell lines derived from monozygotic twins discordant for diagnosis of autism and their nonautistic siblings. Methylation profiling revealed many candidate genes differentially methylated between discordant MZ twins as well as between both twins and nonautistic siblings. Bioinformatics analysis of the differentially methylated genes demonstrated enrichment for high-level functions including gene transcription, nervous system development, cell death/survival, and other biological processes implicated in autism. The methylation status of 2 of these candidate genes, BCL-2 and retinoic acid-related orphan receptor alpha (RORA), was further confirmed by bisulfite sequencing and methylation-specific PCR, respectively. Immunohistochemical analyses of tissue arrays containing slices of the cerebellum and frontal cortex of autistic and age- and sex-matched control subjects revealed decreased expression of RORA and BCL-2 proteins in the autistic brain. Our data thus confirm the role of epigenetic regulation of gene expression via differential DNA methylation in idiopathic autism, and furthermore link molecular changes in a peripheral cell model with brain pathobiology in autism.—Nguyen, A., Rauch, T. A., Pfeifer, G. P., Hu, V. W. Global methylation profiling of lymphoblastoid cell lines reveals epigenetic contributions to autism spectrum disorders and a novel autism candidate gene, RORA, whose protein product is reduced in autistic brain.
Thursday, April 1, 2010
Epigenetic marks identify functional elements
Epigenetic marks identify functional elements
*
Randall H Morse
Journal name:
Nature Genetics
Volume:
42,
Pages:
282–284
Year published:
(2010)
DOI:
doi:10.1038/ng0410-282
Enhancers and transcription factor binding sites that control cell-specific transcription in higher eukaryotes can be found up to hundreds of kilobases from the promoters that they control, making their identification challenging. A new study uses a model based on histone modifications and chromatin dynamics to predict functional elements involved in androgen receptor response.
*
Randall H Morse
Journal name:
Nature Genetics
Volume:
42,
Pages:
282–284
Year published:
(2010)
DOI:
doi:10.1038/ng0410-282
Enhancers and transcription factor binding sites that control cell-specific transcription in higher eukaryotes can be found up to hundreds of kilobases from the promoters that they control, making their identification challenging. A new study uses a model based on histone modifications and chromatin dynamics to predict functional elements involved in androgen receptor response.
Subscribe to:
Posts (Atom)