COMPUTATIONAL APPROACHES TO UNDERSTANDING TRANSCRIPTION REGULATION: FROM NUCLEOTIDE RESOLUTION TO HIGHER ORDER STRUCTURES
Transcription regulation refers to the coordination of numerous processes and protein complexes that results in the production of RNA from DNA. Disruption of the process of transcription has been implicated in disease and developmental disorders, but despite intense study, aspects of transcription regulation continue to remain elusive. Work described here attempts to provide computational approaches with which to further our understanding these events. Studies in Chapter Two investigate the relationship between chromatin structure and transcription regulation. To fully understand gene regulation, we need to understand the processes driving higher order chromatin organization. The high level, territorial structure of interphase chromatin is well established, as are the building blocks of chromatin, the nucleosomes, that form the 10 nm fibers. However, the chromatin folding process that gets us from this basic, 10 nm fiber to the high-level territories is less well understood. I investigate whether changes in chromatin organization are a factor in the response of a cell to stress. To study this, the cell was perturbed via heat-shock and the change in expression measured for selected genes known to respond to heat-shock. Chromatin conformation was then measured, using the Hi-C assay, before and after heat-shock, focusing on the interactions between the enhancers and promoters of these selected genes. Changes in structure that either increase or decrease interactions between specific regions of chromatin after stress was applied would be evidence for these changes driving gene expression. In Chapter Three, I explore machine learning approaches to more fully exploit available precision nuclear run-on and sequencing (PRO-seq) data to improve genome annotations. The start and extent of transcription is very specific. By sequencing RNA transcripts, or by measuring factors that correlate with transcription (such as modifications to histones, or regions in which chromatin is accessible) we can infer the position of elements such as enhancers or promoters. I hypothesized that PRO-seq signal contains subtle patterns that have not been leveraged extensively by previous methods. Here I use two different neural network architectures to see whether inferring patterns in the signal gives my methods an advantage.