<?xml version='1.0' encoding='UTF-8'?><?xml-stylesheet href='static/style.xsl' type='text/xsl'?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T01:03:33Z</responseDate><request verb="GetRecord" identifier="oai:ecommons.cornell.edu:1813/120701" metadataPrefix="dim">https://ecommons.cornell.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:ecommons.cornell.edu:1813/120701</identifier><datestamp>2026-05-15T17:53:44Z</datestamp><setSpec>com_1813_35</setSpec><setSpec>col_1813_47</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="author">Aimuyo, Osayamen</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="chair" lang="en_US">Singh, Rachee</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="committeeMember" lang="en_US">De Sa, Christopher</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="committeeMember" lang="en_US">Guidi, Giulia</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2026-04-02T18:34:23Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2026-04-02T18:34:23Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2025-08</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="other">ProQuest Submission ID: 11955</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="other">ProQuest Publication ID: 30631704</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1813/120701</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="doi">https://doi.org/10.7298/ybdg-yr88</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="bibid">17422617</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">63 pages</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en_US">Distributed Machine Learning (DML) is increasingly recognized as a communication-bound workload, with most existing work aiming to alleviate this bottleneck through communication-computation overlap. However, current approaches—largely reliant on CPU-managed operator scheduling and synchronous communication collectives—leave significant performance on the table. This thesis identifies, analyzes, and addresses the inefficiencies arising from the interplay between CPU-driven collective communication and highly parallel GPU computation. We demonstrate that the standard bulk-synchronous communication model underutilizes GPU interconnect bandwidth and is especially vulnerable to straggler-induced performance degradation. Focusing on dynamic workloads such as Mixture-of-Experts (MoE), we highlight two critical bottlenecks. First, we demonstrate how CPU-driven execution limits the exploitation of task locality and introduces artificial synchronization barriers across distributed GPU tasks, resulting in performance reduction due to straggler effects. Second, we observe payload inefficiency at the application layer, where existing collective primitives force unnecessary padding of GPU communication buffers. To overcome these limitations, we propose a model of complete GPU residency, where inter-GPU communication is integrated directly into GPU kernels. We realize this vision in FlashDMoE: a persistent, in-kernel, actor-style operating system with packet switching that enables complete operator fusion for Distributed MoE (DMoE) into a single kernel, the first of its kind. FlashDMoE features a modular, message-driven architecture that supports lockless execution across tens of thousands of GPU threads and distributed GPUs as well. We demonstrate how FlashDMoE addresses the all-to-all communication bottleneck in expert parallelism and enables high-throughput, GPU-initiated communication. Evaluated against state-of-the-art distributed MoE frameworks, FlashDMoE achieves up to 9x higher GPU utilization, 6x lower latency, 5.7x higher throughput, and 4x better overlap efficiency compared to state-of-the-art baselines—despite FlashDMoE using FP32 while baselines use FP16.</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso">en</dim:field>
   <dim:field mdschema="dc" element="rights">Attribution-NonCommercial-NoDerivatives 4.0 International</dim:field>
   <dim:field mdschema="dc" element="rights" qualifier="uri">https://creativecommons.org/licenses/by-nc-nd/4.0/</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Accelerator</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Distributed Machine Learning</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">GPU</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Machine Learning</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Mixture-of-Experts</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Operator Fusion</dim:field>
   <dim:field mdschema="dc" element="title" lang="en_US">OS INSPIRED COMPLETE KERNEL FUSION</dim:field>
   <dim:field mdschema="dc" element="type" lang="en_US">dissertation or thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="discipline">Computer Science</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="grantor">Cornell University</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="level">Master of Science</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="name">M.S., Computer Science</dim:field>
   <dim:field mdschema="dcterms" element="license">https://hdl.handle.net/1813/59810.2</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="cris" element="virtual" qualifier="collection" authority="https://ecommons.cornell.edu/handle/1813/47" confidence="600">Cornell Theses and Dissertations</dim:field>
   <dim:field mdschema="cris" element="virtual" qualifier="author">Aimuyo, Osayamen</dim:field>
   <dim:field mdschema="cris" element="virtualsource" qualifier="collection">5893a6ea-7af3-41d7-abc6-04bcd26ab5df</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="02e0b327-fe03-4f7a-9187-96e66dc9b78b">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
	&lt;Language>en&lt;/Language>
   	&lt;Title>OS INSPIRED COMPLETE KERNEL FUSION&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2025-08&lt;/PublicationDate>
   	&lt;DOI>https://doi.org/10.7298/ybdg-yr88&lt;/DOI>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>Aimuyo, Osayamen&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;License>https://creativecommons.org/licenses/by-nc-nd/4.0/&lt;/License>
    &lt;Keyword>Accelerator&lt;/Keyword>
    &lt;Keyword>Distributed Machine Learning&lt;/Keyword>
    &lt;Keyword>GPU&lt;/Keyword>
    &lt;Keyword>Machine Learning&lt;/Keyword>
    &lt;Keyword>Mixture-of-Experts&lt;/Keyword>
    &lt;Keyword>Operator Fusion&lt;/Keyword>
   	&lt;Abstract>Distributed Machine Learning (DML) is increasingly recognized as a communication-bound workload, with most existing work aiming to alleviate this bottleneck through communication-computation overlap. However, current approaches—largely reliant on CPU-managed operator scheduling and synchronous communication collectives—leave significant performance on the table. This thesis identifies, analyzes, and addresses the inefficiencies arising from the interplay between CPU-driven collective communication and highly parallel GPU computation. We demonstrate that the standard bulk-synchronous communication model underutilizes GPU interconnect bandwidth and is especially vulnerable to straggler-induced performance degradation. Focusing on dynamic workloads such as Mixture-of-Experts (MoE), we highlight two critical bottlenecks. First, we demonstrate how CPU-driven execution limits the exploitation of task locality and introduces artificial synchronization barriers across distributed GPU tasks, resulting in performance reduction due to straggler effects. Second, we observe payload inefficiency at the application layer, where existing collective primitives force unnecessary padding of GPU communication buffers. To overcome these limitations, we propose a model of complete GPU residency, where inter-GPU communication is integrated directly into GPU kernels. We realize this vision in FlashDMoE: a persistent, in-kernel, actor-style operating system with packet switching that enables complete operator fusion for Distributed MoE (DMoE) into a single kernel, the first of its kind. FlashDMoE features a modular, message-driven architecture that supports lockless execution across tens of thousands of GPU threads and distributed GPUs as well. We demonstrate how FlashDMoE addresses the all-to-all communication bottleneck in expert parallelism and enables high-throughput, GPU-initiated communication. Evaluated against state-of-the-art distributed MoE frameworks, FlashDMoE achieves up to 9x higher GPU utilization, 6x lower latency, 5.7x higher throughput, and 4x better overlap efficiency compared to state-of-the-art baselines—despite FlashDMoE using FP32 while baselines use FP16.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>