Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. Training Paradigms For Deep Residual Networks

Training Paradigms For Deep Residual Networks

File(s)
dms422.pdf (1.88 MB)
Permanent Link(s)
https://doi.org/10.7298/X44B2Z7F
https://hdl.handle.net/1813/44294
Collections
Cornell Theses and Dissertations
Author
Sedra, Daniel
Abstract

Convolutional networks are the current state of the art for image tasks. It has long been known that depth is key for increasing their expressive power, but many challenges rendered training difficult. With the advent of deep residual networks [13], the feasible depth of networks has increased from a few dozen to several hundred. Nevertheless several traditional machine learning problems persist such as overfitting, vanishing gradients, and diminishing feature reuse. Additionally, the training time for large networks is still measured in weeks. This thesis will detail two novel approaches for training deep residual networks that address the aforementioned persistent difficulties and present experimental evidence of their efficacy.

Date Issued
2016-05-29
Keywords
deep learning
•
machine learning
•
deep residual networks
Committee Chair
Weinberger,Kilian Quirin
Committee Member
Van Loan,Charles Francis
Degree Discipline
Computer Science
Degree Name
M.S., Computer Science
Degree Level
Master of Science
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance