Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. Knowledge Gradient Methods for Bayesian Optimization

Knowledge Gradient Methods for Bayesian Optimization

File(s)
Wu_cornellgrad_0058F_10520.pdf (1.44 MB)
Permanent Link(s)
https://doi.org/10.7298/X4610XHC
https://hdl.handle.net/1813/56976
Collections
Cornell Theses and Dissertations
Author
Wu, Jian
Abstract

Bayesian optimization, a framework for global optimization of expensive-to-evaluate functions, has shown success in machine learning and experimental design because it is able to find global optima with a remarkably small number of potentially noisy objective function evaluations. In this dissertation, we study in detail how the concept of the knowledge gradient (KG) can be adopted to design novel Bayesian optimization algorithms. First, we propose a novel parallel Bayesian optimization algorithm by generalizing the concept of KG from the fully sequential setting to the parallel setting (qKG). By construction, this method provides a one-step Bayes-optimal batch of points to sample. We provide an efficient strategy for computing this Bayes-optimal batch of points, and we demonstrate that the parallel knowledge gradient method finds global optima significantly faster than previous batch Bayesian optimization algorithms on both synthetic test functions and when tuning hyperparameters of practical machine learning algorithms, especially when function evaluations are noisy. Second, we present a novel discretization-free strategy to calculate the set of points to be evaluated under the knowledge gradient method when used over a continuous domain. KG methods are widely studied for discrete ranking and selection problems, and provide a one-step Bayes-optimal point to sample, but all the previous efforts generalizing KG to continuous domains rely on a discretized finite approximation due to the computational challenges in calculating KG. However, the discretization introduces error and scales poorly as the dimension of the domain grows. In this chapter, we develop a fast discretization-free knowledge gradient method for Bayesian optimization, which is useful for all settings where KG is used over a continuous domain that overcomes these challenges. Third, we explore the “what, when, and why” of Bayesian optimization with derivative information. We also develop a Bayesian optimization algorithm that effectively leverages gradients. This algorithm accommodates incomplete and noisy gradient observations, can be used in both the sequential and batch settings, and can optionally reduce the computational overhead of inference by selecting the single most valuable directional derivative to retain. For this purpose, we develop a novel acquisition function, called the derivative-enabled knowledge-gradient (dKG). This generalizes the previously proposed batch knowledge gradient method to the derivative setting. We also provide a theoretical analysis of the algorithm: it is one- step Bayes-optimal by construction when derivatives are available, and we show (1) that it provides one-step value greater than in the derivative-free setting; and (2) that its estimator of the global optimum is asymptotically consistent. Fourth, we show some preliminary results on how KG can be adopted to settings where we have some low-fidelity but cheap approximations. To this end, we develop a novel Bayesian optimization algorithm, continuous-fidelity knowledge gradient (cfKG), which can adaptively choose both the fidelity and the desired point to sample by better balancing the trade-off between the information gain vs. the cost when we have some continuous parameters controlling the fidelity of the information source we can query. Some preliminary numerical results are shown.

Date Issued
2017-08-30
Keywords
Statistics
•
Batch-Sequential
•
Discretization-Free
•
Gradient-Enhanced
•
Knowledge Gradient
•
Multi-fidelity
•
Operations research
•
Computer science
•
Bayesian optimization
Committee Chair
Dai, Jiangang
Frazier, Peter
Committee Member
Joachims, Thorsten
Degree Discipline
Operations Research
Degree Name
Ph. D., Operations Research
Degree Level
Doctor of Philosophy
Rights
Attribution-ShareAlike 2.0 Generic
Rights URI
https://creativecommons.org/licenses/by-sa/2.0/
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance