Skip to main navigation Skip to search Skip to main content

Quick Best Action Identification in Linear Bandit Problems

  • School of Electronics and Information Engineering, Harbin Institute of Technology
  • University of California at Davis

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, we consider a best action identification problem in the stochastic linear bandit setup with a fixed confident constraint. In the considered best action identification problem, instead of minimizing the accumulative regret as done in existing works, the learner aims to obtain an accurate estimate of the underlying parameter based on his action and reward sequences. To improve the estimation efficiency, the learner is allowed to select his action based his historical information; hence the whole procedure is designed in a sequential adaptive manner. We first show that the existing algorithms designed to minimize the accumulative regret is not a consistent estimator and hence is not a good policy for our problem. We then characterize a lower bound on the estimation error for any policy. We further design a simple policy and show that the estimation error of the designed policy achieves the same scaling order as that of the derived lower bound.

Original languageEnglish
Title of host publicationConference Record of the 52nd Asilomar Conference on Signals, Systems and Computers, ACSSC 2018
EditorsMichael B. Matthews
PublisherIEEE Computer Society
Pages1312-1316
Number of pages5
ISBN (Electronic)9781538692189
DOIs
StatePublished - 2 Jul 2018
Externally publishedYes
Event52nd Asilomar Conference on Signals, Systems and Computers, ACSSC 2018 - Pacific Grove, United States
Duration: 28 Oct 201831 Oct 2018

Publication series

NameConference Record - Asilomar Conference on Signals, Systems and Computers
Volume2018-October
ISSN (Electronic)2576-2303

Conference

Conference52nd Asilomar Conference on Signals, Systems and Computers, ACSSC 2018
Country/TerritoryUnited States
CityPacific Grove
Period28/10/1831/10/18

Fingerprint

Dive into the research topics of 'Quick Best Action Identification in Linear Bandit Problems'. Together they form a unique fingerprint.

Cite this