Skip to main navigation Skip to search Skip to main content

Thompson Sampling on Asymmetric α-stable Bandits

  • Tsinghua University

Research output: Contribution to journalConference articlepeer-review

Abstract

In algorithm optimization in reinforcement learning, how to deal with the exploration-exploitation dilemma is particularly important. Multi-armed bandit problem can be designed to realize the dynamic balance between exploration and exploitation by changing the reward distribution. Thompson Sampling has been proposed in the literature for the solution of the multi-armed bandit problem by sampling rewards from posterior distributions. Recently, it was used to process non-Gaussian data with heavy tailed distributions. It is a common observation that various real-life data such as social network data and financial data demonstrate not only impulsive but also asymmetric characteristics. In this paper, we consider the Thompson Sampling approach for multi-armed bandit problem, in which rewards conform to an asymmetric a-stable distribution with unknown parameters and explore their applications in modelling financial and recommendation system data.

Original languageEnglish
Pages (from-to)434-441
Number of pages8
JournalInternational Conference on Agents and Artificial Intelligence
Volume3
DOIs
StatePublished - 2023
Externally publishedYes
Event15th International Conference on Agents and Artificial Intelligence, ICAART 2023 - Lisbon, Portugal
Duration: 22 Feb 202324 Feb 2023

Keywords

  • Asymmetric Reward
  • Multi-Armed Bandit Problem
  • Reinforcement Learning
  • Thompson Sampling
  • α-stable Distribution

Fingerprint

Dive into the research topics of 'Thompson Sampling on Asymmetric α-stable Bandits'. Together they form a unique fingerprint.

Cite this