Abstract
In algorithm optimization in reinforcement learning, how to deal with the exploration-exploitation dilemma is particularly important. Multi-armed bandit problem can be designed to realize the dynamic balance between exploration and exploitation by changing the reward distribution. Thompson Sampling has been proposed in the literature for the solution of the multi-armed bandit problem by sampling rewards from posterior distributions. Recently, it was used to process non-Gaussian data with heavy tailed distributions. It is a common observation that various real-life data such as social network data and financial data demonstrate not only impulsive but also asymmetric characteristics. In this paper, we consider the Thompson Sampling approach for multi-armed bandit problem, in which rewards conform to an asymmetric a-stable distribution with unknown parameters and explore their applications in modelling financial and recommendation system data.
| Original language | English |
|---|---|
| Pages (from-to) | 434-441 |
| Number of pages | 8 |
| Journal | International Conference on Agents and Artificial Intelligence |
| Volume | 3 |
| DOIs | |
| State | Published - 2023 |
| Externally published | Yes |
| Event | 15th International Conference on Agents and Artificial Intelligence, ICAART 2023 - Lisbon, Portugal Duration: 22 Feb 2023 → 24 Feb 2023 |
Keywords
- Asymmetric Reward
- Multi-Armed Bandit Problem
- Reinforcement Learning
- Thompson Sampling
- α-stable Distribution
Fingerprint
Dive into the research topics of 'Thompson Sampling on Asymmetric α-stable Bandits'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver