Skip to main navigation Skip to search Skip to main content

Safely Learn to Fly Aircraft from Human: An Offline-Online Reinforcement Learning Strategy and Its Application to Aircraft Stall Recovery

  • Hantao Jiang
  • , Hao Xiong*
  • , Weifeng Zeng
  • , Yiming Ou
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Researchers have made many attempts to apply reinforcement learning (RL) to learn to fly aircraft in recent years. However, existing RL strategies are usually not safe (e.g., can lead to crash) in the initial stage of the training of an RL-based policy. For increasingly complex piloting tasks whose representative models are hard to establish, it is not safe to apply the existing RL strategies to learn an RL-based policy by interacting with an aircraft. To enhance the safety and feasibility of applying an RL-based policy to an aircraft, this study develops an offline-online RL strategy. The offline-online RL strategy learns an effective initialization for an RL-based flight control policy from human pilots without interacting with an aircraft through offline RL. Then, the offline-online RL strategy can further improve the RL-based flight control policy safely without leading to crash by interacting with the aircraft according to regular online RL, requiring no or very little intervention performed by a human pilot. To demonstrate and evaluate the offline-online RL strategy, the strategy is utilized to address the stall recovery problem of aircraft based on a flight simulator.

Original languageEnglish
Pages (from-to)8194-8207
Number of pages14
JournalIEEE Transactions on Aerospace and Electronic Systems
Volume59
Issue number6
DOIs
StatePublished - 1 Dec 2023
Externally publishedYes

Keywords

  • Aircraft
  • human pilot
  • reinforcement learning
  • safety
  • stall recovery

Fingerprint

Dive into the research topics of 'Safely Learn to Fly Aircraft from Human: An Offline-Online Reinforcement Learning Strategy and Its Application to Aircraft Stall Recovery'. Together they form a unique fingerprint.

Cite this