Abstract
Depth estimation plays a crucial role in scene perception and understanding, which aims to predict distances between camera and real-world pixels from single or multiple images. It is a popular research field in computer vision, applying as an important step in many practical tasks such as 3D reconstruction. In recent years, depth estimation methods have drawn increasing attention and intensive research as a low-level vision task. Traditional methods use lidar to obtain high-precision depth information but cannot be widely used in practice due to the high cost of obtaining dense and accurate depth maps. In contrast, image-based depth estimation methods directly estimate depth based on input RGB images without expensive equipment, yielding more favor in applications. Specifically, image-based depth estimation methods can be divided into the multi-view and monocular estimation according to the number of required input pictures. Since the monocular camera has the advantages of low cost, more standard equipment, and convenient image acquisitions, compared with multi-view depth estimation methods, estimating depth information from monocular images is currently a more popular research field. With the rapid development of deep learning, monocular depth estimation based on deep neural networks has been widely studied, and many excellent methods have been proposed. This paper provides an overview of thestate-of-the-art on monocular depth estimation methods. First, the definition of monocular depth estimation, commonly used datasets, evaluation metrics, and applications are introduced. Then, we review representative methods according to different training manners: supervised, unsupervised and semi-supervised. The existing methods based on different learning manners are divided into several types, respectively. We summarize supervised methods and classify them into enhancing framework, introducing auxiliary information, improving loss function, classification-based methods, applying conditional random field, applying generative adversarial network, and methods based on partial depth labels. As for unsupervised methods, we classify them into models trained by image pairs and monocular videos. Improving strategies are divided into mask-based methods, applying visual odometry, applying generative adversarial network, and methods towards fast and light runtime performance. In terms of semi-supervised methods, they are similarly trained on image pairs or monocular videos. The proposed methods are classified into methods that apply a generative adversarial network and introduce semantic information, respectively. The ideas, advantages, and disadvantages of each type of methods are analyzed in detail. Finally, we sort out trends of future development and key technologies of monocular depth estimation methods based on deep learning. We aim to inspire readers to make further breakthroughs based on existing research summarized in this paper. Compared with the previous overviews, our paper mainly focuses on deep learning methods. We elaborated methods more comprehensively, point out ideas and characteristics of each method and use a more fine-grained and more accurate classification criterion, which is more helpful for readers to understand the overall research progress of depth estimation methods.Moreover, we provide detailed comparison results among all kinds of methods for quick investigation in tables. We also discuss the technical difficulties and development trends in monocular depth estimation, hoping to further inspire readers.
| Translated title of the contribution | Deep Learning Based Monocular Depth Estimation: A Survey |
|---|---|
| Original language | Chinese (Traditional) |
| Pages (from-to) | 1276-1307 |
| Number of pages | 32 |
| Journal | Jisuanji Xuebao/Chinese Journal of Computers |
| Volume | 45 |
| Issue number | 6 |
| DOIs | |
| State | Published - Jun 2022 |
| Externally published | Yes |
Fingerprint
Dive into the research topics of 'Deep Learning Based Monocular Depth Estimation: A Survey'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver