TY - GEN
T1 - Voiceprint Diagnosis Method for Urban Rail Power Transformers Based on Mel Spectrogram and Improved Vision Transformer
AU - Zhou, Shangmin
AU - Chen, Liang
AU - Feng, Baiju
AU - Xiao, Jinyu
AU - Zheng, Wei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - The fault diagnosis for urban rail power transformers is crucial for ensuring the safety and reliability of urban rail power systems. Voiceprint-based fault diagnosis methods have gradually become a research hotspot due to their non-contact and uninterrupted advantages, especially with significant progress in the application of deep learning methods in this field. However, challenges such as poor robustness against noise interference and scarcity of real fault samples still persist. Therefore, this paper proposes a voiceprint diagnosis method for urban rail power transformers based on Mel spectrogram and an improved Vision Transformer(ViT). The improved ViT model integrates a CNN encoder and adopts a column-wise patch strategy, building upon the original ViT architecture. Experimental results show that the proposed method, with only 0.47M parameters and 0.149GLOPs, achieves a mere 6.75% accuracy drop when the signal-to-noise ratio decreases from 10dB to -4dB, outperforming other comparative methods.
AB - The fault diagnosis for urban rail power transformers is crucial for ensuring the safety and reliability of urban rail power systems. Voiceprint-based fault diagnosis methods have gradually become a research hotspot due to their non-contact and uninterrupted advantages, especially with significant progress in the application of deep learning methods in this field. However, challenges such as poor robustness against noise interference and scarcity of real fault samples still persist. Therefore, this paper proposes a voiceprint diagnosis method for urban rail power transformers based on Mel spectrogram and an improved Vision Transformer(ViT). The improved ViT model integrates a CNN encoder and adopts a column-wise patch strategy, building upon the original ViT architecture. Experimental results show that the proposed method, with only 0.47M parameters and 0.149GLOPs, achieves a mere 6.75% accuracy drop when the signal-to-noise ratio decreases from 10dB to -4dB, outperforming other comparative methods.
KW - fault diagnosis
KW - mel spectrogram
KW - power transformers
KW - vision transformer
KW - voiceprint recognition
UR - https://www.scopus.com/pages/publications/105007731359
U2 - 10.1109/CSECS64665.2025.11009465
DO - 10.1109/CSECS64665.2025.11009465
M3 - 会议稿件
AN - SCOPUS:105007731359
T3 - CSECS 2025 - Proceedings of 2025 7th International Conference on Software Engineering and Computer Science
BT - CSECS 2025 - Proceedings of 2025 7th International Conference on Software Engineering and Computer Science
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 7th International Conference on Software Engineering and Computer Science, CSECS 2025
Y2 - 21 March 2025 through 23 March 2025
ER -