TY - GEN
T1 - DA-MLAD
T2 - 40th ACM International Conference on Supercomputing, ICS 2026
AU - Tan, Kai
AU - Du, Yangliu
AU - Zhan, Dongyang
AU - Xie, Yang
AU - Yu, Haining
AU - Zhao, Bei
AU - Liu, Hao
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/7/5
Y1 - 2026/7/5
N2 - Large-scale supercomputing systems generate massive, continuously evolving log streams essential for fault diagnosis, but log pattern drift - caused by system upgrades, workload variations, and hardware aging - rapidly degrades detection models, missing emerging failures or flooding operators with false alarms. Existing online methods adapt blindly whenever distribution shifts occur, unable to distinguish genuine fault evolution from superficial format changes - leading to either over-adaptation that discards relevant knowledge, or under-adaptation that misses emerging failures. We present DA-MLAD, a drift-decomposed online meta-learning framework that first decomposes observed drift into its critical sources, then adapts accordingly. Our key insight is that fault concepts (memory errors, network timeouts, I/O stalls) provide a stable semantic view even as log syntax evolves. The core contribution is the Semantic Drift Ratio (SDR), which exploits the aggregation structure of concept mapping to decompose observed drift: by measuring shift at both template and concept levels, SDR disentangles surface-level format changes from semantic-level fault evolution, enabling principled adaptation that preserves knowledge when only formats change while learning aggressively when fault semantics genuinely evolve. Experiments on three production supercomputer logs (BGL, Thunderbird, Spirit) under closed-budget protocol (2.5% labeled templates, no extra supervision for meta-update) demonstrate 3.3-4.4 F1-point improvements over state-of-the-art methods under sustained drift, with 68% less forgetting.
AB - Large-scale supercomputing systems generate massive, continuously evolving log streams essential for fault diagnosis, but log pattern drift - caused by system upgrades, workload variations, and hardware aging - rapidly degrades detection models, missing emerging failures or flooding operators with false alarms. Existing online methods adapt blindly whenever distribution shifts occur, unable to distinguish genuine fault evolution from superficial format changes - leading to either over-adaptation that discards relevant knowledge, or under-adaptation that misses emerging failures. We present DA-MLAD, a drift-decomposed online meta-learning framework that first decomposes observed drift into its critical sources, then adapts accordingly. Our key insight is that fault concepts (memory errors, network timeouts, I/O stalls) provide a stable semantic view even as log syntax evolves. The core contribution is the Semantic Drift Ratio (SDR), which exploits the aggregation structure of concept mapping to decompose observed drift: by measuring shift at both template and concept levels, SDR disentangles surface-level format changes from semantic-level fault evolution, enabling principled adaptation that preserves knowledge when only formats change while learning aggressively when fault semantics genuinely evolve. Experiments on three production supercomputer logs (BGL, Thunderbird, Spirit) under closed-budget protocol (2.5% labeled templates, no extra supervision for meta-update) demonstrate 3.3-4.4 F1-point improvements over state-of-the-art methods under sustained drift, with 68% less forgetting.
KW - Supercomputing reliability
KW - concept drift
KW - continual learning
KW - log anomaly detection
KW - online meta-learning
UR - https://www.scopus.com/pages/publications/105046601505
U2 - 10.1145/3797905.3807873
DO - 10.1145/3797905.3807873
M3 - 会议稿件
AN - SCOPUS:105046601505
T3 - Proceedings of the International Conference on Supercomputing
SP - 675
EP - 686
BT - ICS 2026 - Proceedings of the 40th ACM International Conference on Supercomputing
PB - Association for Computing Machinery
Y2 - 6 July 2026 through 9 July 2026
ER -