Abstract
HFTC_2 is a practical dependable distributed computer system, it works correctly in the existence of arbitrary fault combination with checkpoint and system-level online diagnosis to exclude the faulty nodes from the system. The most important problem of this system is synchronization, diagnosis and reconfiguration. This paper describes how to reduce the cost of system synchronization and diagnosis to eliminate the major performance bottleneck in this fault tolerance system. The logical clock replaces the real clock in the system synchronization, and it overcomes the fundamental limitation of the systemlevel diagnosis algorithm. The algorithm uses only O(Nt) messages to detect the faulty nodes in the system where arbitrary faulty nodes exist.
| Original language | English |
|---|---|
| Pages (from-to) | 21-23 |
| Number of pages | 3 |
| Journal | Chinese Journal of Electronics |
| Volume | 10 |
| Issue number | 1 |
| State | Published - 2001 |
Keywords
- Checkpoint
- Fault tolerance
- System-level diagnosis
Fingerprint
Dive into the research topics of 'The design and implementation of HFTC_2 - A dependable distributed computer system based diagnosis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver