Skip to main navigation Skip to search Skip to main content

The design and implementation of HFTC_2 - A dependable distributed computer system based diagnosis

  • Decheng Zuo*
  • , Xiaozong Yang
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

HFTC_2 is a practical dependable distributed computer system, it works correctly in the existence of arbitrary fault combination with checkpoint and system-level online diagnosis to exclude the faulty nodes from the system. The most important problem of this system is synchronization, diagnosis and reconfiguration. This paper describes how to reduce the cost of system synchronization and diagnosis to eliminate the major performance bottleneck in this fault tolerance system. The logical clock replaces the real clock in the system synchronization, and it overcomes the fundamental limitation of the systemlevel diagnosis algorithm. The algorithm uses only O(Nt) messages to detect the faulty nodes in the system where arbitrary faulty nodes exist.

Original languageEnglish
Pages (from-to)21-23
Number of pages3
JournalChinese Journal of Electronics
Volume10
Issue number1
StatePublished - 2001

Keywords

  • Checkpoint
  • Fault tolerance
  • System-level diagnosis

Fingerprint

Dive into the research topics of 'The design and implementation of HFTC_2 - A dependable distributed computer system based diagnosis'. Together they form a unique fingerprint.

Cite this