TY - GEN
T1 - R-ADMAD
T2 - 23rd International Conference on Supercomputing, ICS'09
AU - Liu, Chuanyi
AU - Gu, Yu
AU - Sun, Linchun
AU - Yan, Bin
AU - Wang, Dongsheng
PY - 2009
Y1 - 2009
N2 - Data de-duplication has become a commodity component in data-intensive systems and it is required that these systems provide high reliability comparable to others. Unfortunately, by storing duplicate data chunks just once, de-duped system improves storage utilization at cost of error resilience or reliability. In this paper, R-ADMAD, a high reliability provision mechanism is proposed. It packs variable-length data chunks into fixed sized objects, and exploits ECC codes to encode the objects and distributes them among the storage nodes in a redundancy group, which is dynamically generated according to current status and actual failure domains. Upon failures, R-ADMAD proposes a distributed and dynamic recovery process. Experimental results show that R-ADMAD can provide the same storage utilization as RAID-like schemes, but comparable reliability to replication based schemes with much more redundancy. The average recovery time of R-ADMAD based configurations is about 2-6 times less than RAID-like schemes. Moreover, R-ADMAD can provide dynamic load balancing even without the involvement of the overloaded storage nodes.
AB - Data de-duplication has become a commodity component in data-intensive systems and it is required that these systems provide high reliability comparable to others. Unfortunately, by storing duplicate data chunks just once, de-duped system improves storage utilization at cost of error resilience or reliability. In this paper, R-ADMAD, a high reliability provision mechanism is proposed. It packs variable-length data chunks into fixed sized objects, and exploits ECC codes to encode the objects and distributes them among the storage nodes in a redundancy group, which is dynamically generated according to current status and actual failure domains. Upon failures, R-ADMAD proposes a distributed and dynamic recovery process. Experimental results show that R-ADMAD can provide the same storage utilization as RAID-like schemes, but comparable reliability to replication based schemes with much more redundancy. The average recovery time of R-ADMAD based configurations is about 2-6 times less than RAID-like schemes. Moreover, R-ADMAD can provide dynamic load balancing even without the involvement of the overloaded storage nodes.
KW - Data de-duplication
KW - Error correcting code
KW - Reliability
UR - https://www.scopus.com/pages/publications/70450060162
U2 - 10.1145/1542275.1542327
DO - 10.1145/1542275.1542327
M3 - 会议稿件
AN - SCOPUS:70450060162
SN - 9781605584980
T3 - Proceedings of the International Conference on Supercomputing
SP - 370
EP - 379
BT - ICS'09 - Proceedings of the 23rd International Conference on Supercomputing
Y2 - 8 June 2009 through 12 June 2009
ER -