TY - GEN
T1 - A bucket index correction based method for compression of genomic sequencing data
AU - Wang, Rongjie
AU - Bai, Yang
AU - Cheng, Qianlong
AU - Zang, Tianyi
AU - Wang, Yadong
N1 - Publisher Copyright:
© 2017 IEEE.
PY - 2017/12/15
Y1 - 2017/12/15
N2 - As high-throughput sequencing technologies are generating vast amounts of data, there is urgent need to develop efficient algorithms for sequencing data compression. Existing methods usually dispatch the similar sequences into the same bucket based on their same minimizer, that is the lexicographical smallest k-mer within the sequence, for data compression. However, when the sequencing error existed in the minimizer area, it could cause sequences to be distributed into the improper buckets, which could result in a negative effect in the following compression process. In this paper, we propose a novel method BIC, a bucket index correction method for sequencing data compression. BIC is the first method to correct sequencing errors in minimizer area, which dispatches more similar sequences into the same buckets, that could effectively compress sequencing data. Compared with three state-of-the-art methods on five different data sets, BIC could reach more compression rate. The codes of BIC are available at https://github.com/rongjiewang/BIC.
AB - As high-throughput sequencing technologies are generating vast amounts of data, there is urgent need to develop efficient algorithms for sequencing data compression. Existing methods usually dispatch the similar sequences into the same bucket based on their same minimizer, that is the lexicographical smallest k-mer within the sequence, for data compression. However, when the sequencing error existed in the minimizer area, it could cause sequences to be distributed into the improper buckets, which could result in a negative effect in the following compression process. In this paper, we propose a novel method BIC, a bucket index correction method for sequencing data compression. BIC is the first method to correct sequencing errors in minimizer area, which dispatches more similar sequences into the same buckets, that could effectively compress sequencing data. Compared with three state-of-the-art methods on five different data sets, BIC could reach more compression rate. The codes of BIC are available at https://github.com/rongjiewang/BIC.
UR - https://www.scopus.com/pages/publications/85045988464
U2 - 10.1109/BIBM.2017.8217727
DO - 10.1109/BIBM.2017.8217727
M3 - 会议稿件
AN - SCOPUS:85045988464
T3 - Proceedings - 2017 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2017
SP - 634
EP - 637
BT - Proceedings - 2017 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2017
A2 - Yoo, Illhoi
A2 - Zheng, Jane Huiru
A2 - Gong, Yang
A2 - Hu, Xiaohua Tony
A2 - Shyu, Chi-Ren
A2 - Bromberg, Yana
A2 - Gao, Jean
A2 - Korkin, Dmitry
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2017 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2017
Y2 - 13 November 2017 through 16 November 2017
ER -