TY - GEN
T1 - Co-clustering of single-cell RNA-seq data based on weighted non-negative matrix tri-factorization combined with consensus clustering
AU - Ren, Tongtong
AU - Wang, Guohua
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Single-cell RNA sequencing (scRNA-seq) has the ability to accurately identify cell types contained in tissues at single-cell resolution. However, the research of identifying imbalanced cell types using scRNA-seq data is still challenging. In this article, we propose scCO2, a method based on weighted non-negative matrix tri-factorization (NMTF) combined with Kmeans-based consensus clustering for the co-clustering of scRNA-seq data. Compared with several popular methods on six real scRNA-seq data with known cell types, scCO2 can achieve comparable or superior cell clustering performance to selected clustering methods. Through the case study on a real scRNA-seq dataset of the human pancreas, scCO2 could obtain good correspondence between gene clusters and cell clusters. Additionally, scCO2 also shows the ability to identify rare cell types. Moreover, by comparing the gene sets from gene clusters to existing known marker genes, we demonstrate that scCO2 has the potential to identify more underlying cell-type-specific genes, and the weights of genes learned by scCO2 could be used as the indicator of gene importance.
AB - Single-cell RNA sequencing (scRNA-seq) has the ability to accurately identify cell types contained in tissues at single-cell resolution. However, the research of identifying imbalanced cell types using scRNA-seq data is still challenging. In this article, we propose scCO2, a method based on weighted non-negative matrix tri-factorization (NMTF) combined with Kmeans-based consensus clustering for the co-clustering of scRNA-seq data. Compared with several popular methods on six real scRNA-seq data with known cell types, scCO2 can achieve comparable or superior cell clustering performance to selected clustering methods. Through the case study on a real scRNA-seq dataset of the human pancreas, scCO2 could obtain good correspondence between gene clusters and cell clusters. Additionally, scCO2 also shows the ability to identify rare cell types. Moreover, by comparing the gene sets from gene clusters to existing known marker genes, we demonstrate that scCO2 has the potential to identify more underlying cell-type-specific genes, and the weights of genes learned by scCO2 could be used as the indicator of gene importance.
KW - co-clustering
KW - consensus clustering
KW - imbalanced cell types
KW - non-negative matrix tri-factorization
KW - single-cell RNA sequencing
UR - https://www.scopus.com/pages/publications/85184900446
U2 - 10.1109/BIBM58861.2023.10385799
DO - 10.1109/BIBM58861.2023.10385799
M3 - 会议稿件
AN - SCOPUS:85184900446
T3 - Proceedings - 2023 2023 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2023
SP - 183
EP - 190
BT - Proceedings - 2023 2023 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2023
A2 - Jiang, Xingpeng
A2 - Wang, Haiying
A2 - Alhajj, Reda
A2 - Hu, Xiaohua
A2 - Engel, Felix
A2 - Mahmud, Mufti
A2 - Pisanti, Nadia
A2 - Cui, Xuefeng
A2 - Song, Hong
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2023 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2023
Y2 - 5 December 2023 through 8 December 2023
ER -