Abstract
The RNA-Seq-based transcriptome sequencing data has a high feature dimension that requires a lot of computing resources when using traditional methods to find phenotype related genes. Moreover, the range of candidate genes obtained by difference analysis is large, and further screening depends on existing a prior knowledge. A transcrip-tome analysis method combining genetic algorithm and XGBoost, GA-XGBoost, was proposed to narrow the range of candidate genes for subsequent analysis by incorporating machine learning algorithm. A comparative experiment and subsequent analysis of the gene-100-kernel weight trait association on a set of high-quality maize datasets showed that, compared with training the XGBoost model directly with whole genes and differentially expressed genes, the candidate gene training XGBoost model obtained by the proposed method had the minimum MSE in predicting the 100-kernel weight of maize. Compared with 1542 differentially expressed genes in the results of differential expression analysis, the range of candidate genes was reduced to 48 by the GA-XGBoost method, which was reduced by 31 times, indicating that the proposed method could effectively improve the ability and efficiency of transcriptome data analysis.
| Original language | English |
|---|---|
| Pages (from-to) | 170-180 |
| Number of pages | 11 |
| Journal | CAAI Transactions on Intelligent Systems |
| Volume | 17 |
| Issue number | 1 |
| DOIs | |
| State | Published - Jan 2022 |
| Externally published | Yes |
Keywords
- 100-kernel weight
- eXtreme gradient boosting
- gene ontology
- genetic algorithm
- kyoto encyclopedia of genes and genomes
- machine learning
- maize
- transcriptome analysis
Fingerprint
Dive into the research topics of 'The method of 100-kernel weight related genes mining in maize mixed with genetic algorithm and XGboost'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver