Biodiv Sci ›› 2026, Vol. 34 ›› Issue (7): 0.  DOI: 10.17520/biods.2026044

    Next Articles

Molecular Identification of Triplophysa Species Using Cyt b Gene-Based Machine Learning Methods

Rigui Yi1,2, Delin Qi1, Shihai Zhu2, Tingting Yang1,2, Dan Liu1, Gao Qiang1, Mingzhe Xia1, Wei Wang1, Miaomiao Nie1, Junmei Jia1, Cunfang Zhang1*   

  1. 1. State Key Laboratory of Plateau Ecology and Agriculture, Qinghai University, Xining 810016, China 

    2. College of Eco‑Environmental Engineering, Qinghai University, Xining 810016, China

  • Received:2026-02-08 Revised:2026-06-12 Online:2026-07-20
  • Contact: Zhang, Cunfang
  • Supported by:
    Innovative Research on First-Class Discipline of Ecology Supported by Qinghai University(2026-ST-05); Chief Scientist Program of Qinghai Province(2024-SF-102)

Abstract:

Aims: The genus Triplophysa is a key component of the fish fauna of the Qinghai-Tibet Plateau and its adjacent regions. Traditional morphological classification and conventional molecular identification encounter considerable challenges attributable to convergent evolution, phenotypic plasticity, and other confounding factors, thereby posing substantial obstacles to precise species delimitation. This study systematically compared the application efficacy of phylogenetic approaches and machine learning techniques based on the mitochondrial cytochrome b (Cyt b) gene in the molecular identification of seven Triplophysa fishes, aiming to establish a reliable technical system for species delimitation of this genus and further provide critical technical support for its biodiversity conservation and quantitative resource assessment. 

Methods: In this study, we integrated 395 newly generated Cyt b gene sequences (experimentally obtained) and 247 sequences retrieved from the NCBI database, constructing a comprehensive dataset comprising 642 sequences across seven Triplophysa species. It systematically compared the identification efficacy of phylogenetic methods (NJ/ML/BI trees, GMYC/bPTP, K2P genetic distance) with machine learning approaches (BLOG and WEKA platform classifiers including SMO, J48, JRip and Naïve Bayes). 

Results: The Cyt b gene sequences exhibited substantial genetic variation, with a haplotype diversity (Hd) of 0.975±0.003 and a nucleotide diversity (π) of 0.107±0.007, as well as a significant A+T base bias (56.5%). These characteristics fully fulfilled the core requirement of genetic polymorphism for molecular identification markers. Due to the interference of recent divergence time, interspecific gene flow, and hybridization events among closely related species, phylogenetic approaches and species delimitation models showed limited performance in delimiting closely related species, only effectively distinguishing 5 out of the 7 studied species, accounting for 62% of all haplotypes. In contrast, machine learning techniques demonstrated superior overall performance: the SMO classifier achieved 100% accuracy in species identification, the BLOG algorithm yielded a correct classification rate of 96.23% on the test set, and the classification accuracies of J48, JRip, and Naïve Bayes classifiers all exceeded 98%. Furthermore, these machine learning methods could identify species-specific diagnostic sites and generate interpretable classification rules. 

Conclusion: This study demonstrates that machine learning methods, characterized by high accuracy and efficiency, can serve as reliable auxiliary tools for the accurate identification of the seven Triplophysa species examined in this study. These methods provide technical support for the biodiversity conservation and resource assessment of these species, and simultaneously offer new perspectives and technical references for the molecular identification of other fish groups facing morphological discrimination challenges.

Key words: Triplophysa, machine learning, Cyt b gene, species identification