Impact of dataset partitioning on the reliability of SERS–machine learning models for antimicrobial resistance classification
Analytica Chimica Acta, cilt.1423, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 1423
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.aca.2026.346185
- Dergi Adı: Analytica Chimica Acta
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Artic & Antarctic Regions, BIOSIS, Chemical Abstracts Core, Chimica, Compendex, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO)
- Erciyes Üniversitesi Adresli: Evet
Özet
Antimicrobial resistance (AMR) represents a major global health challenge and motivates the development of rapid analytical approaches for characterizing clinical bacterial isolates. Surface-enhanced Raman spectroscopy (SERS), combined with machine learning (ML), has emerged as a promising approach for label-free spectral classification of bacterial isolates. However, the field lacks methodological consensus on how hierarchical spectral datasets should be structured, standardized, and partitioned for reliable model evaluation. This gap has led to widespread data leakage and inflated performance reports in SERS–AI studies. Here, we present a systematic benchmarking framework based on 15,000 SERS spectra acquired from 15 clinical Staphylococcus aureus isolates representing distinct resistance phenotypes (MRSA, ERSA, and SSA), evaluating twelve data-partitioning strategies across five machine learning models and three feature selection or dimensionality-reduction methods. Our results show that spectrum-level splits consistently overestimate classification accuracy due to non-independent samples, whereas isolate- or subject-level partitioning provides more realistic estimates of model generalizability. Among the evaluated models, random forest combined with Boruta feature selection produced consistently robust and interpretable performance. Overall, model performance depended more strongly on dataset organization and partitioning strategy than on the choice of machine learning algorithm.