| Keywords [eng] |
microbiome, differential abundance analysis, 16S rRNA data, compositional data, DAA method benchmark, mikrobiomas, diferencinio gausumo analizė, 16S rRNR duomenys, kompoziciniai duomenys, diferencinio gausumo analizės metodų palyginimas |
| Abstract [eng] |
Differential abundance analysis identifies microbial features whose abundance shifts are statistically associated with conditions of interest. However, sequencing count data has strong technical confounding, compositionality, zero inflation, and library size heterogeneity, which complicate inference. Numerous differential abundance analysis tools have been developed, but no single method controls error and maintains power across realistic data settings, motivating new procedures. This thesis presents PURSUE, a reference-normalized two-part regression framework with permutation-based inference. PURSUE was benchmarked against eight DAA methods (ALDEx2, ANCOM-BC2, corncob, LDM, LinDA, LOCOM, MaAsLin3, ZicoSeq) using the SimulateMSeq simulation framework across null and non-null simulation data settings spanning sample size, taxa diversity, number of taxa, signal density, sequencing depth confounding, differential taxa rarity, and confounder adjustment. Under null settings, PURSUE remained well calibrated, with mean false positive rate well below the nominal 0.05 threshold in every setting. Under non-null settings, PURSUE showed high mean precision and intermediate recall, making it precision-favored, with a moderate F1 score and elevated false discovery rate relative to the most conservative methods. PURSUE performed strongest in depth confounded settings, ranking second after ZicoSeq for precision. Low diversity and small sample size settings were its main weaknesses. These findings suggest that PURSUE can serve as a high precision, depth-specialized tool, while highlighting which design components contribute to its strengths and weaknesses, providing a basis for further development. |