Supervised machine learning approaches for phylogenetic analyses

Loading...
Thumbnail Image

Date

Authors

Rosas Puchuri, Ulises

Journal Title

Journal ISSN

Volume Title

Publisher

University of Oklahoma – Graduate College

Item Statistics

  • Total Views: 10
  • Total Downloads: 43
  • Views in the Last Month: 3

Abstract

The following dissertation adds a new level of realism to multiple phylogenetic analyses by using supervised machine learning algorithms, which focus on prediction rather than data pattern discovery. Supervised machine learning has recently gained popularity for its application in large language models. However, applying these ideas to improve performance in phylogenetics remains underexplored. Through three chapters, the present dissertation showcases improvements across different phylogenetic analyses using supervised machine learning algorithms. Chapter 2 addresses the scalability problem in likelihood-based phylogenetic network inference. As genomic datasets continue to grow, phylogenetic network inference faces serious scalability challenges, since the input size increases quartically with respect to the number of species. This chapter proposes a new subsampling algorithm tailored to large datasets, in which subsampling is guided by a sparse machine learning model. To the best of our knowledge, this is the first time a sparse non-parametric model has been used for data subsampling in phylogenetics. Previous approaches for phylogenetic trees relied on parametric models and required the subsample size to be fixed in advance. Chapter 2 shows that both assumptions can be relaxed, allowing the data to determine the optimal subsample size as a function of model complexity—restricted to at most four levels—without compromising estimation accuracy. Chapter 3 tackles the parametric assumptions commonly made in phylogenetic regression. These assumptions break down, for example, when there are complex interactions among species traits or when the response function is more complex than a linear relationship. This chapter reformulates standard phylogenetic regression as kernel ridge regression, a non-parametric machine learning method, such that accommodates both correlation among observations and nonlinear effects. This approach is particularly useful in the subfield of phylogenetic comparative methods, where the correlation structure—phylogenetic relatedness among species derived from a phylogenetic tree—is known. Through simulations and analyses of empirical datasets, this model outperforms standard phylogenetic regression. It also provides an initial framework for non-parametric regression with multispecies data while explicitly accounting for phylogenetic structure. Chapter 4 explores factors influencing error in standard phylogenetic tree inference using deep learning. A deep learning model is trained to predict support for species trees based on a comprehensive set of features typically associated with methodological error. Deep learning outperforms simple linear models and enables model interrogation through feature importance analyses, revealing the dominant factors affecting inference in a given dataset. These insights allow for more informed choices regarding evolutionary models and data preprocessing strategies. The intersection of phylogenetics and machine learning is an active area of research, and this dissertation contributes to this growing field. Throughout the dissertation, we demonstrate that machine learning methods can both improve phylogenetic analyses and provide insight into their behavior. Conversely, incorporating evolutionary information can enhance the performance of machine learning models when analyzing data from multiple species.

Description

Citation

Related file

Notes

Endorsement

Review

Supplemented By

Referenced By

DOI

Collection Detail

# of Isolates from RBM

# of Isolates from TV8