PREDICTING STELLAR AGES FROM CHEMICAL ABUNDANCES WITH XGBOOST: A MACHINE LEARNING APPROACH TO GALACTIC ARCHAEOLOGY
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Estimating stellar ages remains a central challenge in astrophysics, particularly for stars outside the narrow evolutionary phases where traditional methods like isochrone fitting are reliable. This study explores the use of supervised machine learning, specifically, the Extreme Gradient Boosting (XGBoost) algorithm, to infer stellar ages from chemical abundance ratios derived from the GALAH DR4 spectroscopic survey. A training set of 20,303 Main Sequence Turn-Off stars was constructed using strict quality cuts and age labels obtained via Bayesian isochrone fitting. The XGBoost model was trained on 15 elemental abundance ratios and tuned through exhaustive hyperparameter optimization with five-fold cross-validation. The final model achieves a root mean squared error (RMSE) of 1.38 Gyr on a held-out test set, generalizes well across stellar populations not seen during training, and reproduces known chemical evolution trends such as the anti-correlation between [Fe/H] and age and the positive correlation between [$\alpha$/Fe] and age. Feature importance analysis reveals that [Mg/Fe], [Ca/Fe], and [Y/Fe] are the most informative predictors of age, consistent with theoretical nucleosynthetic expectations. Residual analyses expose a tendency for the model to regress predictions toward the mean age in underrepresented regions, particularly for very young and very old stars. These results demonstrate that machine learning can extract robust age information from chemical abundance patterns and extend age-dating to stellar populations for which traditional techniques break down. This work highlights the power of data-driven approaches in galactic archaeology and paves the way for scalable age estimation in next-generation spectroscopic surveys.