Cost-Effective Machine Learning for Automatically Processing Bibliographic Metadata

Loading...
Thumbnail Image

Authors

Huskey, Samuel J.

Journal Title

Journal ISSN

Volume Title

Publisher

Edinburgh University Press

Item Statistics

  • Total Views: 0
  • Total Downloads: 2
  • Views in the Last Month: 0

Abstract

Many digital humanities projects involve tedious and repetitive tasks take time away from the higher-level tasks further down the pipeline that require intelligent decision-making. When funding is available, the tedious and repetitive tasks are often assigned to research assistants, but when funding is scarce, those tasks tend to create bottlenecks that either impede progress or halt it altogether. This paper argues that artificial intelligence and machine learning tools and techniques are worth exploring as cost-effective, accessible solutions to these problems. The Digital Latin Library project provides a case study through its experiments with fine-tuning pretrained transformer language models to process noisy bibliographic metadata. The results show that the models have potential for accelerating this tedious task. But the experiments also had an unexpected, yet positive outcome: the models revealed gaps in the catalog's coverage, helping to focus the efforts of the human experts working on the project.

Description

Citation

Huskey, S. J. (2025). Cost-effective machine learning for automatically processing bibliographic metadata. International Journal of Humanities and Arts Computing, 19(2), 112–126. https://doi.org/10.3366/ijhac.2025.0353

Related file

https://www.euppublishing.com/doi/10.3366/ijhac.2025.0353

Notes

© Edinburgh University Press 2025.

Endorsement

Review

Supplemented By

Referenced By

Collection Detail

# of Isolates from RBM

# of Isolates from TV8