Cost-Effective Machine Learning for Automatically Processing Bibliographic

Loading...
Thumbnail Image

Authors

Huskey, Samuel J.

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Many digital humanities projects involve tedious and repetitive tasks take time away from the higher-level tasks further down the pipeline that require intelligent decision-making. When funding is available, the tedious and repetitive tasks are often assigned to research assistants, but when funding is scarce, those tasks tend to create bottlenecks that either impede progress or halt it altogether. This paper argues that artificial intelligence and machine learning tools and techniques are worth exploring as cost-effective, accessible solutions to these problems. The Digital Latin Library project provides a case study through its experiments with fine-tuning pretrained transformer language models to process noisy bibliographic metadata. The results show that the models have potential for accelerating this tedious task. But the experiments also had an unexpected, yet positive outcome: the models revealed gaps in the catalog's coverage, helping to focus the efforts of the human experts working on the project.

Description

This article discusses how artificial intelligence can be used for routine, repetitive tasks like processing thousands of bibliographic citations in different formats. Using AI in this way, scholars can free up precious time for focusing on higher-level research tasks.

Citation

Related file

Notes

Endorsement

Review

Supplemented By

Referenced By

Collection Detail

# of Isolates from RBM

# of Isolates from TV8