MEDICAL CLASSIFICATION THROUGH MACHINE LEARNING, DOMAIN-SPECIFIC LANGUAGE MODELS AND GENERATIVE AI TECHNIQUES

Loading...
Thumbnail Image

Date

Authors

Gheibi, Reza

Journal Title

Journal ISSN

Volume Title

Publisher

University of Oklahoma – Graduate College

Item Statistics

  • Total Views: 17
  • Total Downloads: 34
  • Views in the Last Month: 11

Abstract

Advancing reliable machine learning and natural language processing methods for classification under data constrained conditions remains a central challenge in computer science. This dissertation addresses core problems related to data scarcity, class imbalance, heterogeneous tabular features, and the integration of structured and unstructured representations. To investigate these challenges, cervical cancer classification is used as an application domain where limited, noisy, and imbalanced datasets make model development particularly difficult. The work first evaluates traditional machine learning models including logistic regression, decision trees, random forests, gradient boosting, and multilayer perceptrons under varying data quality conditions. Multiple data augmentation strategies are examined, such as SMOTE, generative adversarial networks, diffusion based models, large language model, domain specific small language models, and synthetic tabular data generation, demonstrating that appropriate augmentation can substantially improve robustness and predictive performance. The second component introduces a framework that transforms structured medical records into natural language representations, allowing the use of modern language models for classification. Domain-specific and general purpose transformers (e.g., PubMedBERT, BioBERT, GPT-2, LLaMA, Mistral) are fine-tuned and evaluated in supervised, zero-shot, and few-shot settings. The results show that text-based modeling can match or exceed traditional ML performance, particularly in low-resource scenarios, while capturing richer contextual information. Finally, the dissertation extends this language-based methodology to a large-scale clinical taxonomy task. Through unsupervised analysis and systematic refinement, ambiguous specialty labels are consolidated into coherent Super-Specialties, and a fine-tuned PubMedBERT classifier achieves strong multi-label performance. This demonstrates the scalability of the proposed methods for organizing complex clinical data. In general, this dissertation demonstrates that combining structured data modeling, generative data augmentation, and language-based representations provides a flexible and effective framework for cervical cancer classification under data-constrained conditions. The findings highlight the potential of these approaches to support more accessible, scalable, and robust clinical decision-support systems, with broader implications for medical classification tasks beyond cervical cancer.

Description

Citation

Related file

Notes

Endorsement

Review

Supplemented By

Referenced By

DOI

Collection Detail

# of Isolates from RBM

# of Isolates from TV8