Parameter-Free Loss for Imbalanced Classification in Medicare Fraud Detection

Loading...
Thumbnail Image

Date

Authors

Colindres, David

Journal Title

Journal ISSN

Volume Title

Publisher

University of Oklahoma – Graduate College

Item Statistics

  • Total Views: 16
  • Total Downloads: 132
  • Views in the Last Month: 3

Abstract

Healthcare fraud costs the U.S. Medicare system over $100 billion annually; yet, fraudulent providers represent only a small share of all claims, creating a severe class imbalance that conventional classifiers overlook. This thesis presents one of the first applications of Parameter-Free Imbalance-Aware Loss (PF-Loss) for deep learning to Medicare fraud detection and evaluates it alongside eleven other supervised models on a publicly available CMS dataset containing 5,410 providers, 558,211 claims, and 163 engineered features. The study benchmarks eight traditional and ensemble machine learning classifiers and four tabular ResNet variants, each distinguished by its loss function: PF-Loss (PF), Weighted Cross-Entropy (WCE), Focal Loss (FL), and Weighted Label-Distribution-Aware Margin Loss (WLDAM). All models are evaluated under a unified experimental framework with G-mean as the primary performance metric. PF-Loss achieved the highest G-mean among all models (86.66%), outperforming the strongest traditional classifier, XGBoost (85.26%), as well as the other neural network variants. All four ResNet loss variants exceeded the traditional models in G-mean, which indicates that deep learning combined with imbalance-aware loss design is effective for tabular healthcare fraud detection. PF-Loss completed training in 19.7 seconds, making it 3.7 to 11.5 times faster than the other neural network variants, because its parameter-free formulation required only three architecture configurations, compared with 12, 27, and 36 for WCE, Focal Loss, and WLDAM, respectively. A systematic class imbalance analysis further showed that traditional models deteriorated sharply under the original no-resampling condition, whereas PF-Loss maintained a G-mean of 86.66% without external resampling; combining PF-Loss with SMOTE improved performance only marginally to 86.74%. SHAP-based interpretability analysis identified provider-level and physician-level reimbursement amounts, claim counts, admission duration, and deductible amounts as the most consistent fraud indicators across model families, indicating that fraud predictions were driven primarily by billing behavior rather than patient demographics. This thesis also extends classification outputs into a quantitative risk framework based on Basel II/III credit risk methodology, including probability calibration, expected loss estimation, budget-constrained audit optimization formulated as a 0–1 knapsack problem, and Monte Carlo simulation for tail-risk assessment. The results from this study show that PF-Loss is an effective and efficient alternative to conventional imbalance strategies, especially for tabular healthcare fraud detection. The model offers strong predictive performance, a significantly lower tuning burden, and practical application for resource-constrained fraud investigation units. Keywords: Medicare fraud detection, Parameter-free loss, Class imbalance, Tabular ResNet, G-mean, SHAP interpretability, Cost-sensitive learning, Deep learning, Quantitative risk analysis

Description

Citation

Related file

Notes

Collections

Endorsement

Review

Supplemented By

Referenced By

DOI

Collection Detail

# of Isolates from RBM

# of Isolates from TV8