Medical Provider Embeddings: Moving Beyond One-Hot Encoding in Healthcare Fraud Detection

Medical Provider Embeddings for Healthcare Fraud Detection

2021-05-15
Justin M. Johnson, Taghi M. Khoshgoftaar
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces four embedding techniques (GloVe, Med-W2V, HcpcsVec, and RxVec) to convert medical provider specialties into dense semantic vectors for healthcare fraud detection. By leveraging large-scale Medicare Part B and Part D datasets, the authors achieve state-of-the-art results using classifiers like Gradient Boosted Trees (GBT) and Multilayer Perceptrons (MLP).

TL;DR

Healthcare fraud costs the U.S. Medicare program up to $79 billion annually. This paper demonstrates that how we represent a medical provider's "Specialty" is critical for detection. By abandoning sparse One-Hot Encodings in favor of dense Semantic Embeddings (HcpcsVec/RxVec) derived from actual billing patterns, the authors achieved an AUC of 0.873, setting a new benchmark for the field.

Deep Dive into the Motivation

In healthcare informatics, "Medical Provider Type" (e.g., Internal Medicine, Urology) is a high-cardinality categorical variable. Historically, researchers faced a dilemma:

  1. Exclude the feature: Afraid of adding noise.
  2. One-Hot Encoding: Creating a massive, sparse matrix where every specialty is equidistant.

The second approach is mathematically "blind." In a one-hot space, a Certified Registered Nurse Anesthetist is as different from an Anesthesiologist as they are from a Podiatrist. The authors' core insight is that similarity in behavior implies similarity in specialty. If two providers perform the same set of HCPCS procedures or prescribe the same drugs, their embeddings should be close in latent space.

Methodology: The Birth of HcpcsVec and RxVec

While the paper evaluates pre-trained NLP embeddings (GloVe and Med-W2V), the real innovation lies in the data-driven embeddings:

1. The Construction Pipeline

The authors built a transformation pipeline that converts raw claims into specialty "signatures":

  • Aggregation: Summing procedure (HCPCS) or drug occurrences per provider.
  • Normalization: Scaling by maximum occurrence to focus on relative frequency rather than volume.
  • Mean Centroid: Averaging provider vectors within a specialty to create a "Specialty-Procedure Matrix."
  • Dimensionality Reduction: Using PCA to compress these high-dimensional signatures into dense vectors (32, 64, or 128-D).

Methodological Pipeline for Creating Embeddings

Experimental Evidence & Results

The researchers tested four learners: Logistic Regression (LR), Random Forest (RF), GBT, and MLP.

The GBT Breakthrough

The Gradient Boosted Tree (GBT) emerged as the strongest performer. Interestingly, the model's Feature Importance score for "Provider Type" spiked when using semantic embeddings. This proves that the semantic representation makes the feature "readable" for the tree splitting criteria, whereas one-hot features are often ignored due to sparsity.

Performance Comparison on Part B and Part D

Visualizing the Semantic Space

Using t-SNE, the authors validated their "Physical Intuition." In the RxVec (prescription-based) embedding space, dental providers clustered together, as did those involved in childbirth (Midwives, OB/GYN). This confirms that the embeddings successfully captured the medical reality of provider practices without any manual "expert" labeling of the specialties.

t-SNE Visualization of RxVec Embeddings

Critical Analysis & Future Outlook

Takeaway: This work proves that "Behavioral Signatures" (what you do/prescribe) are superior to "Textual Signatures" (how your specialty is named) for fraud detection.

Limitations:

  1. Static Logic: The embeddings are calculated as a pre-processing step (static) rather than being learned end-to-end (dynamic) like modern Entity Embedding layers in PyTorch/TensorFlow.
  2. Temporal Drift: Procedure patterns change over time (e.g., new COVID-19 codes); these embeddings would need periodic re-training.

Future Work: The logical next step is to integrate these embeddings into a Graph Neural Network (GNN) where providers, patients, and procedures form a tripartite graph, allowing the model to capture not just specialty similarity, but also the "guilt by association" often found in fraudulent billing rings.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) to model the relationship between medical providers, procedures (HCPCS), and fraud labels in Medicare data.
  • Which study first introduced the concept of "Entity Embeddings" for categorical variables in tabular data, and how does this paper's HcpcsVec approach differ from automated entity embedding layers in neural networks?
  • Identify research that applies Transformer-based contextual embeddings (like ClinicalBERT or BEHRT) specifically to the task of medical provider behavioral modeling for anomaly detection.
Contents
Medical Provider Embeddings: Moving Beyond One-Hot Encoding in Healthcare Fraud Detection
1. TL;DR
2. Deep Dive into the Motivation
3. Methodology: The Birth of HcpcsVec and RxVec
3.1. 1. The Construction Pipeline
4. Experimental Evidence & Results
4.1. The GBT Breakthrough
4.2. Visualizing the Semantic Space
5. Critical Analysis & Future Outlook