Beyond Co-occurrence: Graph Embeddings for Precision Overtreatment Detection

Abuse detection in healthcare insurance with disease-treatment network embedding

2021-10-19
Jehyuk Lee, Sungzoon Cho
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a three-stage network-based approach for detecting healthcare insurance abuse (overtreatment). It constructs a disease-treatment association network using Relative Risk (RR) and employs metapath2vec graph embedding for link prediction, achieving superior performance in identifying unnecessary medical treatments on real-world HIRA data.

    ## TL;DR
    Overtreatment in healthcare is a multi-billion dollar drain on global insurance systems. This paper moves beyond traditional "scoring" models by building a **disease-treatment association network**. By using **metapath2vec** to embed these entities into a shared vector space, the authors can predict whether a prescription is medically "logical" for a given diagnosis, significantly outperforming traditional rule-based or co-occurrence-based detection systems.

    ## The Core Intuition: Why Simple "Co-occurrence" Fails
    In the world of medical billing, a single claim often lists multiple diseases and multiple treatments. A naive machine learning model might assume that *every* treatment in that claim is linked to *every* disease mentioned. 

    This creates "noisy" data. For example, a patient might have a neck fracture (main disease) and a minor skin rash (sub-disease). If a doctor prescribes lumbar spine imaging, a naive model might incorrectly learn that neck fractures or rashes justify lower-back scans. 

    The authors' key insight is to use **Relative Risk (RR)** to filter these links. A connection only exists in their "Association Network" if the treatment is statistically significantly more likely to occur in the presence of that disease across the entire dataset.

    ## Methodology: The Three-Stage Framework

    ### 1. Constructing the Association Network
    The network is heterogeneous, comprising Refined Diagnosis-Related Groups (RDRG), main diseases, sub-diseases, and three types of treatments: Procedures, Prescriptions, and Materials.
    
    ![Association vs Co-occurrence Network](https://cdn.atominnolab.com/wisdoc/images/20260608-1e2fb6aa-5e23-4fa0-a7f9-edbe7c38cf16/page_008_block_015.png)
    *As shown above, the association network (right) is much sparser and more clinically accurate than the dense, noisy co-occurrence network (left).*

    ### 2. Selecting the Best Graph Embedding
    The paper evaluates several SOTA graph embedding methods (DeepWalk, node2vec, SDNE, etc.). However, because the network contains different *types* of nodes, **metapath2vec**—designed for heterogeneous information networks—emerged as the winner. It uses specific "meta-paths" (e.g., Treatment-Disease-Treatment) to capture the fact that two different drugs are similar if they are used to treat the same disease.

    ### 3. Overtreatment Detection via Link Prediction
    Overtreatment is redefined as a **Link Prediction problem**. If the trained model predicts a "no-link" status between a treatment and *every* disease listed in a claim, that treatment is flagged as an abuse case.

    ![Overtreatment Detection Framework](https://cdn.atominnolab.com/wisdoc/images/20260608-1e2fb6aa-5e23-4fa0-a7f9-edbe7c38cf16/page_006_block_004.png)

    ## Experimental Results: Precision Matters
    Testing on South Korea's HIRA (Health Insurance Review and Assessment) 2017 data, the model showed massive gains over the "Without Embedding" baseline:

    *   **Procedures**: Accuracy reached up to **97.6%**.
    *   **Prescriptions**: Accuracy peaked at **93.8%**.
    
    The model proved particularly robust at identifying "unseen" patterns—cases where a specific disease-treatment pair was not in the training set, but the model could "infer" its legitimacy based on the similarity of the nodes in the latent embedding space.

    ## Critical Analysis & Conclusion
    The value of this work lies in its **Inductive Bias**. By forcing the model to learn medical logic (the relationship between a diagnosis and an action) via a graph structure, it becomes much harder for fraudulent providers to "hide" abnormal billing patterns beneath complex claim filings.

    **Limitations**: The model currently treats overtreatment as a binary "necessary vs. unnecessary" problem. It does not yet account for **over-dosage** (the treatment is correct, but the quantity is too high).
    
    **Future Outlook**: Integrating this approach with external medical Knowledge Graphs (like DrugBank) could further enhance the model's clinical reasoning, allowing it to detect not just fraud, but also potentially harmful drug-drug interactions.

Find Similar Papers

Try Our Examples

  • Search for recent papers using Heterogeneous Graph Neural Networks (HGNNs) or Knowledge Graphs specifically for medical insurance fraud and overtreatment detection.
  • Which study first introduced the use of Relative Risk (RR) for constructing medical association networks, and how does this paper's graph-based refinement compare?
  • Explore the application of metapath2vec or similar HIN embedding techniques in other high-stakes auditing industries, such as financial tax evasion or credit card fraud detection.
Contents
Beyond Co-occurrence: Graph Embeddings for Precision Overtreatment Detection
1. TL;DR
2. The Core Intuition: Why Simple "Co-occurrence" Fails
3. Methodology: The Three-Stage Framework
3.1. 1. Constructing the Association Network
3.2. 2. Selecting the Best Graph Embedding
3.3. 3. Overtreatment Detection via Link Prediction
4. Experimental Results: Precision Matters
5. Critical Analysis & Conclusion