GAN-Impute: Leveraging Generative Adversarial Networks for P2P Lending Risk Assessment

AI-Based Online P2P Lending Risk Assessment On Social Network Data With Missing Value

2019-12-01
Lok Ting Lam, Shun-Wen Hsiao
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a Generative Adversarial Network (GAN) based approach for missing value imputation in online P2P lending risk assessment. By training a Generator to produce plausible data points and a Discriminator to distinguish them from complete records, the model effectively recreates social network data patterns to enhance credit prediction accuracy.

Executive Summary

TL;DR: This research tackles the pervasive issue of missing data in online P2P lending by introducing a GAN-based imputation framework. By training a generator to "fill in the blanks" and a discriminator to maintain data integrity, the method provides a robust pipeline for preparing messy social network data for high-precision risk assessment. It moves beyond simple statistical averages to capture the "hidden physics" of financial datasets.

Positioning: This work represents a transition from traditional statistical cleaning (Mean/Median) to AI-driven data synthesis in the FinTech domain, specifically targeting the high-sparsity nature of P2P lending features.

Problem & Motivation

In the world of online P2P lending, data is rarely pristine. Features such as "listing titles" or social network interactions often contain over 30% missing values.

The authors identify a critical flaw in current industry standards:

  • Mean/Median Imputation: Reduces variability and underestimates variance, Leading to overconfident but inaccurate risk models.
  • Deletion: Wastes valuable samples, which is unacceptable in specialized financial subsets.
  • Linear Regression Imputation: Assumes a perfect linear correlation (1.0) between variables, which fails to capture the non-linear complexities of borrower behavior.

The motivation is clear: if we can simulate the distribution of the missing data using its relationship with existing features, we can build a much more resilient credit scoring model.

Methodology: The Adversarial Engine

The core innovation lies in applying the Generative Adversarial Network (GAN) architecture to tabular data imputation.

The Architecture

The workflow is divided into three distinct phases:

  1. Preparation: Complete data () is passed through a "Random Dropper" to simulate real-world sparsity ().
  2. Adversarial Training:
    • Generator (G): Takes the incomplete record and a noise vector to produce a reconstructed record ().
    • Discriminator (D): Evaluates whether the imputed values "fit" the patterns seen in complete records.
  3. Imputation & Classification: Once G is trained, it generates multiple candidates; the system uses a latent-space k-NN approach to select the most realistic values before feeding the data into a risk classifier.

Model Architecture Fig 1. Detailed GAN architecture for missing value imputation.

The Loss Function

Unlike standard GANs used for images, this model incorporates an MSE-weighted loss. This ensures that the generated data doesn't just "look" real but is also numerically close to the expected values in a financial context.

Experiments & Results

The authors tested their model using the Lending Club dataset from Kaggle, simulating a 20% data loss scenario.

SOTA Comparison & Stability

The experimental results show that the model is highly efficient:

  • Convergence: Both training and testing MSE (Mean Square Error) show a sharp decline, stabilizing after approximately 1,000 epochs.
  • Reliability: The generator successfully learns the internal correlations of the 18 columns (loan amount, interest rate, term, etc.), ensuring the "fake" data is indistinguishable from real records from the perspective of the discriminator.

MSE Train Loss Fig 2. Convergence of the MSE training loss over 5,000 epochs.

Critical Analysis & Conclusion

Takeaway

The study demonstrates that GANs are not just for generating "Deepfakes" or images; they are powerful tools for Data Augmentation and Cleaning. In financial risk assessment, where a single missing feature can lead to a wrong "Default" prediction, the ability of GANs to preserve the joint distribution of features is a game-changer.

Limitations

While effective, the paper relies on a relatively small sample size (1,000 records) and a simplistic "Random Missing" (MCAR) assumption. In real-world P2P lending, data is often "Missing Not At Random" (MNAR)—for example, borrowers with poor credit may intentionally hide certain information.

Future Outlook

The next step for this technology is its application in Unbalanced Data. In credit risk, defaults are rare events. Combining GAN imputation with SMOTE-like oversampling could yield a new generation of highly robust risk assessment engines.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Generative Adversarial Imputation Nets (GAIN) specifically for financial risk assessment or credit scoring.
  • What are the original theoretical foundations of using GANs for tabular data imputation, and how does this study's HINT-based discriminator modification compare to the GAIN framework?
  • Explore research that applies GAN-based missing value imputation to multi-modal datasets combining tabular P2P data with textual "listing titles".
Contents
GAN-Impute: Leveraging Generative Adversarial Networks for P2P Lending Risk Assessment
1. Executive Summary
2. Problem & Motivation
3. Methodology: The Adversarial Engine
3.1. The Architecture
3.2. The Loss Function
4. Experiments & Results
4.1. SOTA Comparison & Stability
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook