Beyond K-Means: Mastering Process Complexity with Context-Aware Trace Clustering

A New Trace Clustering Algorithm Based on Context in Process Mining

2018-01-01
Hong-Nhung Bui, Tri-Thanh Nguyen, Thi-Cham Nguyen, Quang-Thuy Ha
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces ContextTracClus, a novel trace clustering algorithm for process mining that utilizes "trace contexts" (common prefixes) to group event logs. Unlike traditional methods like K-means, it automatically determines the number of clusters and avoids iterative convergence, significantly reducing computational complexity while improving process discovery precision.

TL;DR

Process mining often suffers from "Spaghetti Models"—complex, unreadable diagrams generated from messy event logs. While clustering is the standard solution, traditional methods like K-means are often a "square peg in a round hole." This paper introduces ContextTracClus, a tree-based algorithm that uses "Trace Contexts" (common prefixes) to automatically partition logs into clean, logical sub-processes with higher precision and lower overhead than ever before.

The Problem: The "Spaghetti" Nightmare

Process discovery aims to turn event logs into Petri nets. When logs are small, the result is elegant. However, real-world logs are chaotic. Standard K-means clustering treats these logs as flat vectors, losing the sequential "flow" that defines a business process. Furthermore, someone has to guess the value of (number of clusters), and the iterative nature of these algorithms makes them slow for massive datasets.

The Spaghetti Model Problem Fig 1. Difference between a structured model (left) and the dreaded 'Spaghetti' model (right).

The Insight: Context is Key

The researchers argue that every business process follows discrete "procedures" (e.g., personal loans vs. corporate loans). These procedures naturally reveal themselves through Common Prefixes. If a group of traces starts with the same sequence of activities, they likely belong to the same logical variant.

The Context-Tree

To find these prefixes efficiently, the authors designed a Context-Tree. This is a compact data structure where:

  • Each node represents an activity.
  • Each branch represents a sequence.
  • Nodes store a count of how many traces passed through that specific path.

Concept of Context Tree Fig 2. The Context-Tree: Mapping traces into shared branches to identify logical contexts.

Methodology: ContextTracClus

The algorithm operates in two elegant phases:

  1. Cluster Discovery: It builds the Context-tree in one pass. For any trace, its "Context" is the longest sequence of nodes from the root where the count > 1. This effectively identifies shared procedural pathways.
  2. Adjustment: Small "micro-clusters" (under a specific threshold) are merged into their nearest contextually similar neighbor.

Key Advantage: Zero iterations. No waiting for centroids to converge. It works directly on the traces without converting them into high-dimensional binary vectors.

Experimental Performance

The authors tested the algorithm against K-means across three different event logs (Lfull, prAm6, prHm6).

Precision is the Winner

While Fitness (the ability to replay logs) remained high across all methods, Precision saw a massive boost. Because ContextTracClus groups traces that share specific structural paths, the resulting Petri nets are much tighter and don't allow for behaviors that never actually occurred in the data.

Experimental Results Table 1. ContextTracClus consistently outperforms K-means, especially in Precision.

Critical Insight: Why it Works

Traditional clustering (like K-means with k-grams) is probabilistic. It looks at the frequency of activity pairs. ContextTracClus is structural. It respects the "Start-to-Finish" integrity of a business process. By using a tree structure, it preserves the causality of the first few critical steps of a process, which are often the strongest indicators of which "variant" a case belongs to.

Conclusion & Future Outlook

ContextTracClus proves that specialized data structures (like the Focus-tree) can outperform general-purpose ML algorithms by baking domain logic into the architecture. For organizations dealing with massive event logs, this offers a path toward Automated Process Discovery that doesn't require a data scientist to tune hyperparameters.

Future work could involve expanding this "Context" beyond just prefixes to include Resource-based context (who is doing the work) or Time-based context (holiday seasons), creating a multi-dimensional view of process variants.

Find Similar Papers

Try Our Examples

  • Find recent research papers that use Prefix Trees or FP-growth structures for trace clustering in process mining beyond 2020.
  • Which paper first introduced the concept of "Spaghetti Models" in process mining, and how have subsequent clustering methods evolved to address this?
  • Investigate if the ContextTracClus approach has been adapted for real-time stream process mining or concept drift detection.
Contents
Beyond K-Means: Mastering Process Complexity with Context-Aware Trace Clustering
1. TL;DR
2. The Problem: The "Spaghetti" Nightmare
3. The Insight: Context is Key
3.1. The Context-Tree
4. Methodology: ContextTracClus
5. Experimental Performance
5.1. Precision is the Winner
6. Critical Insight: Why it Works
7. Conclusion & Future Outlook