Revealing the Pulse of Society: Decoding Human Behavior via DNS Traffic

Revealing User Behavior by Analyzing DNS Traffic

2020-01-01
Martín Panza, Diego Madariaga, Javier Bustos-Jiménez
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a two-stage machine learning framework to reveal human behavior patterns by analyzing DNS traffic time series from the '.cl' ccTLD. It utilizes Partitioning Around Medoids (PAM) with Shape-Based Distance for clustering and the Apriori algorithm for mining association rules between different domain categories.

TL;DR

Can the "phonebook of the Internet" tell us when a nation sleeps, works, or shops? This paper demonstrates that DNS traffic is a rich mirror of human activity. By combining time-series clustering (PAM + SBD) with transaction mining (Apriori), researchers successfully grouped Chilean domains into semantically coherent "behavioral clusters" and extracted rules that define how different sectors of society influence each other online.

Problem & Motivation: The Social Signal in the Wire

Every time you visit a website or use an app, a DNS query is triggered. Aggregated at the scale of a Country Code Top-Level Domain (ccTLD) like Chile’s .cl, this data forms a massive time series. The challenge is that this raw data is noisy. While simple "day vs. night" cycles are obvious, the nuanced differences—such as how a university's traffic differs from a bank's or a delivery service's—require more than just visual inspection.

The authors' intuition was that the shape of traffic is a proxy for the human lifecycle. If two domains have similar query shapes, they likely serve the same human need at the same time.

Methodology: The Two-Stage Extraction Pipeline

1. Clustering via Shape-Based Distance (SBD)

Instead of standard Euclidean distance, which fails if traffic is shifted slightly in time, the authors used Shape-Based Distance (SBD). SBD focuses on the "morphology" of the traffic waves.

  • Pre-processing: Results were smoothed using Simple Moving Average (SMA) and normalized via Z-Score to isolate shape from absolute volume.
  • Algorithm: Partitioning Around Medoids (PAM) was chosen for its robustness against outliers compared to k-means.

Overall Architecture Fig 1: Visualization of the 6 identified clusters, showing distinct patterns for weekends and nights.

2. Association Rule Mining

To understand how these clusters interact (e.g., "If shopping traffic peaks, does banking traffic also peak?"), the authors applied the Apriori algorithm.

  • Symbolic Conversion: Time series were converted into symbols (a, b, c, d, e) using SAX.
  • Feature Engineering: They added a "delta" feature to capture whether traffic was increasing or decreasing, not just its current level.

Experimental Results & Insights

The method achieved a Davies-Bouldin Index of 0.292 at , indicating high-quality, well-separated clusters.

  • Semantic Coherence: Cluster 6 captured almost all Educational (ED) domains. Cluster 2 became a "Mass Media" hub, grouping Television, Radio, and Newspapers.
  • The Weekend Effect: Cluster 1 (Banking, Government, Postal) showed massive drops on weekends, while Cluster 3 (Tourism/Search) maintained or even peaked during the break.

Experimental Results Table 1: Association rules showing how Cluster 2 (Media) and Cluster 4 (Banking) often move in lockstep.

One of the most striking findings was Rule #3: When Banking and Mass Media clusters both see a sharp increase (magnitude 2), the Government/Postal cluster (Cluster 1) follows with a Lift of 28. This suggests a highly synchronized start to the Chilean business day across multiple human sectors.

Critical Analysis & Conclusion

Takeaway

This work transforms DNS logs from a technical necessity into a sociological sensor. By proving that clustering can group domains by human intent without knowing the domain's content beforehand, it opens doors for automated traffic forecasting and infrastructure optimization.

Limitations

  • Data Granularity: The study uses 10-minute intervals; finer granularity might reveal even more specific "micro-behaviors."
  • Domain Selection: The study relied on the top 82 domains. Pattern mining for mid-tail or niche domains remains unexplored.

Future Outlook

As the Internet becomes more integrated into IoT and smart cities, applying these techniques to encrypted DNS (DoH/DoT) could be the next frontier for privacy-preserving behavioral research.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Shape-Based Distance (SBD) or k-Shape clustering for network traffic anomaly detection.
  • What are the state-of-the-art methods for transforming DNS time-series data into symbolic representations beyond SAX for behavior analysis?
  • Explore research that applies association rule mining to correlate ccTLD DNS traffic with national-level socio-economic events or public holidays.
Contents
Revealing the Pulse of Society: Decoding Human Behavior via DNS Traffic
1. TL;DR
2. Problem & Motivation: The Social Signal in the Wire
3. Methodology: The Two-Stage Extraction Pipeline
3.1. 1. Clustering via Shape-Based Distance (SBD)
3.2. 2. Association Rule Mining
4. Experimental Results & Insights
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook