SysDroid: Mastering Dynamic Android Malware Detection via SAILS Feature Selection
SysDroid: a dynamic ML-based android malware analyzer using system call traces
SysDroid is a dynamic Android malware analyzer that leverages system call traces and N-gram sequences (unigrams, bigrams, trigrams) to detect malicious applications. The core contribution is a novel feature selection mechanism called SAILS (Selection of relevant Attributes for Improving Locally extracted features using classical feature Selectors), which consistently achieves detection accuracies between 95% and 99.4% across various ML and Deep Learning models.
TL;DR
SysDroid presents a robust dynamic analysis framework that utilizes system call traces transformed into N-gram sequences. By introducing a novel feature selection algorithm named SAILS, the researchers achieved near-perfect detection rates (up to 99.4% accuracy). This work highlights the power of XGBoost and DNNs in behavioral analysis while exposing a critical vulnerability: the susceptibility of ML classifiers to adversarial evasion via benign feature injection.
Background: Why Dynamic Analysis?
The Android ecosystem remains a primary target for banking Trojans and ransomware. Static analysis—examining code without execution—is increasingly inadequate because modern malware utilizes encryption and dynamic execution to hide its payload. Dynamic analysis, which monitors the application during runtime using tools like strace, provides a "ground truth" of what the app actually does at the kernel level.
The Core Innovation: SAILS Methodology
The primary challenge in machine learning for security is Feature Selection. Including too many system calls introduces noise; including too few misses the "signal."
The authors proposed SAILS (Selection of relevant Attributes for Improving Locally extracted features using classical feature Selectors).
- Local Scoring: It first uses standard metrics (Mutual Information, GSS, DFS) to score how well a system call identifies "Malware" and "Benign" apps separately.
- Global Ranking: Utilizing a Max-Heap structure, it interleaves the top-ranked calls from both lists. This ensures the model learns not just what malware does, but also what makes an app legitimate, creating a more discriminative boundary.
Figure 1: The SysDroid Workflow - From Data Collection to Adversarial Evaluation
Experimental Insights: XGBoost vs. DNN
The study conducted exhaustive tests on the Drebin dataset (2474 malware, 2475 benign).
- N-Grams matter: Bigrams and Trigrams consistently outperformed Unigrams, as they capture the sequence of operations, which is far more indicative of malicious intent than standalone calls.
- Champion Models: XGBoost emerged as a powerhouse for Bigrams (99.4% accuracy), while Logistic Regression (LR) surprisingly excelled in high-dimensional Trigram spaces due to its linear behavior in complex manifolds.
- DL Regularization: In the Deep Learning tracks, the authors found that Dropout (rates 0.3-0.6) was essential to prevent the model from simply memorizing specific malware families.
Table 1: Accuracy comparison across different classifiers and N-gram lengths.
The Achilles' Heel: Adversarial Evasion
Perhaps the most significant contribution of this paper is the Adversarial Machine Learning (AML) evaluation. The authors simulated an "Evasion Attack" where a malware sample is injected with a small percentage (1%-3%) of system calls typically found only in benign apps (e.g., UI swipes or battery status checks).
The Result? A catastrophic decline in Recall. Some classifiers saw their true positive rate drop from 99% to below 25%.
Figure 2: The sharp decline in detection capability after benign call poisoning.
Takeaway and Future Directions
While SysDroid and the SAILS mechanism provide a powerful toolkit for identifying today's malware, the research proves that high accuracy is a "glass cannon." A hacker doesn't need to rewrite their virus; they just need to make it look a bit more like a calculator.
Future research must focus on Adversarial Training—including these "poisoned" samples in the training set—to build models that can see through the benign camouflage.
