SSL-IDS — Anomaly Detection & Representation Learning
PyTorchPythonMachine LearningNetwork SecurityResearch
Overview
This repository contains the codebase and experimental pipeline for the research paper:
"Contrastive Representation Learning for Network Flow Anomaly Detection Under Distribution Shift" (Presented at CCIDSA 2026).
The research evaluates contrastive self-supervised learning (SSL) for flow-level network intrusion detection, characterizing the learned latent manifold under distribution shift and temporal drift rather than simple classification leaderboards.
Key Research Findings
- Cross-Dataset Transferability: Under a strict Leave-One-Attack-Type-Out (LOATO) protocol evaluating transfer between the
CIC-IDS2017andCSE-CIC-IDS2018datasets, contrastive embeddings yield the lowest performance degradation (a ΔAUROC of −0.430) compared to reconstruction-based (−0.527) and partition-based (−0.444) baselines. - Augmentation Selection: Tabular network flow data is semantically constrained. Implementing simple Gaussian Jitter is highly effective (+5.5% DoS AUROC), whereas aggressive combined augmentations destroy flow semantics.
- Calibration Collapse Under Temporal Drift: While aggregate AUROC remains stable, thresholded true positive rates (TPR) collapse from 0.120 to 0.000 at a 1% False Positive Rate (FPR) under temporal shift. This proves that representation robustness and score calibration are distinct operational challenges.
Tech Stack
PyTorch Scikit-Learn InfoNCE Loss Pandas CUDA Data Science