Cube Sampled K-Prototype Clustering for Featured Data

Jain, Seemandhar; Shastri, Aditya; Ahuja, Kapil; Busnel, Yann; Singh, Navneet Pratap

doi:10.48550/arxiv.2108.10262

Cited by 1 publication

(1 citation statement)

References 8 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Prior to using unequal probability sampling, an inclusion probability must be determined. There have been efforts to assess the probability of inclusion for large datasets, such as Jain et al [13] and Nigam et al [16]. This is not a feasible strategy for dealing with multi-dimensional data.…”

Section: Related Workmentioning

confidence: 99%

An Efficient Anomaly Detection Approach using Cube Sampling with Streaming Data

Jain,

Srivastava

2021

Preprint

Self Cite

View full text Add to dashboard Cite

Anomaly detection is critical in various fields, including intrusion detection, health monitoring, fault diagnosis, and sensor network event detection. The isolation forest (or iForest) approach is a well-known technique for detecting anomalies. It is, however, ineffective when dealing with dynamic streaming data, which is becoming increasingly prevalent in a wide variety of application areas these days. In this work, we extend our previous work by proposed an efficient iForest based approach for anomaly detection using cube sampling that is effective on streaming data. Cube sampling is used in the initial stage to choose nearly balanced samples, significantly reducing storage requirements while preserving efficiency. Following that, the streaming nature of data is addressed by a sliding window technique that generates consecutive chunks of data for systematic processing. The novelty of this paper is in applying Cube sampling in iForest and calculating inclusion probability. The proposed approach is equally successful at detecting anomalies as existing stateof-the-art approaches, requiring significantly less storage and time complexity. We undertake empirical evaluations of the proposed approach using standard datasets and demonstrate that it outperforms traditional approaches in terms of Area Under the ROC Curve (AUC-ROC) and can handle high-dimensional streaming data.

show abstract

Section: Related Workmentioning

confidence: 99%