[Paper Review] Focusing on a Probability Element: Parameter Selection of Message Importance Measure in Big Data
This paper proposes a parameter selection method for Message Importance Measure (MIM) that focuses on a specific probability element in big data distributions by setting the importance coefficient π_j = 1/p_j, enabling enhanced detection of rare or anomalous events. The method improves anomaly detection by emphasizing low-probability events, with theoretical analysis and simulations confirming its effectiveness in minority subset detection under statistical models.
Message importance measure (MIM) is applicable to characterize the importance of information in the scenario of big data, similar to entropy in information theory. In fact, MIM with a variable parameter can make an effect on the characterization of distribution. Furthermore, by choosing an appropriate parameter of MIM, it is possible to emphasize the message importance of a certain probability element in a distribution. Therefore, parametric MIM can play a vital role in anomaly detection of big data by focusing on probability of an anomalous event. In this paper, we propose a parameter selection method of MIM focusing on a probability element and then present its major properties. In addition, we discuss the parameter selection with prior probability, and investigate the availability in a statistical processing model of big data for anomaly detection problem.
Motivation & Objective
- To address the limitation of existing MIM parameter selection methods that focus only on the minimum probability event, ignoring other rare events.
- To develop a practical and adaptive parameter selection strategy for MIM that emphasizes a specific probability element in a distribution.
- To enable more effective minority subset detection in big data by tailoring MIM to focus on anomalous events with small but non-minimal probabilities.
- To validate the proposed method in statistical models relevant to real-world big data anomaly detection applications.
Proposed method
- Defines a parametric MIM with importance coefficient π_j = 1/p_j, where p_j is the target probability element to emphasize.
- Derives the MIM expression as L_j(p, π_j) = log(β p_i * exp( (1/p_j)(1 - p_i) )), which adjusts sensitivity to p_j.
- Analyzes key properties of the MIM under this parameterization, including monotonicity and behavior under uniform vs. non-uniform distributions.
- Applies first-order and second-order Taylor approximations to estimate the expected value and variance of the empirical MIM under sampling.
- Uses Chebyshevβs inequality to analyze convergence properties of the empirical MIM estimator as sample size increases.
- Validates the method using numerical simulations on binomial, Poisson, and geometric distributions.
Experimental results
Research questions
- RQ1How can the MIM parameter be selected to focus on a specific low-probability event rather than just the minimum probability event?
- RQ2What are the fundamental properties of the MIM when the importance coefficient is set as π_j = 1/p_j?
- RQ3How does the MIM perform in detecting minority subsets when applied to different probability distributions?
- RQ4Can the empirical MIM estimator converge reliably under sampling, and how does sample size affect its accuracy and variance?
- RQ5Is the proposed MIM parameterization effective in distinguishing rare events from uniform or majority distributions?
Key findings
- The MIM with π_j = 1/p_j is strictly decreasing with respect to p_j across all tested distributions, confirming its sensitivity to low-probability events.
- For non-uniform distributions, the MIM value focusing on a specific p_j exceeds that of the uniform distribution, validating its ability to detect deviations from uniformity.
- The expected value of the empirical MIM estimator E(πΏΜ_i) converges toward the true MIM value as sample size N_i increases, especially when p is small.
- The variance D(πΏΜ_i) of the empirical MIM estimator decreases with increasing N_i, and is monotonically decreasing with ΞΌ = p when 0 < p < 1/2.
- The MIM focusing on a probability element shows strong convergence in empirical settings, making it suitable for statistical anomaly detection in big data.
- Numerical results confirm that the MIM with π_j = 1/p_j effectively highlights rare events, outperforming uniform distribution baselines in detection capability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card Β· Free plan available
This review was created by AI and reviewed by human editors.