Skip to main content
QUICK REVIEW

[Paper Review] Automatic Classification of X-rated Videos using Obscene Sound Analysis based on a Repeated Curve-like Spectrum Feature

Jae-Deok Lim, Byeong-Cheol Choi|arXiv (Cornell University)|Dec 9, 2011
Video Analysis and Summarization16 references3 citations
TL;DR

This paper proposes a repeated curve-like spectrum feature to automatically classify X-rated videos by detecting obscene sounds—such as sexual moans and screams—using audio analysis. The method achieves 96.6% F1-score on clean audio and 92.6% on noisy 5dB SNR data, demonstrating high accuracy in identifying explicit content using only audio cues.

ABSTRACT

This paper addresses the automatic classification of X-rated videos by analyzing its obscene sounds. In this paper, obscene sounds refer to audio signals generated from sexual moans and screams during sexual scenes. By analyzing various sound samples, we determined the distinguishable characteristics of obscene sounds and propose a repeated curve-like spectrum feature that represents the characteristics of such sounds. We constructed 6,269 audio clips to evaluate the proposed feature, and separately constructed 1,200 X-rated and general videos for classification. The proposed feature has an F1-score, precision, and recall rate of 96.6%, 98.2%, and 95.2%, respectively, for the original dataset, and 92.6%, 97.6%, and 88.0% for a noisy dataset of 5dB SNR. And, in classifying videos, the feature has more than a 90% F1-score, 97% precision, and an 84% recall rate. From the measured performance, X-rated videos can be classified with only the audio features and the repeated curve-like spectrum feature is suitable to detect obscene sounds.

Motivation & Objective

  • To address the challenge of automating the detection of X-rated content in videos using audio signals alone.
  • To identify distinctive acoustic patterns in obscene sounds such as sexual moans and screams.
  • To develop a robust spectral feature that captures the repetitive, curve-like nature of these sounds for classification.
  • To evaluate the feature’s performance on both clean and noisy audio datasets to ensure real-world applicability.

Proposed method

  • The authors analyze 6,269 audio clips extracted from X-rated and general videos to identify recurring spectral patterns in obscene sounds.
  • They propose a novel repeated curve-like spectrum feature that models the periodic, harmonic-like structure of sexual moans and screams.
  • The feature is derived from spectral envelope analysis, emphasizing repeated, smooth, curve-like patterns across frequency bands.
  • Classification is performed using machine learning models trained on the proposed spectral feature, with evaluation on both original and noisy audio datasets.
  • The method focuses exclusively on audio, avoiding reliance on visual or textual content.
  • Performance is evaluated using standard metrics: F1-score, precision, and recall on both clean and 5dB SNR degraded audio.

Experimental results

Research questions

  • RQ1Can obscene sounds in X-rated videos be reliably distinguished from general audio using spectral features?
  • RQ2Does a repeated curve-like spectrum feature effectively capture the acoustic characteristics of sexual moans and screams?
  • RQ3How well does the proposed feature generalize to noisy audio conditions typical in real-world video content?
  • RQ4Can automatic classification of X-rated videos be achieved with high accuracy using only audio features?
  • RQ5What is the performance of the method on both clean and degraded audio datasets?

Key findings

  • The proposed repeated curve-like spectrum feature achieves an F1-score of 96.6%, precision of 98.2%, and recall of 95.2% on the original (clean) audio dataset.
  • On a noisy dataset with 5dB signal-to-noise ratio, the feature maintains strong performance with an F1-score of 92.6%, precision of 97.6%, and recall of 88.0%.
  • In full video classification, the method achieves an F1-score above 90%, precision of 97%, and recall of 84%.
  • The results confirm that audio-only classification of X-rated content is feasible using the proposed spectral feature.
  • The feature is robust to noise, maintaining high performance even under degraded audio conditions.
  • The study demonstrates that obscene sounds exhibit consistent spectral patterns that can be effectively modeled and detected using the proposed method.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.