Skip to main content
QUICK REVIEW

[论文解读] Automatic Classification of X-rated Videos using Obscene Sound Analysis based on a Repeated Curve-like Spectrum Feature

Jae-Deok Lim, Byeong-Cheol Choi|arXiv (Cornell University)|Dec 9, 2011
Video Analysis and Summarization参考文献 16被引用 3
一句话总结

本文提出了一种重复曲线状谱特征,通过音频分析自动分类X级视频,检测诸如性呻吟和尖叫等低俗声音。该方法在干净音频上的F1得分为96.6%,在信噪比5dB的嘈杂数据上为92.6%,展示了仅使用音频线索识别明确内容的高精度。

ABSTRACT

This paper addresses the automatic classification of X-rated videos by analyzing its obscene sounds. In this paper, obscene sounds refer to audio signals generated from sexual moans and screams during sexual scenes. By analyzing various sound samples, we determined the distinguishable characteristics of obscene sounds and propose a repeated curve-like spectrum feature that represents the characteristics of such sounds. We constructed 6,269 audio clips to evaluate the proposed feature, and separately constructed 1,200 X-rated and general videos for classification. The proposed feature has an F1-score, precision, and recall rate of 96.6%, 98.2%, and 95.2%, respectively, for the original dataset, and 92.6%, 97.6%, and 88.0% for a noisy dataset of 5dB SNR. And, in classifying videos, the feature has more than a 90% F1-score, 97% precision, and an 84% recall rate. From the measured performance, X-rated videos can be classified with only the audio features and the repeated curve-like spectrum feature is suitable to detect obscene sounds.

研究动机与目标

  • 解决仅使用音频信号自动检测视频中X级内容的挑战。
  • 识别性呻吟和尖叫等低俗声音中的独特声学模式。
  • 开发一种能够捕捉这些声音重复性、曲线状特性的鲁棒谱特征以用于分类。
  • 在干净和嘈杂音频数据集上评估该特征的性能,以确保其在真实场景中的适用性。

提出的方法

  • 作者分析了从X级视频和普通视频中提取的6,269段音频剪辑,以识别低俗声音中的重复谱模式。
  • 提出一种新颖的重复曲线状谱特征,用于建模性呻吟和尖叫的周期性、类似谐波的结构。
  • 该特征源自谱包络分析,强调在频带中重复、平滑的曲线状模式。
  • 使用在所提谱特征上训练的机器学习模型进行分类,并在原始和嘈杂音频数据集上进行评估。
  • 该方法仅依赖音频,不依赖视觉或文本内容。
  • 使用标准指标(F1得分、精确率和召回率)在干净和5dB信噪比降质音频上评估性能。

实验结果

研究问题

  • RQ1能否通过谱特征可靠地区分X级视频中的低俗声音与普通音频?
  • RQ2重复曲线状谱特征是否能有效捕捉性呻吟和尖叫的声学特征?
  • RQ3所提出的特征在典型真实视频内容中常见的嘈杂音频条件下泛化能力如何?
  • RQ4能否仅使用音频特征实现高精度的X级视频自动分类?
  • RQ5该方法在干净和降质音频数据集上的性能如何?

主要发现

  • 所提出的重复曲线状谱特征在原始(干净)音频数据集上达到96.6%的F1得分、98.2%的精确率和95.2%的召回率。
  • 在信噪比为5dB的嘈杂数据集上,该特征仍保持优异性能,F1得分为92.6%,精确率为97.6%,召回率为88.0%。
  • 在完整视频分类中,该方法的F1得分超过90%,精确率为97%,召回率为84%。
  • 结果证实,仅使用所提出的谱特征即可实现音频驱动的X级内容自动分类。
  • 该特征对噪声具有鲁棒性,在音频质量下降的条件下仍能保持高性能。
  • 本研究证明,低俗声音表现出一致的谱模式,可被所提方法有效建模与检测。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。