[Paper Review] A Drowsiness Detection Scheme Based on Fusion of Voice and Vision Cues.
This paper proposes a non-contact drowsiness detection system fusing voice and vision cues to assess driver drowsiness levels in real time. By combining audio features (e.g., speech patterns) and visual features (e.g., eye closure, head posture), the system achieves accurate drowsiness estimation without invasive methods, with results cross-validated against brain signal data for reliability.
Drowsiness level detection of an individual is very important in many safety critical applications such as driving. There are several invasive and contact based methods such as use of blood biochemical, brain signals etc. which can estimate the level of drowsiness very accurately. However, these methods are very difficult to implement in practical scenarios, as they cause discomfort to the user. This paper presents a combined voice and vision based drowsiness detection system well suited to detect the drowsiness level of an automotive driver. The vision and voice based detection, being non-contact methods, has the advantage of their feasibility of implementation. The authenticity of these methods have been cross-validated using brain signals.
Motivation & Objective
- To develop a non-invasive, real-time drowsiness detection system suitable for automotive environments.
- To improve detection accuracy by fusing complementary voice and vision-based features.
- To validate the system's reliability using brain signals as a gold standard.
- To enable practical deployment through contact-free sensing methods.
- To reduce user discomfort compared to invasive physiological monitoring techniques.
Proposed method
- The system captures audio and video streams from the driver in real time.
- Voice features are extracted using speech processing techniques, focusing on changes in speech rhythm and pitch.
- Vision features are derived from facial landmarks, particularly eye closure and head pose, using computer vision algorithms.
- A fusion model combines voice and vision features to estimate drowsiness levels.
- The system's performance is validated by comparing its output with drowsiness levels inferred from brain signals.
- Feature fusion is performed using a weighted combination strategy to optimize detection accuracy.
Experimental results
Research questions
- RQ1Can voice and vision cues be effectively combined to detect drowsiness without physical contact?
- RQ2How does the fusion of audio and visual features improve drowsiness detection accuracy compared to single-modality approaches?
- RQ3To what extent does the system's output correlate with actual drowsiness levels as measured by brain signals?
- RQ4Can the system be reliably deployed in real-world driving conditions?
Key findings
- The proposed fusion system achieves higher drowsiness detection accuracy than single-modality approaches.
- The system demonstrates strong correlation with drowsiness levels derived from brain signals, validating its reliability.
- Voice and vision cues provide complementary information, enhancing detection robustness.
- The non-contact nature of the system ensures user comfort and practical deployability in vehicles.
- The fusion model effectively reduces false positives and improves detection consistency.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.