Skip to main content
QUICK REVIEW

[Paper Review] DogFLW: Dog Facial Landmarks in the Wild Dataset

George Martvel, Greta Abele|arXiv (Cornell University)|May 19, 2024
Wildlife Ecology and ConservationEnvironmental Science3 citations
TL;DR

This paper introduces DogFLW, a dataset of 3,274 wild-dog images annotated with 46 anatomy-based facial landmarks, enabling automated analysis of dog facial expressions for affective computing. The study benchmarks landmark detection using an ensemble learning approach, achieving a normalized mean error (NME) of 5.31 on the full test set, with higher errors in ear regions due to breed-specific morphological variation.

ABSTRACT

Affective computing for animals is a rapidly expanding research area that is going deeper than automated movement tracking to address animal internal states, like pain and emotions. Facial expressions can serve to communicate information about these states in mammals. However, unlike human-related studies, there is a significant shortage of datasets that would enable the automated analysis of animal facial expressions. Inspired by the recently introduced Cat Facial Landmarks in the Wild dataset, presenting cat faces annotated with 48 facial anatomy-based landmarks, in this paper, we develop an analogous dataset containing 3,274 annotated images of dogs. Our dataset is based on a scheme of 46 facial anatomy-based landmarks. The DogFLW dataset is available from the corresponding author upon a reasonable request.

Motivation & Objective

  • To address the critical lack of large-scale, diverse datasets for automated dog facial expression analysis in uncontrolled environments.
  • To develop a standardized, anatomy-based landmark scheme (46 points) aligned with DogFACS for consistent facial feature representation.
  • To enable the training of robust, generalizable models for facial landmark detection across diverse dog breeds and morphologies.
  • To support research in canine affective computing, including emotion recognition, health monitoring, and human-dog interaction studies.
  • To provide a benchmark dataset that facilitates progress in AI-driven analysis of dog facial expressions for welfare and cognitive science.

Proposed method

  • Anatomically guided landmark annotation using 46 points based on dog facial musculature and DogFACS, ensuring consistency with affective computing standards.
  • Data collection from diverse sources to include real-world conditions (e.g., varying lighting, poses, backgrounds), enhancing dataset realism.
  • Application of an ensemble learning model combining multiple deep learning architectures to improve landmark detection robustness.
  • Training and evaluation on a balanced split of 3,274 images, with 2,500 for training and 774 for testing, using normalized mean error (NME) as the primary metric.
  • Subgroup analysis by ear type (erect, floppy, half-floppy) and breed to evaluate detection performance variation across morphological subtypes.
  • Use of pre-cropped images for inference to isolate landmark detection performance from face detection errors.
Figure 1 : Annotated Dog’s Face. Image of a dog with a face bounding box and 46 facial landmarks.
Figure 1 : Annotated Dog’s Face. Image of a dog with a face bounding box and 46 facial landmarks.

Experimental results

Research questions

  • RQ1How does the performance of facial landmark detection vary across different facial regions in dogs, particularly in the ears?
  • RQ2To what extent do breed-specific morphological traits (e.g., fur length, ear type) affect landmark detection accuracy?
  • RQ3Can a single, standardized landmark scheme based on canine facial anatomy support reliable and generalizable facial expression analysis across diverse dog breeds?
  • RQ4How does dataset composition (e.g., breed distribution, ear type prevalence) influence model generalization and error rates?
  • RQ5What are the key challenges in detecting facial landmarks in dogs with extreme morphological variation, and how can they be mitigated?

Key findings

  • The overall normalized mean error (NME) for landmark detection on the full test set was 5.31, indicating strong baseline performance for a real-world dataset.
  • Ear regions exhibited the highest error (NME = 13.19), significantly higher than eyes (2.24) and nose (4.22), due to morphological diversity and occlusion.
  • Dogs with floppy ears showed the highest detection error (NME = 15.28), followed by half-floppy (12.77) and erect ears (9.79), confirming breed morphology as a major challenge.
  • Breeds with long, dense fur covering facial features (e.g., Irish Water Spaniel, Basset Hound) had notably higher detection errors due to occlusion of landmarks.
  • Breeds with short, smooth fur (e.g., Chihuahua, Miniature Pinscher) achieved the lowest detection errors, indicating clearer feature visibility improves model performance.
  • The model’s performance was significantly affected by breed-specific traits such as ear position variability (e.g., Collie, Ibizan Hound) and obscured features (e.g., Chow), highlighting the need for more diverse training data.
Figure 2 : Examples of annotated images from the DogFLW.
Figure 2 : Examples of annotated images from the DogFLW.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.