This study evaluates the ability of large vision-language models, such as GPT-4o and Gemini, to accurately identify dog emotions from images. Researchers discovered that these AI systems frequently rely on superficial contextual cues, such as the background environment, rather than the animal’s actual biological signals. When tested on scientifically controlled datasets featuring cropped facial expressions and minimal scenery, the models’ performance plummeted to near-chance levels. These results suggest that current AI carries significant anthropocentric biases and often misinterprets canine internal states based on human-centric assumptions. Consequently, the authors argue for a more interdisciplinary approach that integrates validated behavioral science to improve animal-centered artificial intelligence. Through background manipulation and prompt testing, the paper highlights the technical and ethical risks of using general-purpose models for veterinary or animal welfare assessments.
Supervised Neural Style Transfer as an Augmentation Technique for Facial Landmark Detection
A major challenge in training accurate facial landmark detection models