Projecte llegit
Títol: Can General-Purpose Generative AI Models Match the Performance of Specialized Facial Emotion Recognition Systems?
Estudiants que han llegit aquest projecte:
LILLO CASTILLO, ALEX (data lectura: 23-07-2026)- Cerca aquest projecte a Bibliotècnica
LILLO CASTILLO, ALEX (data lectura: 23-07-2026)Director/a: TARRÉS RUIZ, FRANCESC
Departament: TSC
Títol: Can General-Purpose Generative AI Models Match the Performance of Specialized Facial Emotion Recognition Systems?
Data inici oferta: 16-07-2026 Data finalització oferta: 16-07-2026
Estudis d'assignació del projecte:
GR ENG TELEMÀTICA
| Tipus: Individual | |
| Lloc de realització: ERASMUS | |
| Paraules clau: | |
| Facial Emotion Recognition (FER), Artificial Intelligence, Generative Artificial Intelligence, Deep Learning, Emotion Classification | |
| Descripció del contingut i pla d'activitats: | |
| Overview (resum en anglès): | |
| Facial Emotion Recognition (FER) has become an important research field within Artificial Intelligence due to its potential applications in healthcare, human-computer interaction, marketing, education and automotive safety. Traditionally, FER has relied on specialized deep learning models trained exclusively for emotion classification. However, the rapid development of multimodal generative Artificial Intelligence systems raises the question of whether these general-purpose models can achieve comparable performance despite not being specifically designed for facial emotion recognition.
The objective of this bachelor's thesis is to compare the performance of a specialized FER system with modern multimodal generative AI models. Specifically, DeepFace was selected as the representative specialized model and compared against ChatGPT and Gemini under identical experimental conditions using the RAF-DB dataset, a widely adopted benchmark containing real-world facial images annotated with seven basic emotion categories. The evaluation methodology consisted of three experiments of increasing complexity, ranging from an initial exploratory analysis to a balanced evaluation including all seven emotions. Model predictions were processed through a Python evaluation pipeline to compute overall accuracy and generate confusion matrices, enabling a detailed comparison of classification performance and error patterns. The results show that specialized FER models remain the most reliable solution for facial emotion recognition tasks. DeepFace consistently achieved the highest accuracy across the experiments, while ChatGPT obtained competitive results despite not being specifically trained for this task. Gemini presented the lowest performance together with higher variability between experiments. Error analysis revealed that all models experienced greater difficulty distinguishing subtle emotions such as fear, surprise and neutral expressions, whereas more distinctive expressions, particularly happiness, were classified more accurately. Overall, this work demonstrates that modern multimodal generative AI models already possess remarkable facial emotion recognition capabilities and, in some controlled evaluation scenarios, can approach the performance of specialized FER systems. Nevertheless, dedicated FER models continue to provide the highest reliability and consistency. Finally, the study discusses methodological limitations, including the possibility of data leakage in foundation models, and outlines future research directions involving larger and completely unseen datasets to further validate these findings. |
|