Abstracts Language, Literature, and Linguistic

Add abstract

Want to add your dissertation abstract to this database? It only takes a minute!

Search abstract

Search for abstracts by subject, author or institution

Share this abstract

MulTweEmo: A New Resource and Experiments on CLIP-based Multimodal Emotion Recognition

by ROBERTO CANNARELLA

Institution: Pisa University
Department: INFORMATICA UMANISTICA
Degree: FILOLOGIA, LETTERATURA E LINGUISTICA
Year: 2022
Keywords: multimodal semantics; grounded cognition; emotion recognition; multimodal sentiment analysis; CLIP language model; language resources
Posted: 3/25/2025
Record ID: 2286424
Full text PDF: http://etd.adm.unipi.it/theses/available/etd-08292022-105130/


Abstract

This work focuses on the image-text emotion recognition (ITER) task, which consists in training NLP models that, thanks to the combination of visual and textual information, can predict the emotion associated with a multimodal document. Since research on ITER is still scarce, the first aim of this work is to better frame it within a solid theoretical framework. To do that, Chapter 1 presents contributions from cognitive linguistics and communication studies arguing that our language use is inherently multimodal and grounded on perceptual experience. The chapter also revises literature on how extra-linguistic information can be integrated into typical word embeddings, which yields the so-called visual-semantic embeddings. As a task, ITER is deeply connected with a few others: textual emotion recognition (ER), object recognition, and image emotion recognition (IER). For each of them, previous research, methods, and theoretical considerations are described in Chapter 2. Chapter 3, then, focuses on image-polarity classification, which is a task deeply connected with ITER, and finally reports the little existing research on ITER. The rest of the work reports the creation of a new multimodal resource, MulTweEmo, and discusses the related experiments. Chapter 4 describes the creation process, including how labels have been defined, both automatically and, for a portion of the data, through crowdsourcing. Chapter 5 describes several experiments involving the use of the dataset: since the resource contains texts associated with pictures, it has been used to feed classifiers with both unimodal (textual or visual) and multimodal embeddings. All embeddings have been created by using the CLIP model. Experiments show that multimodal classifiers outperform the others across all experimental setups, suggesting that multimodality can be advantageous in tasks involving the analysis of the affective value of documents.

Add abstract

Want to add your dissertation abstract to this database? It only takes a minute!

Search abstract

Search for abstracts by subject, author or institution

Share this abstract

Relevant publications

Book cover thumbnail image
Harriet Beecher Stowe's "Uncle Tom's Cabin" in Ara... Challenges of Cross-Cultural Translation
by AL-Sarrani, Abeer Abdulaziz
   
Book cover thumbnail image
An Analysis of the Knowledge and Use of English Co...
by Kurosaki, Shino
   
Book cover thumbnail image
Literature and Education Proposal of an English Literature Program for Prim...
by Puebla, Esther de la Peña
   
Book cover thumbnail image
Putting Assessment for Learning (AfL) into Practic...
by White, Edmund
   
Book cover thumbnail image
A Case Study on the Impact of Weblogs on the Writi...
by Higginson, Simon
   
Book cover thumbnail image
Thucydides and US Foreign Policy Debates after the...
by Bloxham, John A.
   
Book cover thumbnail image
Quantificational Modification The Semantics of Totality and Proportionality
by Tsouhlaris, Zaina Hafiz
   
Book cover thumbnail image
Language Choice in Interracial Marriages The Case of Filipino-Malaysian Couples
by Dumanig, Francisco Perlas