Add abstract
Want to add your dissertation abstract to this database? It only takes a minute!
Search abstract
Search for abstracts by subject, author or institution
Want to add your dissertation abstract to this database? It only takes a minute!
Search for abstracts by subject, author or institution
MulTweEmo: A New Resource and Experiments on CLIP-based Multimodal Emotion Recognition
by ROBERTO CANNARELLA
| Institution: | Pisa University |
|---|---|
| Department: | INFORMATICA UMANISTICA |
| Degree: | FILOLOGIA, LETTERATURA E LINGUISTICA |
| Year: | 2022 |
| Keywords: | multimodal semantics; grounded cognition; emotion recognition; multimodal sentiment analysis; CLIP language model; language resources |
| Posted: | 3/25/2025 |
| Record ID: | 2286424 |
| Full text PDF: | http://etd.adm.unipi.it/theses/available/etd-08292022-105130/ |
This work focuses on the image-text emotion recognition (ITER) task, which consists in training NLP models that, thanks to the combination of visual and textual information, can predict the emotion associated with a multimodal document. Since research on ITER is still scarce, the first aim of this work is to better frame it within a solid theoretical framework. To do that, Chapter 1 presents contributions from cognitive linguistics and communication studies arguing that our language use is inherently multimodal and grounded on perceptual experience. The chapter also revises literature on how extra-linguistic information can be integrated into typical word embeddings, which yields the so-called visual-semantic embeddings. As a task, ITER is deeply connected with a few others: textual emotion recognition (ER), object recognition, and image emotion recognition (IER). For each of them, previous research, methods, and theoretical considerations are described in Chapter 2. Chapter 3, then, focuses on image-polarity classification, which is a task deeply connected with ITER, and finally reports the little existing research on ITER. The rest of the work reports the creation of a new multimodal resource, MulTweEmo, and discusses the related experiments. Chapter 4 describes the creation process, including how labels have been defined, both automatically and, for a portion of the data, through crowdsourcing. Chapter 5 describes several experiments involving the use of the dataset: since the resource contains texts associated with pictures, it has been used to feed classifiers with both unimodal (textual or visual) and multimodal embeddings. All embeddings have been created by using the CLIP model. Experiments show that multimodal classifiers outperform the others across all experimental setups, suggesting that multimodality can be advantageous in tasks involving the analysis of the affective value of documents.
Want to add your dissertation abstract to this database? It only takes a minute!
Search for abstracts by subject, author or institution
|
|
Harriet Beecher Stowe's "Uncle Tom's Cabin" in Ara...
Challenges of Cross-Cultural Translation
|
|
|
An Analysis of the Knowledge and Use of English Co...
|
|
|
Literature and Education
Proposal of an English Literature Program for Prim...
|
|
|
Putting Assessment for Learning (AfL) into Practic...
|
|
|
A Case Study on the Impact of Weblogs on the Writi...
|
|
|
Thucydides and US Foreign Policy Debates after the...
|
|
|
Quantificational Modification
The Semantics of Totality and Proportionality
|
|
|
Language Choice in Interracial Marriages
The Case of Filipino-Malaysian Couples
|