Reference model
Baseline
Grad-TTS with emotion embedding
Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces
Nara Institute of Science and Technology, Japan
The same emotional expression can be perceived differently across listeners and cultures. We personalize arousal–valence (A–V) coordinates through listener feedback, adapting emotional speech without retraining the TTS backbone.
Compare Grad-TTS with an emotion embedding against the proposed emotion controller and ground-truth speech, before interactive personalization.
Headphones recommended. Playing a sample pauses the previous one.
Reference model
Grad-TTS with emotion embedding
Our method
Grad-TTS with emotion controller
Recorded speech
Reference recording
Listen to the averaged U.S. reference mapping alongside an individual listener’s personalized A–V mapping. Each pair demonstrates a different target emotion and listener background.
Population mapping
A–V values from the U.S.-based reference dataset
Individual mapping
A–V values adapted to an individual listener
Population mapping
A–V values from the U.S.-based reference dataset
Individual mapping
A–V values adapted to an individual listener
Population mapping
A–V values from the U.S.-based reference dataset
Individual mapping
A–V values adapted to an individual listener
Explore the same target emotion with culture-specific A–V mappings. Country labels refer to participant groups; the speech samples are in English.
Participant group
Culture-specific A–V mapping
Participant group
Culture-specific A–V mapping
Participant group
Culture-specific A–V mapping