Interspeech 2026 · Audio demos

Beyond
One-Size-Fits-All.

Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

Wangzixi Zhou · Bagus Tris Atmaja · Sakriani Sakti

Nara Institute of Science and Technology, Japan

Emotion is personal.

The same emotional expression can be perceived differently across listeners and cultures. We personalize arousal–valence (A–V) coordinates through listener feedback, adapting emotional speech without retraining the TTS backbone.

01 /

Speech quality & emotional similarity

Compare Grad-TTS with an emotion embedding against the proposed emotion controller and ground-truth speech, before interactive personalization.

Headphones recommended. Playing a sample pauses the previous one.

Confused

Target emotion

Reference model

Baseline

Grad-TTS with emotion embedding

Recorded speech

Ground truth

Reference recording

02 /

Personalized emotional expression

Listen to the averaged U.S. reference mapping alongside an individual listener’s personalized A–V mapping. Each pair demonstrates a different target emotion and listener background.

Angry

Target emotion

Population mapping

Averaged reference

A–V values from the U.S.-based reference dataset

Sad

Target emotion

Population mapping

Averaged reference

A–V values from the U.S.-based reference dataset

Neutral

Target emotion

Population mapping

Averaged reference

A–V values from the U.S.-based reference dataset

03 /

Cross-cultural emotional expression

Explore the same target emotion with culture-specific A–V mappings. Country labels refer to participant groups; the speech samples are in English.

Happy

Target emotion

Participant group

Indonesia

Culture-specific A–V mapping

Participant group

Japan

Culture-specific A–V mapping

Participant group

China

Culture-specific A–V mapping