Perception of synthetic emotion expressions in speech: Categorical and dimensional annotations
conference paper
In this paper, both categorical and dimensional annotations have been made of neutral and emotional speech synthesis (anger, fear, sad, happy and relaxed). With various prosodic emotion manipulation techniques we found emotion classification rates of 40%, which is significantly above chance level (17%). The classification rates are higher for sentences that have a semantics matching the synthetic emotion. By manipulating the pitch and duration, differences in arousal were perceived whereas differences in valence were hardly perceived. Of the investigated emotion manipulation methods, EmoFilt and EmoSpeak performed very similar, except for the emotion fear. Copy synthesis did not perform well, probably caused by suboptimal alignments and the use of multiple speakers. ©2009 IEEE.
TNO Identifier
346444
ISBN
9781424447992
Article nr.
No.: 5349594
Source title
3rd International Conference on Affective Computing and Intelligent Interaction and Workshops, ACII 2009, 10 September 2009 through 12 September 2009, Amsterdam, Conference code: 79475
Files
To receive the publication files, please send an e-mail request to TNO Repository.