On the Generalization of Shannon Entropy for Speech Recognition - Sorbonne Université Access content directly
Conference Papers Year : 2012

On the Generalization of Shannon Entropy for Speech Recognition


This paper introduces an entropy-based spectral representation as a measure of the degree of noisiness in audio signals, complementary to the standard MFCCs for audio and speech recognition. The proposed representation is based on the Rényi entropy, which is a generalization of the Shannon entropy. In audio signal representation, Rényi entropy presents the advantage of focusing either on the harmonic content (prominent amplitude within a distribution) or on the noise content (equal distribution of amplitudes). The proposed representation outperforms all other noisiness measures - including Shannon and Wiener entropies - in a large-scale classification of vocal effort (whispered-soft/normal/loud-shouted) in the real scenario of multi-language massive role-playing video games. The improvement is around 10% in relative error reduction, and is particularly significant for the recognition of noisy speech - i.e., whispery/breathy speech. This confirms the role of noisiness for speech recognition, and will further be extended to the classification of voice quality for the design of an automatic voice casting system in video games.
Fichier principal
Vignette du fichier
SLT12_NO_ML.pdf (1.45 Mo) Télécharger le fichier
Origin Files produced by the author(s)

Dates and versions

hal-00737653 , version 1 (02-10-2012)


  • HAL Id : hal-00737653 , version 1


Nicolas Obin, Marco Liuni. On the Generalization of Shannon Entropy for Speech Recognition. IEEE workshop on Spoken Language Technology, Dec 2012, United States. ⟨hal-00737653⟩
288 View
484 Download


Gmail Mastodon Facebook X LinkedIn More