Abstract
We describe an architecture for speech recognition based interactive toys and discuss the strategies we have adopted to deal with the requirements for the speech recognizer imposed by this application. In particular, we focus on the fact that speech recognizers used in interactive toys must deal with users whose age ranges from children to adults. The large variations in vocal tract length between children and adults can significantly degrade the performance of speech recognizers. We compare two approaches to vocal tract length normalization: feature-based VTLN and model-based VTLN. We describe why intuitively, one might expect that due to the coarser frequency information used by the model-based approach, that feature-based VTLN would outperform modelbased VTLN. However, our results indicate that there is very little difference in performance between the two schemes.
| Original language | English |
|---|---|
| Title of host publication | Active Media Technology - 6th International Computer Science Conference, AMT 2001, Proceedings |
| Editors | Jiming Liu, Pong C. Yuen, Chun-hung Li, Joseph Ng, Toru Ishida |
| Publisher | Springer Verlag |
| Pages | 134-143 |
| Number of pages | 10 |
| ISBN (Electronic) | 9783540430353 |
| DOIs | |
| Publication status | Published - 2001 |
| Event | 6th International Computer Science Conference on Active Media Technology, AMT 2001 - Hong Kong, China Duration: 18 Dec 2001 → 20 Dec 2001 |
Publication series
| Name | Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) |
|---|---|
| Volume | 2252 |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | 6th International Computer Science Conference on Active Media Technology, AMT 2001 |
|---|---|
| Country/Territory | China |
| City | Hong Kong |
| Period | 18/12/01 → 20/12/01 |
Bibliographical note
Publisher Copyright:© Springer-Verlag Berlin Heidelberg 2001.