ISCApad #164 |
Saturday, February 11, 2012 by Chris Wellekens |
5-1-1 | Robert M. Gray, Linear Predictive Coding and the Internet Protocol Linear Predictive Coding and the Internet Protocol, by Robert M. Gray, a special edition hardback book from Foundations and Trends in Signal Processing (FnT SP). The book brings together two forthcoming issues of FnT SP, the first being a survey of LPC, the second a unique history of realtime digital speech on packet networks.
Volume 3, Issue 3 A Survey of Linear Predictive Coding: Part 1 of LPC and the IP By Robert M. Gray (Stanford University) http://www.nowpublishers.com/product.aspx?product=SIG&doi=2000000029
Volume 3, Issue 4
A History of Realtime Digital Speech on Packet Networks: Part 2 of LPC and the IP By Robert M. Gray (Stanford University) http://www.nowpublishers.com/product.aspx?product=SIG&doi=2000000036
The links above will take you to the article abstracts.
| ||
5-1-2 | M. Embarki and M. Ennaji, Modern Trends in Arabic Dialectology Modern Trends in Arabic Dialectology,
| ||
5-1-3 | Gokhan Tur , R De Mori, Spoken Language Understanding: Systems for Extracting Semantic Information from Speech Title: Spoken Language Understanding: Systems for Extracting Semantic Information from Speech Editors: Gokhan Tur and Renato De Mori Web: http://www.wiley.com/WileyCDA/WileyTitle/productCd-0470688246.html Brief Description (please use as you see fit): Spoken language understanding (SLU) is an emerging field in between speech and language processing, investigating human/ machine and human/ human communication by leveraging technologies from signal processing, pattern recognition, machine learning and artificial intelligence. SLU systems are designed to extract the meaning from speech utterances and its applications are vast, from voice search in mobile devices to meeting summarization, attracting interest from both commercial and academic sectors. Both human/machine and human/human communications can benefit from the application of SLU, using differing tasks and approaches to better understand and utilize such communications. This book covers the state-of-the-art approaches for the most popular SLU tasks with chapters written by well-known researchers in the respective fields. Key features include: Presents a fully integrated view of the two distinct disciplines of speech processing and language processing for SLU tasks. Defines what is possible today for SLU as an enabling technology for enterprise (e.g., customer care centers or company meetings), and consumer (e.g., entertainment, mobile, car, robot, or smart environments) applications and outlines the key research areas. Provides a unique source of distilled information on methods for computer modeling of semantic information in human/machine and human/human conversations. This book can be successfully used for graduate courses in electronics engineering, computer science or computational linguistics. Moreover, technologists interested in processing spoken communications will find it a useful source of collated information of the topic drawn from the two distinct disciplines of speech processing and language processing under the new area of SLU.
| ||
5-1-4 | Jody Kreiman, Diana Van Lancker Sidtis ,Foundations of Voice Studies: An Interdisciplinary Approach to Voice Production and Perception Foundations of Voice Studies: An Interdisciplinary Approach to Voice Production and Perception
| ||
5-1-5 | G. Nick Clements and Rachid Ridouane, Where Do Phonological Features Come From?
Where Do Phonological Features Come From?
Edited by G. Nick Clements and Rachid Ridouane CNRS & Sorbonne-Nouvelle This volume offers a timely reconsideration of the function, content, and origin of phonological features, in a set of papers that is theoretically diverse yet thematically strongly coherent. Most of the papers were originally presented at the International Conference 'Where Do Features Come From?' held at the Sorbonne University, Paris, October 4-5, 2007. Several invited papers are included as well. The articles discuss issues concerning the mental status of distinctive features, their role in speech production and perception, the relation they bear to measurable physical properties in the articulatory and acoustic/auditory domains, and their role in language development. Multiple disciplinary perspectives are explored, including those of general linguistics, phonetic and speech sciences, and language acquisition. The larger goal was to address current issues in feature theory and to take a step towards synthesizing recent advances in order to present a current 'state of the art' of the field.
| ||
5-1-6 | Dorothea Kolossa and Reinhold Haeb-Umbach: Robust Speech Recognition of Uncertain or Missing DataTitle: Robust Speech Recognition of Uncertain or Missing Data Editors: Dorothea Kolossa and Reinhold Haeb-Umbach Publisher: Springer Year: 2011 ISBN 978-3-642-21316-8 Link: http://www.springer.com/engineering/signals/book/978-3-642-21316-8?detailsPage=authorsAndEditors Automatic speech recognition suffers from a lack of robustness with respect to noise, reverberation and interfering speech. The growing field of speech recognition in the presence of missing or uncertain input data seeks to ameliorate those problems by using not only a preprocessed speech signal but also an estimate of its reliability to selectively focus on those segments and features that are most reliable for recognition. This book presents the state of the art in recognition in the presence of uncertainty, offering examples that utilize uncertainty information for noise robustness, reverberation robustness, simultaneous recognition of multiple speech signals, and audiovisual speech recognition. The book is appropriate for scientists and researchers in the field of speech recognition who will find an overview of the state of the art in robust speech recognition, professionals working in speech recognition who will find strategies for improving recognition results in various conditions of mismatch, and lecturers of advanced courses on speech processing or speech recognition who will find a reference and a comprehensive introduction to the field. The book assumes an understanding of the fundamentals of speech recognition using Hidden Markov Models.
| ||
5-1-7 | Mohamed Embarki et Christelle Dodane: La coarticulation
LA COARTICULATION Mohamed Embarki et Christelle Dodane La parole est faite de gestes articulatoires complexes qui se chevauchent dans l’espace et dans le temps. Ces chevauchements, conceptualisés par le terme coarticulation, n’épargnent aucun articulateur. Ils sont repérables dans les mouvements de la mâchoire, des lèvres, de la langue, du voile du palais et des cordesvocales. La coarticulation est aussi attendue par l’auditeur, les segments coarticulés sont mieux perçus. Elle intervient dans les processus cognitifs et linguistiques d’encodage et de décodage de la parole. Bien plus qu’un simple processus, la coarticulation est un domaine de recherche structuré avec des concepts et des modèles propres. Cet ouvrage collectif réunit des contributions inédites de chercheurs internationaux abordant lacoarticulation des points de vue moteur, acoustique, perceptif et linguistique. C’est le premier ouvrage publié en langue française sur cette question et le premier à l’explorer dans différentes langues.
Collection : Langue & Parole, L'Harmattan ISBN : 978-2-296-55503-7 • 25 € • 260 pages
Mohamed Embarki Christelle Dodane
| ||
5-1-8 | Ben Gold, Nelson Morgan, Dan Ellis :Speech and Audio Signal Processing: Processing and Perception of Speech and Music [Digital]Speech and Audio Signal Processing: Processing and Perception of Speech and Music [2nd edition] Ben Gold, Nelson Morgan, Dan EllisDigital copy: http://www.amazon.com/Speech-Audio-Signal-Processing-Perception/dp/product-description/1118142888 Hardcopy available: http://www.amazon.com/Speech-Audio-Signal-Processing-Perception/dp/0470195363/ref=sr_1_1?s=books&ie=UTF8&qid=1319142964&sr=1-1
| ||
5-1-9 | Video Proceedings ERMITES 2011Actes vidéo des journées ERMITES 2011 'Décomposition Parcimonieuse, Contraction et Structuration pour l'Analyse de Scènes', sont en ligne sur : http://glotin.univ-tln.fr/ERMITES11 On y retrouve (en .mpg) la vingtaine d'heure des conférences de : Y. Bengio, Montréal «Apprentissage Non-Supervisé de Représentations Profondes » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_Y_Bengio_1sur4.mp4 ... S. Mallat, Paris « Scattering & Matching Pursuit for Acoustic Sources Separation » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_Mallat_1sur3.mp4 ... J.-P. Haton, Nancy « Analyse de Scène et Reconnaissance Stochastique de la Parole » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_JP_Haton_1sur4.mp4 ... M. Kowalski, Paris « Sparsity and structure for audio signal: a *-lasso therapy » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_Kowalski_1sur5.mp4 ... O. Adam, Paris « Estimation de Densité de Population de Baleines par Analyse de leurs Chants » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_Adam.mp4 X. Halkias, New-York « Detection and Tracking of Dolphin Vocalizations » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_Halkias.mp4 J. Razik, Toulon « Sparse coding : from speech to whales » http://lsis.univ-tln.fr/~glotin/ERMITES_2011_Razik.mp4 H. Glotin, Toulon « Suivi & reconstruction du comportement de cétacés par acoustique passive » ps : ERMITES 2012 portera sur la vision (Y. Lecun, Y. Thorpe, P. Courrieu, M Perreira, M. Van Gerven,...)
| ||
5-1-10 | Zeki Majeed Hassan and Barry Heselwood (Eds): Instrumental Studies in Arabic Phonetics Instrumental Studies in Arabic Phonetics
|