ISCA Services

ISCA - International Speech
Communication Association

ISCApad Archive » 2010 » ISCApad #150 » Resources

ISCApad #150

Tuesday, December 07, 2010 by Chris Wellekens

5 Resources

5-1 Books

5-1-1

Spoken Language Processing

Spoken Language Processing, edited by Joseph Mariani (IMMI and
LIMSI-CNRS, France). ISBN: 9781848210318. January 2009. Hardback 504 pp

Publisher ISTE-Wiley

Speech processing addresses various scientific and technological areas. It includes speech analysis and variable rate coding, in order to store or transmit speech. It also covers speech synthesis, especially from text, speech recognition, including speaker and language identification, and spoken language understanding. This book covers the following topics: how to realize speech production and perception systems, how to synthesize and understand speech using state-of-the-art methods in signal processing, pattern recognition, stochastic modeling, computational linguistics and human factor studies.

More on its content can be found at
http://www.iste.co.uk/index.php?f=a&ACTION=View&id=150

Back

Top

5-1-2

L'imagerie medicale pour l'etude de la parole

Alain Marchal, Christian Cave

Eds Hermes Lavoisier

99 euros • 304 pages • 16 x 24 • 2009 • ISBN : 978-2-7462-2235-9

Du miroir laryngé à la vidéofibroscopie actuelle, de la prise d'empreintes statiques à la palatographie dynamique, des débuts de la radiographie jusqu'à l'imagerie par résonance magnétique ou la magnétoencéphalographie, cet ouvrage passe en revue les différentes techniques d'imagerie utilisées pour étudier la parole tant du point de vue de la production que de celui de la perception. Les avantages et inconvénients ainsi que les limites de chaque technique sont passés en revue, tout en présentant les principaux résultats acquis avec chacune d'entre elles ainsi que leurs perspectives d'évolution. Écrit par des spécialistes soucieux d'être accessibles à un large public, cet ouvrage s'adresse à tous ceux qui étudient ou abordent la parole dans leurs activités professionnelles comme les phoniatres, ORL, orthophonistes et bien sûr les phonéticiens et les linguistes.

Back

Top

5-1-3

Korpusbasierte Sprachverarbeitung

Author: Christoph Draxler
Title: Korpusbasierte Sprachverarbeitung
Publisher: Narr Francke Attempto Verlag Tübingen
Year: 2008
Link: http://www.narr.de/details.php?catp=&p_id=16394

Summary: Spoken language is a major area of linguistic research and speech technology development. This handbook presents an introduction to the technical foundations and shows how speech data is collected, annotated, analysed, and made accessible in the form of speech databases. The book focuses on web-based procedures for the recording and processing of high quality speech data, and it is intended as a desktop reference for practical recording and annotation work. A chapter is devoted to the Ph@ttSessionz database, the first large-scale speech data collection (860+ speakers, 40 locations in Germany) performed via the Internet. The companion web site (http://www.narr-studienbuecher.de/Draxler/index.html) contains audio examples, software tools, solutions to the exercises, important links, and checklists.

Back

Top

5-1-4

Linear Predictive Coding and the Internet Protocol, by Robert M. Gray

Linear Predictive Coding and the Internet Protocol, by Robert M. Gray, a special edition hardback book from Foundations and Trends in Signal Processing (FnT SP). The book brings together two forthcoming issues of FnT SP, the first being a survey of LPC, the second a unique history of realtime digital speech on packet networks.

Volume 3, Issue 3

A Survey of Linear Predictive Coding: Part 1 of LPC and the IP

By Robert M. Gray (Stanford University)

http://www.nowpublishers.com/product.aspx?product=SIG&doi=2000000029

Volume 3, Issue 4

A History of Realtime Digital Speech on Packet Networks: Part 2 of LPC and the IP

By Robert M. Gray (Stanford University)

http://www.nowpublishers.com/product.aspx?product=SIG&doi=2000000036

The links above will take you to the article abstracts.

Back

Top

5-2 Database

5-2-1

Bell System Technical Journal (1922-1983) available .

I received this very good news from Joseph P. Campbell (MIT Lincoln
Laboratory):

The entire Bell System Technical Journal from 1922--1983 is now
available on line! It is offered in PDF format sorted by year, volume, and
issue; and it is searchable:

http://bstj.bell-labs.com/

BSTJ holds a wealth of consolidated information of outstanding contributions
from Bell Labs over the years. For example, check out Fletcher's 1922
article 'The Nature of Speech and Its Interpretation', BSTJ, vol 1, no 1.
Also see Shannon's landmark paper and much, much more!

As Joe pointed out, the current generation of researchers didn't grow up
with BSTJs under their bed like we both did, but for the older generation
that remembers these great articles, this is indeed very good news.

--Isabel Trancoso

Back

Top

5-2-2

ELRA Language Resources Catalogue Update (June 2010)

ELRA is happy to announce that 2 new Speech Desktop/Microphone resources, 1 new Terminological Resource and 1 Written Corpus are now available in its catalogue:

ELRA-S0305 EPAC Corpus: orthographic transcriptions

This corpus consists of approx. 100 hours of manual orthographic transcriptions, which were produced from 1,677 hours of non transcribed recordings from the ESTER Evaluation Campaign (Technolangue programme). This corpus also consists of automatic transcriptions of the full 1,677 hours.

For more information, see: http://catalog.elra.info/product_info.php?products_id=1119

ELRA-S0307 BABEL Polish database

The BABEL Polish Database is a speech database that was produced by a research consortium funded by the European Union under the COPERNICUS programme (COPERNICUS Project 1304). It consists of the basic 'common' set which contains the Many Talker Set (30 males, 30 females), the Few Talker Set (5 males, 5 females), the Very Few Talker Set (1 male, 1 female).

For more information, see: http://catalog.elra.info/product_info.php?products_id=1120

ELRA-T0374 Terminology database of natural sciences

This dictionary covers the three kingdoms: Animal, Vegetal, Mineral. It contains 50,000 species with numerous synonyms in French, English and Latin and many breeds and varieties. Minerals are given with their chemical formula. About 7,900 definitions in French are included. It also includes synonyms and linguistic variants.

For more information, see: http://catalog.elra.info/product_info.php?products_id=1121

ELRA-W0053 Catalan-Spanish Parallel Corpus

This corpus contains more than 100 million words and it contains 10 years of bilingual articles from “El Periódico de Catalunya”. The data are aligned at sentence level and stored in text files, in a one sentence per line basis. The data are provided in plain text, with no encoding whatsoever.

For more information, see: http://catalog.elra.info/product_info.php?products_id=1122

******

Moreover, please note that the content of the following 3 Terminological Resources has been updated and their prices have been revised:

ELRA-T0102 Terminology database of expressions

This resource comprises over about 26,000-30,000 expressions, such as sayings, proverbs, idioms, slogans, citations, exclamations, onomatopoeias and figurative expressions of French and English. Several grammatical topics that are included in some sentences are also handled. This resource contains synonyms. The DISCIPLINE field refers to the expression category: proverbs, idioms, postposition verbs.

For more information, see: http://catalog.elra.info/product_info.php?products_id=114

ELRA-T0103 Terminology database of finance

For more information, see: http://catalog.elra.info/product_info.php?products_id=115

ELRA-T0367 Terminology database of telecommunication

This resource comprises over 89,200 entries in the field of telecommunication. It also contains many synonyms and abbreviations in both languages, as well as meaning, case or applications for polysemic terms.

For more information, see: http://catalog.elra.info/product_info.php?products_id=659

For more information on the catalogue, please contact Valérie Mapelli mailto:mapelli@elda.org

Visit our On-line Catalogue: http://catalog.elra.info

Visit the Universal Catalogue: http://universal.elra.info

Archives of ELRA Language Resources Catalogue Updates: http://www.elra.info/LRs-Announcements.html

*****************************************************************

ELRA - Language Resources Catalogue - Update

*****************************************************************

In the framework of our ongoing campaign for updating and reducing the prices of the language resources distributed in the ELRA catalogue, ELRA is happy to announce that the prices for the following resources have been substantially reduced:

ELRA-S0074 British English SpeechDat(II) MDB-1000

This speech database contains the recordings of 1,000 British speakers recorded over the British mobile telephone network. Each speaker uttered around 40 read and spontaneous items.

For more information, see: http://catalog.elra.info/product_info.php?products_id=723

ELRA-S0075 Welsh SpeechDat(II) FDB-2000

This speech database contains the recordings of 2,000 Welsh speakers recorded over the British fixed telephone network. Each speaker uttered around 40 read and spontaneous items.

For more information, see: http://catalog.elra.info/product_info.php?products_id=557

ELRA-S0101 Spanish SpeechDat(II) FDB-1000

This speech database contains the recordings of 1,000 Castillan Spanish speakers recorded over the Spanish fixed telephone network. Each speaker uttered around 40 read and spontaneous items.

This database is a subset of the Spanish SpeechDat(II) FDB-4000 (ref. ELRA-S0102).

For more information, see: http://catalog.elra.info/product_info.php?products_id=726

ELRA-S0102 Spanish SpeechDat(II) FDB-4000

This speech database contains the recordings of 4,000 Castillan Spanish speakers recorded over the Spanish fixed telephone network. Each speaker uttered around 40 read and spontaneous items.

This database includes the Spanish SpeechDat(II) FDB-1000 (ref. ELRA-S0101).

For more information, see: http://catalog.elra.info/product_info.php?products_id=727

ELRA-S0140 Spanish SpeechDat-Car database

The Spanish SpeechDat-Car database contains the recordings in a car of 306 speakers, who uttered around 120 read and spontaneous items. Recordings have been made through 5 different channels, of which 4 were in-car microphones (1 close-talk microphone, 3 far-talk microphones) and 1 channel over the GSM network.

For more information, see: http://catalog.elra.info/product_info.php?products_id=690

ELRA-S0141 SALA Spanish Venezuelan Database

This speech database contains the recordings of 1,000 Venezuelan speakers recorded over the Venezuelan fixed telephone network. Each speaker uttered around 50 read and spontaneous items.

For more information, see: http://catalog.elra.info/product_info.php?products_id=736

ELRA-S0297 Hungarian Speecon database

The Hungarian Speecon database comprises the recordings of 555 adult Hungarian speakers and 50 child Hungarian speakers who uttered respectively over 290 items and 210 items (read and spontaneous).

For more information, see: http://catalog.elra.info/product_info.php?products_id=1094

ELRA-S0298 Czech Speecon database

The Czech Speecon database comprises the recordings of 550 adult Czech speakers and 50 child Czech speakers who uttered respectively over 290 items and 210 items (read and spontaneous).

For more information, see: http://catalog.elra.info/product_info.php?products_id=1095

For more information on the catalogue, please contact Valérie Mapelli mailto:mapelli@elda.org

Visit our On-line Catalogue: http://catalog.elra.info

Visit the Universal Catalogue: http://universal.elra.info

Archives of ELRA Language Resources Catalogue Updates: http://www.elra.info/LRs-Announcements.html

Back

Top

5-2-3

LDC Newsletter (November 2010)

Early Renewal Discounts for Membership Year (MY) 2011 -

New publications:

LDC2010T13
- Arabic Treebank: Part 1 v 4.1 -

LDC2010T21
- NIST 2008 Open Machine Translation (OpenMT) Evaluation -

Early Renewal Discounts for Membership Year (MY) 2011

LDC values the significant contribution LDC members make through their continued support of the consortium. We would like to invite new members, as well as all current and previous members of LDC, to renew for Membership Year (MY) 2011. For MY2011, LDC is pleased to maintain membership fees at last year’s rates – membership fees will not increase. Additionally, for the third straight year, LDC will extend discounts on membership fees to members who keep their membership current and who join early in the year.

The details of our Early Renewal Discounts for MY2011 are as follows:

Organizations who joined for MY2010, will receive a 5% discount when renewing. This discount will apply throughout 2011, regardless of time of renewal. MY2010 members renewing before March 1, 2011 will receive an additional 5% discount, for a total 10% discount off the membership fee.
New members as well as organizations who did not join for MY2010, but who held membership in any of the previous MYs (1993-2009), will also be eligible for a 5% discount provided that they join/renew before March 1, 2011.

The following table provides exact pricing information.

		MY2011 Fee	MY2011 Fee with 5% Discount *	MY2011 Fee with 10% Discount **
Not-for-Profit
	Standard	US$2400	US$2280	US$2160
	Subscription	US$3850	US$3657.50	US$3465
For-Profit
	Standard	US$24000	US$22800	US$21600
	Subscription	US$27500	US$26125	US$24750

* For new members, MY2010 Members renewing for MY2011, and any previous year Member who renews before March 1, 2011

** For MY2010 Members renewing before March 1, 2011

Publications for MY2011 are still being planned and here are the working titles of data sets we intend to provide:

Arabic Gigaword Fifth Edition	English Gigaword Fifth Edition
Chinese Gigaword Fifth Edition	Indian Language POS Tagset: Sanskrit
Digital Archive of Southern Speech	OntoNotes 4.0

In addition to receiving new publications, current year members of the LDC also enjoy the benefit of licensing older data at reduced costs; current year for-profit members may use most data for commercial applications.

This past year, the LDC members who joined early or kept their membership current saved almost US$60,000 collectively on membership fees. In fact, almost 90% of our members for MY2010 didn't pay full price for membership! Be sure to keep an eye on your mail - all previous and current LDC members have been sent an invitation to join letter and renewal invoice for MY2011. Renew early for MY2011 to save today!

New Publications

(1) Arabic Treebank: Part 1 v 4.1 was developed at LDC. It consists of 734 newswire stories from Agence France Presse with part-of-speech , morphology, gloss and syntactic treebank annotation in accordance with the Penn Arabic Treebank (PATB) Guidelines developed in 2008 and 2009. This release represents a significant revision of LDC's previous ATB1 publications: Arabic Treebank: Part 1 v 2.0 (LDC2003T06) and Arabic Treebank: Part 1 v 3.0 (POS with full vocalization + syntactic analysis) (LDC2005T02).

The ongoing PATB project supports research in Arabic-language natural language processing and human language technology development. The methodology and work leading to the release of this publication are described in detail in the documentation accompanying this corpus and in two research papers: Enhancing the Arabic Treebank: A Collaborative Effort toward New Annotation Guidelines and Consistent and Flexible Integration of Morphological Annotation in the Arabic Treebank.

ATB1 v 4.1 contains a total of 145,386 tokens before clitics are split, and 167,280 tokens after clitics are separated for the treebank annotation. Arabic Treebank: Part 1 v 4.1 is distributed on one DVD-ROM.

2010 Subscription Members will automatically receive two copies of this corpus. 2010 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for US$4500.

(2) NIST 2008 Open Machine Translation (OpenMT) Evaluation is a package containing source data, reference translations and scoring software used in the NIST 2008 OpenMT evaluation. It is designed to help evaluate the effectiveness of machine translation systems. The package was compiled and scoring software was developed by researchers at NIST, making use of broadcast, newswire and web data and reference translations collected and developed by LDC.

The 2008 task was to evaluate translation from Arabic to English, Chinese to English, English to Chinese (newswire only) and Urdu to English. Selected human reference translations and system translations for the NIST MT08 test sets are contained in NIST Open Machine Translation 2008 Evaluation (MT08) Selected Reference and System Translations LDC2010T01.

This release contains of 494 documents with corresponding sets of four separate human expert reference translations. The source data is comprised of Arabic, Chinese, English and Urdu news wire, broadcast and weblog and newsgroup data collected by LDC in 2007. The news wire and broadcast material are from Asharq Al-Awsat (Arabic), Agence France-Presse (Arabic, Chinese, English), Al-Ahram (Arabic), Al Hayat (Arabic), Assabah (Arabic), An Nahar (Arabic), Al-Quds Al-Arabi (Arabic), Xinhua News Agency (Arabic, Chinese, English), Central News Service (Chinese), Guangming Daily (Chinese), People's Daily (Chinese), People's Liberation Army Daily (Chinese), British Broadcasting Corporation (Urdu), Daily Jang (Urdu), Pakistan News Service (Urdu), Voice of America (Urdu), Associated Press (English), New York Times (English) and Los Angeles Times/Washington Post Newswire Service (English).

This evaluation kit includes a single Perl script (mteval-v11b.pl) that may be used to produce a translation quality score for one (or more) MT systems.

Additional information about these evaluations may be found at the NIST Open Machine Translation (OpenMT) Evaluation web site.

NIST 2008 Open Machine Translation (OpenMT) Evaluation is distributed via web download.

2010 Subscription Members will automatically receive two copies of this corpus on disc. 2010 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for US$150.

Back

Top

5-2-4

ELDA Distribution Campaign 2010

ELDA Distribution Campaign 2010
*****************************************************************

ELDA is launching a special distribution campaign offering very favorable conditions for the language resources acquisition,
including discounts on public prices, from the ELRA Catalogue of Language Resources (see http://catalog.elra.info).

This offer will be open until the end of December 2010.

For more information on this offer, please contact Valérie Mapelli (mapelli@elda.org)

Visit our On-line Catalogue: http://catalog.elra.info
Visit the Universal Catalogue: http://universal.elra.info
Archives of ELRA Language Resources Catalogue Updates: http://www.elra.info/LRs-Announcements.html

Back

Top

5-3 Software

Organisation	Events	Membership	Help
> Board	> Interspeech	> Join - renew	> Sitemap
> Legal documents	> Workshops	> Membership directory	> Contact
> Logos			> FAQ
			> Privacy policy