ISCApad #208 |
Saturday, October 10, 2015 by Chris Wellekens |
Speechocean – update (August 2015):
Speechocean: A global language resources and data services supplier
Speechocean is one of the world well-known language related resources & services provider in the fields of Human Computer Interaction and Human Language Technology. At present, we can provide data services with 110+ languages and dialects across the world.
KingLine Data Center ---Data Sharing Platform Kingline Data Center is operated and supervised by Speechocean, which is mainly focused on language resources creating and providing for research and development of human language technology. These diversified corpora are widely used for the research and development in the fields of Speech Recognition, Speech Synthesis, Natural Language Processing, Machine Translation, Web Search, etc. All corpora are openly accessible for users all over the world, including users from scientific research institutions, enterprises or individuals. For more detailed information, please visit our website: http://kingline.speechocean.com
New released corpora: ID: King-ASR-143 This is a 3-channel Mexican Spanish mobile speech database, which is collected over three mobile phone simultaneously (android mobiles, iPhone and windows phones) in Mexico. This database was performed in a quiet environment.
ID: King-ASR-281 This is a 4-channel Spanish desktop speech database, which is collected over 4 different microphones simultaneously. The project was performed in Argentina; cover all the cities, for example: BuenosAires, Cordoba, Lanus, Cordoba... Each Speaker was recorded around 300 sentences which were selected from a pool of phonetically rich sentences in approximate 80 minutes as natural as possible. The recording was performed in a quiet office environment. This database is performed in quiet office environment. The corpus contains the recordings of 236,232 utterances of Spanish speech data which were from 200 speakers. The pure recording time is about 358 hours (4-channel), including the leading silence (about 500 ms) and the trailing silence (about 500 ms). The total size of this database is 141 GB. A pronunciation lexicon with a phonemic transcription in SAMPA was carefully made by covering all the words in the transcription files.
ID: King-ASR-290 This is a 3-channel Chilean Spanish speech database, which is collected over 3 different mobile operating systems: iOS, Android and Windows Phone platform. The project was performed in Chile, cover all the main cities. For example: Santiago, Rancagua, Antofagasta and Viña.
Contact Information Xianfeng Cheng VP Tel: +86-10-62660928; +86-10-62660053 ext.8080 Mobile: +86 13681432590 Skype: xianfeng.cheng1 Email: chengxianfeng@speechocean.com; cxfxy0cxfxy0@gmail.com Website: www.speechocean.com
|
Back | Top |