Conference Proceeding

Gathering a corpus of multimodal computer-mediated meetings with focus on text and audio interaction

Details

Citation

Luz S, Bouamrane M & Masoodian M (2006) Gathering a corpus of multimodal computer-mediated meetings with focus on text and audio interaction. In: Proceedings of The fifth international conference on Language Resources and Evaluation, LREC 2006. The fifth international conference on Language Resources and Evaluation, LREC 2006, Genoa, 22.05.2006-28.05.2006. European Language Resources Association (ELRA), p. 6. http://www.lrec-conf.org/proceedings/lrec2006/pdf/510_pdf.pdf

Abstract
In this paper we describe the gathering of a corpus of synchronised speech and text interaction over the network. The data collection scenarios characterise audio meetings with a significant textual component. Unlike existing meeting corpora, the corpus described in this paper emphasises temporal relationships between speech and text media streams. This is achieved through detailed logging and time stamping of text editing operations, actions on shared user interface widgets and gesturing, as well as generation of speech activity profiles. A set of tools has been developed specifically for these purposes which can be used as a data collection platform for the development of meeting browsers. The data gathered to date consists of nearly 30 hours of recorded audio and time stamped editing operations and gestures.

Keywords
Corpus, Multimedia Meeting Recordings, Speech, Text, Meeting

StatusPublished
Publication date22/05/2006
Publication date online22/05/2006
PublisherEuropean Language Resources Association (ELRA)
Publisher URLhttp://www.lrec-conf.org/…/pdf/510_pdf.pdf
ConferenceThe fifth international conference on Language Resources and Evaluation, LREC 2006
Conference locationGenoa
Dates

People (1)

Professor Matt-Mouley Bouamrane

Professor Matt-Mouley Bouamrane

Professor in Health/Social Informatics, Computing Science