Papers by Carl Vogel
Multilingual Word Segmentation: Training Many Language-Specific Tokenizers Smoothly Thanks to the Universal Dependencies Corpus (L18-1)
Copied to clipboard
| Challenge: | Towards language scalability, major progress has been achieved in multilingual language technology in recent years. |
| Approach: | They propose a tokenizer that can be trained from any Universal Dependencies corpus dataset . they argue that tokenization should be seen as a supervised task and scalability requires a software engineering process across languages. |
| Outcome: | The proposed tokenizer can be trained from any dataset in the corpus UD2 . the proposed software tool relies on elephant to perform the training . |
A Diachronic Corpus for Literary Style Analysis (L18-1)
Copied to clipboard
| Challenge: | Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data . |
| Approach: | They propose a resource for diachronic style analysis in particular the analysis of literary authors over time. |
| Outcome: | The proposed resource can be used to analyze literary authors over time. |
Chats and Chunks: Annotation and Analysis of Multiparty Long Casual Conversations (L18-1)
Copied to clipboard
| Challenge: | dyadic conversations are attracting more interest with attempts to build more friendly and natural spoken dialog systems. |
| Approach: | They describe the collection, organization, and annotation of structural chat and chunk phases in three existing corpora and analyse their preliminary results to find that chunk dominates as conversations get longer. |
| Outcome: | The results show that chunk dominates conversations as they get longer . |
Speech Rate Calculations with Short Utterances: A Study from a Speech-to-Speech, Machine Translation Mediated Map Task (L18-1)
Copied to clipboard
| Challenge: | Computer mediated multi-lingual communication is becoming more frequent. |
| Approach: | They propose a method to verify if an utterance within a corpus is pronounced at a fast or slow pace. |
| Outcome: | The proposed method provides a value for the utterance speech rate in a corpus of short utterations. |
Mutual Gaze and Linguistic Repetition in a Multimodal Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | a study of linguistic repetitions and mutual understanding is conducted . we find no compelling correlation between mutual gaze and duration of the event . |
| Approach: | They investigate the correlation between mutual gaze and linguistic repetition, a form of alignment, which they take as evidence of mutual understanding. |
| Outcome: | The proposed method is based on the Multisimo corpus, a multimodal corpus which provides authentic task-based interactions among three participants. |
An Evaluation Method for Diachronic Word Sense Induction (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to detect semantic shifts across time are based on time-stamped annotated biomedical data . dynamic behaviour of words contributes to semantic ambiguity, which is a challenge in many NLP tasks. |
| Approach: | They propose an evaluation method based on large-scale time-stamped biomedical data . they propose a model which represents the temporal dimension of the task . |
| Outcome: | The proposed method is applied to two recent DWSI systems . it provides an in-depth analysis of the models . |
Is It Dish Washer Safe? Automatically Answering “Yes/No” Questions Using Customer Reviews (N19-3)
Copied to clipboard
| Challenge: | Using Amazon reviews, we find that the answer to a question is only in 45% of cases. |
| Approach: | They combine Amazon reviews with consumer reviews and manually analyse 400 questions from four domains to find that reviews directly contain the answer to the question . they then compare QA systems that use reviews in addition to the questions to see if they can be useful for other question types. |
| Outcome: | The proposed system outperforms the chance baseline but not by a large margin. |
English Machine Reading Comprehension Datasets: A Survey (2021.emnlp-main)
Copied to clipboard
| Challenge: | a survey of English Machine Reading Comprehension datasets is carried out . the aim is to provide a concise yet informative overview of the landscape . |
| Approach: | They survey 60 English Machine Reading Comprehension datasets to provide a resource for other researchers interested in this problem. |
| Outcome: | The proposed survey covers 60 English MRC datasets with a view to providing a resource for other researchers interested in the problem. |
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups. |
| Approach: | They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks. |
| Outcome: | The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants. |