Papers by Gareth Jones
MultiMWE: Building a Multi-lingual Multi-Word Expression (MWE) Parallel Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research . |
| Approach: | They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora. |
| Outcome: | The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs . |
Tempo-Lexical Context Driven Word Embedding for Cross-Session Search Task Extraction (N18-1)
Copied to clipboard
| Challenge: | Existing work on task extraction has focused on identifying tasks within a single session . but, we aim to identify tasks that span across multiple sessions. |
| Approach: | They propose to embed query words into query vectors to capture task semantics . they propose to use query vector embedding to predict whether a session is a part of a broader search task . |
| Outcome: | The proposed method improves task extraction efficiency over existing methods . it can predict whether a session is part of a broader complex search task . |
100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)
Copied to clipboard
Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, Rosie Jones
| Challenge: | Podcasts are a large and growing repository of spoken audio. |
| Approach: | They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain . |
| Outcome: | The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles . |
Word-Node2Vec: Improving Word Embedding with Document-Level Non-Local Word Co-occurrences (N19-1)
Copied to clipboard
| Challenge: | Existing word embedding algorithms make a strong assumption that words are semantically related only if they co-occur locally within a window of fixed size. |
| Approach: | They propose a graph-based word embedding method that relies on locality to capture the semantic association between words that co-occur frequently but non-locally within documents. |
| Outcome: | The proposed method outperforms word2vec and glove on a range of different tasks, such as predicting word-pair similarity, word analogy and concept categorization. |
Development of an Annotated Multimodal Dataset for the Investigation of Classification and Summarisation of Presentations using High-Level Paralinguistic Features (L18-1)
Copied to clipboard
| Challenge: | Existing summarisation methods take no account of multimodal high-level paralinguistic features which form part of audio-visual presentations. |
| Approach: | They propose to use audiovisual recordings to extract paralinguistic features from audio recordings . they use manual annotations to help users find relevant material . |
| Outcome: | The proposed method can identify the most important or emphasised material within a presentation. |
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)
Copied to clipboard
| Challenge: | Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed . |
| Approach: | They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues . |
| Outcome: | The proposed method is highly reliable while remaining feasible and low cost. |