| Challenge: | a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab. |
| Approach: | They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion. |
| Outcome: | The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus. |
Similar Papers
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups. |
| Approach: | They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks. |
| Outcome: | The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants. |
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)
Copied to clipboard
| Challenge: | Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Approach: | They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Outcome: | The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora. |
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)
Copied to clipboard
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner
| Challenge: | Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets. |
| Approach: | They propose to document a dataset created by applying filters to a single snapshot of Common Crawl. |
| Outcome: | The proposed dataset shows that blocklist filtering removes text from minority individuals and patents. |
The AICO Multimodal Corpus – Data Collection and Preliminary Analyses (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on human multimodal behaviour in interactions with a human or a robot partner are limited. |
| Approach: | They describe the first explorative research on the AICO Multimodal Corpus, which contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions. |
| Outcome: | The AICO Multimodal Corpus contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)
Copied to clipboard
Ruochen Zhao, Hailin Chen, Weishi Wang, Fangkai Jiao, Xuan Long Do, Chengwei Qin, Bosheng Ding, Xiaobao Guo, Minzhi Li, Xingxuan Li, Shafiq Joty
| Challenge: | Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities. |
| Approach: | They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs. |
| Outcome: | The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content. |
A Repository of Corpora for Summarization (L18-1)
Copied to clipboard
| Challenge: | Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task. |
| Approach: | They propose a repository containing corpora available to train and evaluate automatic summarization systems. |
| Outcome: | The proposed system is based on a repository of corpora available for summarization tasks. |