Compilation of Corpora for the Study of the Information Structure–Prosody Interface (L18-1)
Copied to clipboard
| Challenge: | empirical studies on the Information Structure-prosody interface are scarce . thematicity defines how content is packaged in terms of "what is being talked about" a different view on thematicality is advocated by I. Mel'uk in the context of the MTT. |
| Approach: | They propose a method for the compilation of annotated corpora to study the correspondence between Information Structure and prosody. |
| Outcome: | The proposed method is applied to a corpus of read speech in English annotated with hierarchical thematicity and automatically extracted prosodic parameters. |
Similar Papers
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)
Copied to clipboard
| Challenge: | Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Approach: | They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Outcome: | The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora. |
An Empirical Evaluation of Annotation Practices in Corpora from Language Documentation (2020.lrec-1)
Copied to clipboard
| Challenge: | Language documentation projects have produced substantial amounts of primary data from a wide variety of endangered languages. |
| Approach: | They propose to use common annotation conventions in existing corpora to facilitate their future processing. |
| Outcome: | The proposed formats are based on the common ELAN and Toolbox formats and are used to facilitate their future processing. |
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences. |
| Approach: | They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs. |
| Outcome: | The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering . |
Quantifying the redundancy between prosody and text (2023.emnlp-main)
Copied to clipboard
Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar Regev
| Challenge: | Existing studies suggest partial redundancy between prosody and linguistic information. |
| Approach: | They use large language models to estimate how much information is redundant between prosody and the words themselves. |
| Outcome: | The proposed model can predict prosodic features across prosodic features, including intensity, duration, pauses, and pitch contours. |
What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels (2026.acl-long)
Copied to clipboard
| Challenge: | Prosody—the melody of speech—conveys critical information often not captured by the words or text of a message. |
| Approach: | They propose an information-theoretic approach to quantify how much is conveyed by prosody that is not recoverable from text alone. |
| Outcome: | The proposed framework can quantify how much is conveyed by prosody that is not recoverable from text alone and crucially, what prosody conveys. |
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)
Copied to clipboard
| Challenge: | Terms are notoriously difficult to identify, both automatically and manually. |
| Approach: | They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information . |
| Outcome: | The proposed method provides a tool for evaluation and rich source of information about terms. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The GermaParl Corpus of Parliamentary Protocols (L18-1)
Copied to clipboard
| Challenge: | Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making. |
| Approach: | They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations. |
| Outcome: | The proposed corpus provides a valuable resource for research and teaching purposes. |
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)
Copied to clipboard
| Challenge: | Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications. |
| Approach: | They analyze available dialogue corpora and propose possible approaches to create new ones. |
| Outcome: | The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed . |
TextEssence: A Tool for Interactive Analysis of Semantic Shifts Between Corpora (2021.naacl-demos)
Copied to clipboard
| Challenge: | Existing studies using distributional embeddings to study language use have focused on quantitative measurement of change, rather than inter-corpus analysis. |
| Approach: | They propose a system that allows comparative analysis of corpora using embeddings. |
| Outcome: | The proposed system can be used for categorical and comparative analysis of text corpora. |