Challenge: Discourse marker inventories are important tools for the development of discourse parsers and corpora with discourse annotations.
Approach: They explore the potential of multilingual lexical knowledge graphs to induce multilingual discourse marker lexicons using concept propagation methods previously developed in translation inference across dictionaries.
Outcome: The proposed method can induce multilingual discourse marker lexicons using multilingual knowledge graphs.

Similar Papers

ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers (2022.lrec-1)

Copied to clipboard

Challenge: Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemic stance of speaker.
Approach: They propose an ISO-based annotated multilingual parallel corpus for discourse markers . they propose an annotation scheme for discourse relations with a plug-in to ISO 24617-2 .
Outcome: The proposed language resource is based on an ISO-based annotated multilingual parallel corpus of discourse markers.
DiscSense: Automated Semantic Analysis of Discourse Markers (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for predicting discourse markers have been used to study link between markers and semantic relations .
Approach: They use a model trained to predict discourse markers between sentence pairs to predict plausible markers between sentences with a known semantic relation.
Outcome: The proposed method predicts markers between sentence pairs with a known semantic relation . the resulting dataset, named DiscSense, is publicly available .
Mining Discourse Markers for Unsupervised Sentence Representation Learning (N19-1)

Copied to clipboard

Challenge: Current state of the art systems in NLP heavily rely on manually annotated datasets, which are expensive to obtain and are ineffective to extract.
Approach: They propose to automatically discover sentence pairs with relevant discourse markers and apply it to massive amounts of data.
Outcome: The proposed method can learn transferable sentence embeddings from 174 discourse markers even for rare markers such as “coincidentally” or “amazingly”.
Distributed Marker Representation for Ambiguous Discourse Markers and Entangled Relations (2023.acl-long)

Copied to clipboard

Challenge: Discourse markers are natural representations of discourse in our daily language.
Approach: They propose to use unlimited discourse marker data to learn a Distributed Marker Representation by bridging markers with sentence pairs.
Outcome: The proposed model outperforms existing models on the implicit discourse relation recognition task and provides strong interpretability.
Discourse Graph Guided Document Translation with Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Recent agentic machine translation systems mitigate context window constraints but require substantial computational resources and are sensitive to memory retrieval strategies.
Approach: They propose a framework that explicitly models inter-chunk relationships through structured discourse graphs and selectively conditions each translation segment on relevant graph neighbourhoods rather than sequential or exhaustive context.
Outcome: The proposed framework surpasses strong baselines in translation quality and terminology consistency while incurring significantly lower token overhead.
Adapters for Enhanced Modeling of Multilingual Knowledge and Text (2022.findings-emnlp)

Copied to clipboard

Challenge: Large language models learn facts from text corpora, but knowledge graphs contain facts in an explicit triple format, restricting their research and application.
Approach: They propose to enhance multilingual language models with knowledge from multilingual knowledge graphs . they propose to use cross-lingual entity alignment and facts from MLKGs to improve performance .
Outcome: The proposed model improves MLLMs with cross-lingual entity alignment and facts from multilingual knowledge graphs for many languages while maintaining performance on other general language tasks.
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set (2025.acl-long)

Copied to clipboard

Challenge: Existing work on discourse understanding is constrained by framework-dependent discourse representations.
Approach: They examine whether large language models capture discourse knowledge that generalizes across languages and frameworks.
Outcome: The proposed model can generalize discourse information across languages and frameworks.
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)

Copied to clipboard

Challenge: Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications.
Approach: They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages.
Outcome: The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages.
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)

Copied to clipboard

Challenge: lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels.
Approach: They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense.
Outcome: The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations