Papers by Stephan Oepen

7 papers
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a new massive multilingual dataset is available for language modeling and machine translation training.
Approach: They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora .
Outcome: The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 .
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) (2025.acl-long)

Copied to clipboard

Challenge: a large number of textual data is needed to train state-of-the-art large language models.
Approach: They propose a collection of monolingual and parallel corpora from the Internet Archive . they document the entire data pipeline and release the code to reproduce it .
Outcome: The proposed collection of monolingual and parallel corpora is based on the HPLT v2 dataset . it includes 8T tokens covering 193 languages and 380M sentence pairs covering 51 languages .
Direct parsing to sentiment graphs (2022.acl-short)

Copied to clipboard

Challenge: Existing methods for structured sentiment analysis (SSA) focus on subcomponents of sentiment graphs without explicitly expressing their relations or the polarity.
Approach: They propose a graph-based semantic parser which directly predicts sentiment graphs from text without reliance on lossy conversions to intermediate dependency representations.
Outcome: The proposed model performs on 4 out of 5 standard benchmark sets and compares with dependency-based models on the more structurally complex datasets.
Graph-Based Meaning Representations: Design and Processing (P19-4)

Copied to clipboard

Challenge: This tutorial focuses on representing and processing sentence meaning in the form of labeled directed graphs.
Approach: This tutorial will briefly review relevant background in formal and linguistic semantics . it will also briefly define a unified abstract view on different flavors of semantic graphs - and associated terminology .
Outcome: The tutorial will briefly review relevant background in formal and linguistic semantics .
A Tale of Three Parsers: Towards Diagnostic Evaluation for Meaning Representation Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Empirical results suggest that the proposed methodology can be meaningfully applied to parsing into graph-structured target representations, uncovering hitherto unknown properties of the different approaches.
Approach: They propose to map from natural language utterances to graph-based encodings of its semantic structure using contrastive and diagnostic evaluation techniques.
Outcome: The proposed method can be meaningfully applied to parsing into graph-structured target representations, uncovering hitherto unknown properties of the different systems that can inform future development and cross-fertilization across approaches.
Structured Sentiment Analysis as Dependency Graph Parsing (2021.acl-long)

Copied to clipboard

Challenge: Structured sentiment analysis attempts to extract full opinion tuples from a text, but has been subdivided into smaller and smaller sub-tasks, e.g., target extraction or targeted polarity classification.
Approach: They propose a framework which jointly predicts all elements of an opinion tuple and their relations by using dependency graph parsing.
Outcome: The proposed framework improves on five datasets in English, Norwegian, Basque, and Catalan and refining the sentiment graphs with syntactic dependency information further improves results.
Transfer and Multi-Task Learning for Noun–Noun Compound Interpretation (D18-1)

Copied to clipboard

Challenge: In computational linguistics, nounnoun compound interpretation is approached as an automatic classification problem.
Approach: They empirically evaluate the utility of transfer and multi-task learning on a challenging semantic classification task.
Outcome: The proposed methods improve the accuracy of a neural classifier and its F1 scores on the less frequent, but more difficult relations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations