Papers by Simon Hengchen

4 papers
Superlim: A Swedish Language Understanding Evaluation Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we present a multi-task benchmark for Swedish language models . we address methodological challenges, such as mitigating the Anglocentric bias when creating datasets for a less-resourced language .
Approach: They propose a multi-task NLP benchmark for Swedish language models . they propose to use superlim to evaluate Swedish language model performance .
Outcome: The proposed benchmark does not approach ceiling performance on any of the tasks, suggesting it is difficult to implement.
Time-Out: Temporal Referencing for Robust Modeling of Lexical Semantic Change (P19-1)

Copied to clipboard

Challenge: State-of-the-art lexical semantic change detection models suffer from noise stemming from vector space alignment.
Approach: They propose a method to simulate lexical semantic change and control for possible biases by avoiding alignment.
Outcome: The proposed method outperforms state-of-the-art models on a synthetic task and a manual testset.
DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for graded contextual word meaning annotation have not been implemented yet.
Approach: They propose a multi-round incremental annotation process and a clustering algorithm to group usages into senses to create a large-scale dataset.
Outcome: The proposed method is the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments.
Dataset for Temporal Analysis of English-French Cognates (2020.lrec-1)

Copied to clipboard

Challenge: Using computational techniques to study language evolution has gained much attention . comparing two or more languages can shed light on how they co-evolve .
Approach: They propose to use a dataset to investigate the similarity in evolution between languages by comparing cognates across time.
Outcome: The proposed dataset is the first to use computational approaches and large data to make a cross-language diachronic analysis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations