Papers by Simon Hengchen
Superlim: A Swedish Language Understanding Evaluation Benchmark (2023.emnlp-main)
Copied to clipboard
Aleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi
| Challenge: | In this paper, we present a multi-task benchmark for Swedish language models . we address methodological challenges, such as mitigating the Anglocentric bias when creating datasets for a less-resourced language . |
| Approach: | They propose a multi-task NLP benchmark for Swedish language models . they propose to use superlim to evaluate Swedish language model performance . |
| Outcome: | The proposed benchmark does not approach ceiling performance on any of the tasks, suggesting it is difficult to implement. |
Time-Out: Temporal Referencing for Robust Modeling of Lexical Semantic Change (P19-1)
Copied to clipboard
| Challenge: | State-of-the-art lexical semantic change detection models suffer from noise stemming from vector space alignment. |
| Approach: | They propose a method to simulate lexical semantic change and control for possible biases by avoiding alignment. |
| Outcome: | The proposed method outperforms state-of-the-art models on a synthetic task and a manual testset. |
DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for graded contextual word meaning annotation have not been implemented yet. |
| Approach: | They propose a multi-round incremental annotation process and a clustering algorithm to group usages into senses to create a large-scale dataset. |
| Outcome: | The proposed method is the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments. |
Dataset for Temporal Analysis of English-French Cognates (2020.lrec-1)
Copied to clipboard
| Challenge: | Using computational techniques to study language evolution has gained much attention . comparing two or more languages can shed light on how they co-evolve . |
| Approach: | They propose to use a dataset to investigate the similarity in evolution between languages by comparing cognates across time. |
| Outcome: | The proposed dataset is the first to use computational approaches and large data to make a cross-language diachronic analysis. |