Papers by Mats Wirén
A Multi-word Expression Dataset for Swedish (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing data on compositionality of multi-word expressions is limited and only available for high resource languages. |
| Approach: | They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression . |
| Outcome: | The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality. |
Identifying Speakers and Addressees in Dialogues Extracted from Literary Fiction (L18-1)
Copied to clipboard
Adam Ek, Mats Wirén, Robert Östling, Kristina N. Björkenstam, Gintarė Grigonytė, Sofia Gustafson Capková
| Challenge: | Using a sequence labeling approach, it is possible to identify speakers and addressees in dialogues extracted from literary fiction using a small amount of training data. |
| Approach: | They propose to use a sequence labeling approach applied to a given set of characters to identify speakers and addressees in dialogues extracted from literary fiction. |
| Outcome: | The proposed method allows for enriched search facilities and construction of social networks from the corpora. |
Evaluation of Really Good Grammatical Error Correction (2024.lrec-main)
Copied to clipboard
| Challenge: | emergence of large language models has highlighted the shortcomings of evaluation methods . evaluators often use grammatical error correction (GEC) to correct language errors at multiple levels . |
| Approach: | They perform a comprehensive evaluation of various GEC systems using Swedish learner texts . they suggest using human post-editing to analyze amount of change required to reach native-level human performance . |
| Outcome: | The proposed evaluations outperform existing methods for grammatical error correction in Swedish . the results highlight the shortcomings of existing evaluation methods . |