Papers by Murathan Kurfalı

4 papers
A Multi-word Expression Dataset for Swedish (2020.lrec-1)

Copied to clipboard

Challenge: Existing data on compositionality of multi-word expressions is limited and only available for high resource languages.
Approach: They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression .
Outcome: The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality.
An Assessment of Explicit Inter- and Intra-sentential Discourse Connectives in Turkish Discourse Bank (L18-1)

Copied to clipboard

Challenge: Discourse parsing is a challenging task for NLP.
Approach: They propose to add a new set of explicit intra-sentential connectives to Turkish Discourse Bank 1.1 . they propose to evaluate the converb sense annotations and compare them to other Turkish corpus .
Outcome: The proposed annotations show that the subordinators tend to select certain senses not selected by explicit inter- and intra-sentential discourse connectives in the data.
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)

Copied to clipboard

Challenge: Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications.
Approach: They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages.
Outcome: The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages.
Evaluation of Really Good Grammatical Error Correction (2024.lrec-main)

Copied to clipboard

Challenge: emergence of large language models has highlighted the shortcomings of evaluation methods . evaluators often use grammatical error correction (GEC) to correct language errors at multiple levels .
Approach: They perform a comprehensive evaluation of various GEC systems using Swedish learner texts . they suggest using human post-editing to analyze amount of change required to reach native-level human performance .
Outcome: The proposed evaluations outperform existing methods for grammatical error correction in Swedish . the results highlight the shortcomings of existing evaluation methods .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations