Assessing the Quality of an Italian Crowdsourced Idiom Corpus:the Dodiom Experiment (2022.lrec-1)
Copied to clipboard
| Challenge: | a crowdsourcing experiment has been used to collect idiom-related language resources . the data were collected through a game-with-a-purpose . |
| Approach: | They propose to use a game-with-a-purpose to collect idiom-related language resources . they use criteria adopted for the data annotation and evaluation process . |
| Outcome: | The proposed project evaluated idiom-related language resources from a game-with-a-purpose . the results and future work are presented. |
Similar Papers
Work Hard, Play Hard: Collecting Acceptability Annotations through a 3D Game (2022.lrec-1)
Copied to clipboard
| Challenge: | Corpus-based studies on acceptability judgements have always been popular thanks to the release of the CoLA corpus, a large-scale corpus of sentences extracted from linguistic handbooks as examples of acceptable/non acceptable phenomena in English. |
| Approach: | They present a 3D video game that was used to collect acceptability judgments on italian sentences and compare them with experts’ acceptability judgements. |
| Outcome: | The proposed game compares the annotations of Italian sentences with those of experts and shows that they are more reliable than crowd-sourced annotations. |
MAGPIE: A Large Corpus of Potentially Idiomatic Expressions (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora cover less than 5,000 instances of less than 100 different idiom types . large corpus allows for better evaluation of assumptions about idiomatic expressions . |
| Approach: | They propose to build the largest-to-date corpus of idioms for English using crowdsourcing methods. |
| Outcome: | The proposed corpus is larger than existing resources and contains rich metadata and is made publicly available. |
LIdioms: A Multilingual Linked Idioms Data Set (L18-1)
Copied to clipboard
| Challenge: | Recent studies have focused on linguistic data sets that are bilingual on the Linguistic Linked Open Data (LLOD) 1 . |
| Approach: | They describe a multilingual RDF representation of idioms currently containing five languages . they use a model to structure the data and a method to link the data to well-known multilingual data sets such as BabelNet. |
| Outcome: | The proposed model complies with best practices according to Linguistic Linked Open Data Community. |
Towards an Automatic Assessment of Crowdsourced Data for NLU (L18-1)
Copied to clipboard
| Challenge: | Recent development of spoken dialog systems aims at allowing a natural input style. |
| Approach: | They investigate how crowdsourced data can be assessed with respect to its naturalness and usefulness by using a word based language model to identify valid data. |
| Outcome: | The proposed methods show that valid data can be identified with the help of a word based language model. |
Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing (L18-1)
Copied to clipboard
| Challenge: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing . |
| Approach: | They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners. |
| Outcome: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy . |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Designing a Russian Idiom-Annotated Corpus (L18-1)
Copied to clipboard
| Challenge: | a pilot experiment using the idiom-annotated corpus of Russian is described . corpora that could be used for training idiomatic classifiers are scarce, especially if one turns to other languages. |
| Approach: | They describe the development of an idiom-annotated corpus of Russian . the corpus is compiled from freely available online resources . |
| Outcome: | The proposed corpus is based on an online corpus of Russian texts . it is available for research purposes and can be used for linguistic studies and pedagogy . |
Improving Crowdsourcing-Based Annotation of Japanese Discourse Relations (L18-1)
Copied to clipboard
| Challenge: | Discourse parsing is an important task in natural language processing, but few languages have corpora annotated with discourse relations . crowdsourcing-based annotations are of poor quality and require expensive and time-consuming . et al. (2009) evaluated the quality of annotations using expert annotations. |
| Approach: | They construct a Japanese corpus with discourse annotations through crowdsourcing . they propose improvement techniques based on language tests . |
| Outcome: | The proposed methods improve the quality of the annotations, and will make them publicly available. |
AStitchInLanguageModels: Dataset and Methods for the Exploration of Idiomaticity in Pre-Trained Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets are limited to providing the degree of idiomaticity of expressions along with the literal and, where applicable, (a single) non-literal interpretation of MWEs. |
| Approach: | They propose to use a dataset to test the effectiveness of a language model in generating representations of sentences containing idioms. |
| Outcome: | The proposed model performs reasonably well on the one-shot and few-shot scenarios, but there is scope for improvement in the zero-shot scenario. |
Crowdsourcing in the Development of a Multilingual FrameNet: A Case Study of Korean FrameNet (2020.lrec-1)
Copied to clipboard
| Challenge: | Using current methods, the construction of multilingual FrameNets is expensive and complex. |
| Approach: | They evaluated whether crowdsourcing approaches captured cross-cultural and cross-linguistic meanings . they found that crowd workers made intuitive choices comparable to trained FrameNet experts . |
| Outcome: | The results are now available in Korean FrameNet 1.1. |