Humans Meet Models on Object Naming: A New Dataset and Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Existing object naming datasets that use only images with a bounding box are noisy . a human-like model behavior is not stable across domains, a study finds . |
| Approach: | They use MN v2 to verify object naming datasets with dozens of valid names per object . they find that human-like model behavior is not stable across domains . |
| Outcome: | The proposed model confuses people and clothing objects more frequently than humans do. |
Similar Papers
Object Naming in Language and Vision: A Survey and a New Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Object naming has been studied in Psycholinguistics, but has received little attention in Computational Linguistics. |
| Approach: | They propose a dataset that provides 36 name annotations for each of 25K objects in images selected from VisualGenome. |
| Outcome: | The proposed dataset shows that people choose certain names for objects, on average. |
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs (2024.acl-short)
Copied to clipboard
| Challenge: | Recent work has highlighted that speakers display a wide range of variability when asked to utter sentences, resulting in inter-speaker variability but also variability over time for the same speaker. |
| Approach: | They evaluate Vision & Language Large Language Models (VLLMs) on three categories where humans show great subjective variability concerning the distribution over plausible labels. |
| Outcome: | The proposed models can mimic human distributions over plausible labels, but fail to assign quantifiers, a task that requires more accurate, high-level reasoning. |
MultiCoNER v2: a Large Multilingual dataset for Fine-grained and Noisy Named Entity Recognition (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a core task in Natural Language Processing. |
| Approach: | They present a dataset for fine-grained Named Entity Recognition covering 33 entity classes across 12 languages in monolingual and multilingual settings. |
| Outcome: | The proposed dataset covers 33 entity classes across 12 languages in monolingual and multilingual settings. |
The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation (2022.emnlp-main)
Copied to clipboard
| Challenge: | a paper argues that human label variation impacts all stages of the ML pipeline . human label variations are often considered noise due to disagreement, subjectivity in annotation or multiple plausible answers. |
| Approach: | They propose to reconcile different notions of human label variation and propose a repository of publicly-available datasets with un-aggregated labels. |
| Outcome: | The proposed approaches are compared with publicly available datasets with un-aggregated labels and identify gaps. |
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)
Copied to clipboard
| Challenge: | Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process. |
| Approach: | They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions. |
| Outcome: | The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input. |
Why do objects have many names? A study on word informativeness in language use and lexical systems (2024.emnlp-main)
Copied to clipboard
| Challenge: | lexical systems contain many different words that can be assigned to the same object . studies on language use have explored how speakers adapt their referring expressions to communicate in context without tackling in-context communication. |
| Approach: | They propose a simple measure of informativeness for words and lexical systems, grounded in a visual space, and analyze color naming data for English and Mandarin Chinese. |
| Outcome: | The proposed system allows for a soft mapping between referents and words, taking into account both in-context communication and the structure of the lexical system. |
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them . |
| Approach: | They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition . |
| Outcome: | The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs . |
Synonym relations affect object detection learned on vision-language data (2024.findings-naacl)
Copied to clipboard
| Challenge: | a recent study shows that vision-language models that accept textual input are not robust to variations in how input is provided. |
| Approach: | They propose two approaches to improve vision-language object detectors' performance . they use back-translation and class embedding enrichment to improve their models . |
| Outcome: | The proposed approaches improve performance on synonyms from mAP@0.3=33.87% to 37.93%. |
CoNLL#: Fine-grained Error Analysis and a Corrected Test Set for CoNLL-03 English (2024.lrec-main)
Copied to clipboard
| Challenge: | a glass ceiling for named entity recognition systems has been suggested for 2021 . however, the performance of the most popular NER benchmarks has plateaued since then . we investigate what NER models are still struggling with . |
| Approach: | They perform a fine-grained evaluation of the model outputs by adding document annotations to the CoNLL-03 English dataset to identify lingering errors. |
| Outcome: | The proposed model is able to correct errors and guide future work. |
Characterizing Human and Zero-Shot GPT-3.5 Object-Similarity Judgments (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent advances in large language models have yielded few-shot, human-comparable performance on a range of tasks, but studies of LLM annotation accuracy and behavior are sparse. |
| Approach: | They characterize OpenAI’s GPT-3.5’s judgment on a behavioral task for implicit object categorization and give similarities and differences between them. |
| Outcome: | The proposed model augments human responses with LLMs for domains where data is sparse or compute resources are low. |