When Beards Start Shaving Men: A Subject-object Resolution Test Suite for Morpho-syntactic and Semantic Model Introspection (2020.coling-main)
Copied to clipboard
| Challenge: | In German, subject-object resolution is the second most error-prone task in syntactic parsing, after prepositional phrase attachment. |
| Approach: | They introduce the SORTS Subject-Object Resolution Test Suite of German minimal sentence pairs for model introspection. |
| Outcome: | The proposed test suite is based on the annotations of 8,502 transitive clauses with 8 word order patterns, 5 morphological and syntactic and 11 semantic property classes. |
Similar Papers
COMPS: Conceptual Minimal Pair Sentences for testing Robust Property Knowledge and its Inheritance in Pre-trained Language Models (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing pre-trained language models (PLMs) lack robustness in demonstrating simple reasoning, despite having the prerequisite knowledge. |
| Approach: | They propose to test pre-trained language models' ability to attribute properties to concepts and their ability to demonstrate property inheritance behavior. |
| Outcome: | The proposed model can easily distinguish between concepts on the basis of a property when they are trivially different, but find it relatively difficult when concepts are related on the base of nuanced knowledge representations. |
jiant: A Software Toolkit for Research on General-Purpose Text Understanding Models (2020.acl-demos)
Copied to clipboard
Yada Pruksachatkun, Phil Yeres, Haokun Liu, Jason Phang, Phu Mon Htut, Alex Wang, Ian Tenney, Samuel R. Bowman
| Challenge: | jiant is an open source toolkit for conducting multitask and transfer learning experiments on English NLU tasks. |
| Approach: | They introduce jiant, an open source toolkit for conducting multitask and transfer learning experiments on English NLU tasks. |
| Outcome: | The proposed toolkit reproduces published performance on GLUE and SuperGLUE tasks. |
What BERT Is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models (2020.tacl-1)
Copied to clipboard
| Challenge: | Pretraining by language modeling has become popular but we have yet to understand what language models learn during that process. |
| Approach: | They propose diagnostics that ask questions about information used by language models for generating predictions in context. |
| Outcome: | The proposed diagnostics can be used to study the popular BERT model . they show that the model can distinguish good from bad completions, but struggles with inference and role-based event prediction. |
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) can explain quantum mechanics and write sophisticated code, yet fail to parse sentences like "The cat that the mouse feared chased meowed" |
| Approach: | They propose a framework to distinguish structural understanding from semantic pattern matching . they use a set of 9,720 comprehension questions on center-embedded sentences . |
| Outcome: | a new framework shows that models lose performance when they abandon structural analysis for semantic associations. |
What’s Wrong with Hebrew NLP? And How to Make it Right (D19-3)
Copied to clipboard
| Challenge: | Sub-optimal performance of many morphologically rich languages (MRLs) is due to errors in early morphology disambiguation decisions, that cannot be recovered later on in the pipeline, yielding incoherent annotations on the whole. |
| Approach: | They propose to use a joint morpho-syntactic infrastructure for processing Modern Hebrew texts to provide rich and expressive annotations. |
| Outcome: | The proposed pipelines are based on a morpho-syntactic infrastructure for processing Modern Hebrew texts. |
Dependency resolution at the syntax-semantics interface: psycholinguistic and computational insights on control dependencies (2023.acl-long)
Copied to clipboard
| Challenge: | Using psycholinguistic and computational experiments, we compare the ability of humans and several pre-trained masked language models to correctly identify control dependencies in Spanish sentences. |
| Approach: | They compare the ability of humans and several pre-trained masked language models to correctly identify control dependencies in Spanish sentences such as ‘José le prometió/ordenó a Mara ser ordenado/a’. |
| Outcome: | The models fail to identify the correct antecedent in non-adjacent dependencies, showing their reliance on linearity. |
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (D18-2)
Copied to clipboard
| Challenge: | 77 submissions were received for the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) 4 of the 73 valid submissions received were either invalid or withdrawn by the authors. |
| Approach: | The volume contains papers from the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) 4 of the 77 submissions were either invalid or withdrawn by the authors. |
| Outcome: | The system demonstrations session included papers from the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) 4 of the 73 valid submissions were either invalid or withdrawn by the authors. |
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations (D19-3)
Copied to clipboard
| Challenge: | Proceedings of the system demonstrations session were presented at the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) EMNMP-IjCNLP 2019 has a Best Demo Award for the first time . |
| Approach: | Proceedings of the system demonstrations session are available online . they were presented at the conference on empirical methods in natural language processing . |
| Outcome: | The system demonstrations session received 110 submissions, 22 of which were either invalid or withdrawn by the authors. |
GerEO: A Large-Scale Resource on the Syntactic Distribution of German Experiencer-Object Verbs (2022.lrec-1)
Copied to clipboard
| Challenge: | Psych verbs and their properties in multiple languages have ignited discussions among linguists for several decades . Psych-verbs are often considered syntactically deviant, although this has occasionally been called into question . |
| Approach: | They propose to use a large-scale database of more than 10,000 examples for 64 verbs from a newspaper corpus annotated for several syntactic and semantic features relevant for their analysis. |
| Outcome: | The proposed database contains 10,000 examples for 64 verbs from a newspaper corpus and includes syntactic construction, semantic stimulus type, and form of possible stimulus preposition. |
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts (2025.emnlp-tutorials)
Copied to clipboard
| Challenge: | EMNLP 2025 tutorials will cover seven cutting-edge topics . the process of soliciting, reviewing and selecting tutorials was a collaborative effort . |
| Approach: | EMNLP 2025 will feature tutorials on seven cutting-edge topics . the process of soliciting, reviewing and selecting tutorials was a collaborative effort . |
| Outcome: | EMNLP 2025 will feature tutorials on seven cutting-edge topics . the process of soliciting, reviewing and selecting tutorials was a collaborative effort . |