Challenge: In this paper, we examine the ability of large language models (LLMs) to identify different meanings in sentences that are superficially similar.
Approach: They propose a challenge dataset for NLP with large lexical overlap which minimises the possibility of models discerning entailment solely based on token distinctions.
Outcome: The proposed model fails to distinguish between constructions with three classes of adjectives which cannot be distinguished by surface features.

Similar Papers

Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones? (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities, but still suffer from inconsistency issues.
Approach: They develop a ConsisEval benchmark to evaluate LLMs' inconsistency . they find that LLM models can paradoxically fail at easier problems .
Outcome: The proposed model achieves highest consistency score but inconsistent to specific questions due to distraction by redundant information, misinterpretation of questions, etc.
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations (N18-5)

Copied to clipboard

Challenge: iii: 39 outstanding papers accepted for presentation at NAACL HLT 2018 in New Orleans, Louisiana . iv: 20 outstanding papers will be displayed during the conference .
Approach: iii: 39 outstanding papers accepted for presentation at NAACL HLT 2018 in new orleans, la . iv: demonstrations session is an opportunity for researchers and developers to present their systems .
Outcome: the demonstrations session at NAACL HLT 2018 in New Orleans, Louisiana, USA, received 39 outstanding papers .
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations (D19-3)

Copied to clipboard

Challenge: Proceedings of the system demonstrations session were presented at the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) EMNMP-IjCNLP 2019 has a Best Demo Award for the first time .
Approach: Proceedings of the system demonstrations session are available online . they were presented at the conference on empirical methods in natural language processing .
Outcome: The system demonstrations session received 110 submissions, 22 of which were either invalid or withdrawn by the authors.
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners (2024.emnlp-main)

Copied to clipboard

Challenge: a new hypothesis-testing framework is developed to assess whether large language models possess genuine reasoning abilities or primarily depend on token bias.
Approach: They propose a framework to assess whether large language models have genuine reasoning abilities or primarily depend on token bias.
Outcome: The proposed framework outlines a list of hypotheses where token biases are readily identifiable . the results suggest that most LLMs still struggle with logical reasoning .
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2023.acl-demo)

Copied to clipboard

Challenge: 58 papers were selected for inclusion in the program, while a small number received only two reviews.
Approach: the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023) will be held in london from July 9-14, 2023 . 58 submissions were selected for inclusion in the program, with an acceptance rate of 37%)
Outcome: the system demonstration track received a record number of submissions . 58 papers were selected for inclusion in the program .
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2024.acl-demos)

Copied to clipboard

Challenge: ACL 2024 System Demonstration Track invites submissions describing system demonstrations . submissions will undergo a single-blind review process .
Approach: the ACL 2024 System Demonstration Track invites submissions . papers will be published in a companion volume of the conference proceedings . submissions will undergo a single-blind review process .
Outcome: the Demonstration Track at ACL 2024 is a venue for papers describing system demonstrations . publicly available open-source or open-access systems are of special interest . submissions will undergo a single-blind review process .
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2025.acl-demo)

Copied to clipboard

Challenge: ACL 2025 System Demonstration Track accepted 64 papers based on reviews . short-listed 7 papers for Best System Demo award .
Approach: the ACL 2025 System Demonstration Track is a conference for papers describing system demonstrations . the track received a record 187 submissions, of which 178 papers were valid with required materials .
Outcome: the ACL 2025 System Demonstration Track received 187 submissions . 178 papers were valid with required materials .
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2026.acl-demo)

Copied to clipboard

Challenge: ACL 2026 System Demonstration Track accepted 85 papers . one paper received Best Demo award .
Approach: the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) took place from July 2-7, 2026 in San Diego, California.
Outcome: the ACL 2026 System Demonstration Track accepted 85 papers based on the submitted reviews . one paper received the best demo award: The olmOCR Project: Building Fully Open OCR using VLMs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations