Identifying Emerging Concepts in Large Corpora (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for text analysis are not specifically designed for identifying emergent concepts, instead applying general-purpose techniques that do not account for distinct temporal patterns associated with conceptual emergence.
Approach: They propose a method to identify emerging concepts in large text corpora by analyzing changes in the heatmaps of the underlying embedding space.
Outcome: The proposed method outperforms existing methods by analyzing speeches in the U.S. Senate from 1941 to 2015.

Similar Papers

Findings of the Association for Computational Linguistics: EMNLP 2020 (2020.findings-emnlp)

Copied to clipboard

Challenge: . - (EN)
Approach: . - (EN)
Outcome: . - (EN)
On the Distribution of Deep Clausal Embeddings: A Large Cross-linguistic Study (P19-1)

Copied to clipboard

Challenge: Empirical evidence on the prevalence and limits of embeddings has been based on either laboratory setups or corpus data of relatively limited size.
Approach: They use large, dependency-parsed corpora to capture clausal embedding through dependency graphs and assess their distribution.
Outcome: The results show that there is no evidence for hard constraints on embedding depth . they also show that sentences with many embeddable clauses do not display a bias towards less deep embedded sentences.
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications.
Approach: They analyze available dialogue corpora and propose possible approaches to create new ones.
Outcome: The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed .
Findings of the Association for Computational Linguistics: EMNLP 2021 (2021.findings-emnlp)

Copied to clipboard

Challenge: . - (EN)
Approach: . - (EN)
Outcome: . - (EN)
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
An Analysis of Negation in Natural Language Understanding Corpora (2022.acl-short)

Copied to clipboard

Challenge: Using annotator-generated examples, one can evaluate systems with synthetic language that is not representative of language in the wild.
Approach: They analyze negation in eight popular corpora spanning six natural language understanding tasks.
Outcome: The proposed corpora have few negations compared to general-purpose English and are often unimportant . state-of-the-art transformers obtain significantly worse results with instances that contain negation, especially if the negations are important.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.
Findings of the Association for Computational Linguistics: EMNLP 2025 (2025.findings-emnlp)

Copied to clipboard

Challenge: null
Approach: null
Outcome: null
Findings of the Association for Computational Linguistics: EMNLP 2022 (2022.findings-emnlp)

Copied to clipboard

Challenge: null
Approach: null
Outcome: null
HistLens: Mapping Idea Change across Concepts and Corpora (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to diachronic semantics and discourse analysis focus on a single concept or corpus, argues a new paper.
Approach: They propose a framework for multi-concept, multi-corpus conceptual-history analysis that decomposes concept representations into interpretable features and tracks activation dynamics over time and across sources.
Outcome: The proposed framework decomposes concept representations into interpretable features and tracks their activation dynamics over time and across sources.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations