Papers with English-speaking

8 papers
PreCo: A Large-scale Dataset in Preschool Vocabulary for Coreference Resolution (D18-1)

Copied to clipboard

Challenge: Existing methods for coreference resolution are based on word2vec-like representations of entities.
Approach: They propose a large-scale English dataset for coreference resolution . they use 38K documents and 12.5M words from English-speaking preschoolers .
Outcome: The proposed dataset is more efficient with higher training-test overlap than OntoNotes . the study also shows that mention detection and clustering are more efficient on PreCo .
PEEP-Talk: A Situational Dialogue-based Chatbot for English Education (2023.acl-demo)

Copied to clipboard

Challenge: Existing chatbots lack realistic practice scenarios for English learners . existing platforms employ hand-crafted and patternmatching rules, limiting communication ability and responding appropriately to out-of-situation utterances.
Approach: They propose a real-world situational dialogue-based chatbot for English education . it generates appropriate responses in various real-life situations while providing accurate feedback .
Outcome: The proposed chatbot generates appropriate responses in various real-life situations while providing accurate feedback to learners.
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns (2024.acl-srw)

Copied to clipboard

Challenge: Existing methods for multilingual framing differ from those used in English-speaking world . framers often use loaded vocabularies to create political images or favor a particular point of view .
Approach: They use eight years of Russian-backed disinformation campaigns to examine framing . they find that disinformation campaign consistently favors specific framers .
Outcome: The proposed method underperforms and shows high disagreements in Russian-language articles . the proposed method is based on eight years of Russian-backed disinformation campaigns .
Massively Multi-Lingual Event Understanding: Extraction, Visualization, and Search (2023.acl-demo)

Copied to clipboard

Challenge: Using only English training data, ISI-Clear makes global events available on-demand in 100 languages . Using a fixed task, events may still shift from day to day .
Approach: They propose a cross-lingual zero-shot event extraction system that makes global events available on-demand in 100 languages.
Outcome: The proposed system can extract events from non-English documents in 100 languages.
Testing Low-Resource Language Support in LLMs Using Language Proficiency Exams: the Case of Luxembourgish (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used in research and society at large, but are mostly developed with English-speaking users in mind.
Approach: They investigate the viability of language proficiency exams as evaluation tools for Luxembourgish . large models such as Claude and DeepSeek-R1 typically achieve high scores .
Outcome: The proposed models can predict performance in Luxembourgish language tests.
Exploring the Impact of Language Switching on Personality Traits in LLMs (2025.coling-main)

Copied to clipboard

Challenge: Using three personality tests, we examine the extent to which LLMs align with humans when personality shifts are associated with language changes.
Approach: They propose to use the Eysenck Personality Questionnaire-Revised to examine whether LLMs align with humans when personality shifts are associated with language changes.
Outcome: The results show that language-switching affects personality traits in multilingual individuals, and that it is not translation-related.
Human Interest Framing across Cultures: A Case Study on Climate Change (2025.coling-main)

Copied to clipboard

Challenge: Human Interest (HI) framing is a narrative strategy that injects news stories with a relatable, emotional angle and a human face to engage the audience.
Approach: They perform a systematic analysis of HI stories to understand its role in climate change reporting in English-speaking countries from four continents.
Outcome: The proposed approach has shown to capture and retain readership and enhance political engagement of the population.
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual machine translation (MT) benchmarks are widely used to evaluate the capabilities of modern MT systems.
Approach: They propose to use a multilingual machine translation benchmark to assess the capabilities of modern machine translation systems.
Outcome: The FLORES+ benchmark claims to maintain a translation quality score of over 90% . however, the data in four languages falls short of the 90% quality standard .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations