Proceedings of the 31st International Conference on Computational Linguistics: System Demonstrations

22 papers
PolyMinder: A Support System for Entity Annotation and Relation Extraction in Polymer Science Documents (2025.coling-demos)

Copied to clipboard

Challenge: Automated Named Entity Recognition (NER) and Relation Extraction (RE) models are tailored to the polymer domain.
Approach: They propose to automate the annotation process by providing a web-based interface where users can visualize, verify, and refine the extracted information before finalizing the annotations.
Outcome: The proposed system streamlines the annotation process by providing a web-based interface where users can visualize, verify, and refine the extracted information before finalizing the annotations.
Streamlining Biomedical Research with Specialized LLMs (2025.coling-demos)

Copied to clipboard

Challenge: Using large language models, we can generate accurate, context-aware responses with minimal prompts.
Approach: They propose a system that integrates domain-specific large language models with advanced information retrieval techniques to deliver comprehensive and context-aware responses.
Outcome: The proposed system improves quality of dialogue generation and improves efficiency in the biomedical and pharmaceutical domains.
LENS: Learning Entities from Narratives of Skin Cancer (2025.coling-demos)

Copied to clipboard

Challenge: Learning entities from narratives of skin cancer (LENS) is an automatic entity recognition system built on colloquial writings from skin cancer-related forums.
Approach: They propose to use reddit forums to create an automatic entity recognition system that can be used to predict skin cancer outcomes.
Outcome: LENS achieves an overall entity-level F1 score of 0.561 . other notable results include “CANC_T” (0.747), “STG” (0.888), “POB” (0.914), “GENDER” (0.750), “A/G” (00.646), “EMO” (0.619), and “MHD” (0.503).
Loki: An Open-Source Tool for Fact Verification (2025.coling-demos)

Copied to clipboard

Challenge: Loki is an open-source fact-checking tool designed to address the growing problem of misinformation.
Approach: They propose a tool that breaks down the fact-checking task into five steps . they propose LOKI, which offers a semiautomated, human-in-the-loop approach .
Outcome: a new open-source tool is designed to address the growing problem of misinformation . the tool breaks down the fact-checking task into five steps to assist human judgment .
UnifiedGEC: Integrating Grammatical Error Correction Approaches for Multi-languages with a Unified Framework (2025.coling-demos)

Copied to clipboard

Challenge: Existing tools for GEC have been developed to support research on grammatical errors, but there is no comprehensive evaluation on these models.
Approach: They propose an open-source framework for Grammatical Error Correction that integrates 5 widely-used GEC models and compares their performance on 7 datasets in different languages.
Outcome: The proposed framework compares 5 widely-used models on 7 datasets in different languages.
Reliable, Reproducible, and Really Fast Leaderboards with Evalica (2025.coling-demos)

Copied to clipboard

Challenge: Using open-source evaluation tools, we create reliable and reproducible model leaderboards with human and machine feedback.
Approach: They propose an open-source evaluation toolkit that facilitates the creation of reliable and reproducible model leaderboards.
Outcome: The evaluation tool facilitates the creation of reliable and reproducible model leaderboards.
BeefBot: Harnessing Advanced LLM and RAG Techniques for Providing Scientific and Technology Solutions to Beef Producers (2025.coling-demos)

Copied to clipboard

Challenge: Generic Large Language Models (LLMs) are useful for information retrieval but often hallucinate and fail to deliver tailored solutions to the specific needs of beef producers.
Approach: They propose to use Retrieval-Augmented Generation and fine-tuning to build a chatbot for beef producers that retrieves latest agricultural technologies and scientific insights.
Outcome: The proposed chatbot retrieves latest agricultural technologies, practices and scientific insights to provide rapid, domain-specific advice.
AI-Press: A Multi-Agent News Generating and Feedback Simulation System Powered by Large Language Models (2025.coling-demos)

Copied to clipboard

Challenge: We introduce AI-Press, an automated news drafting and polishing system based on multi-agent collaboration and Retrieval-Augmented Generation.
Approach: They introduce AI-Press, an automated news drafting and polishing system based on multi-agent collaboration and Retrieval-Augmented Generation.
Outcome: The proposed system generates public responses considering demographic distributions.
A Probabilistic Toolkit for Multi-grained Word Segmentation in Chinese (2025.coling-demos)

Copied to clipboard

Challenge: Existing tools for word segmentation are based on different linguistic theories or target different scenarios.
Approach: They propose a probabilistic toolkit for multi-grained word segmentation in Chinese . they adopt semi-Markov CRF for single-grain word segmenting (SWS) .
Outcome: The proposed approach can produce marginal probabilities of words during inference and significantly improve performance in the cross-domain scenario.
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs (2025.coling-demos)

Copied to clipboard

Challenge: Existing open-source evaluation models lack a user-friendly visualization tool and are not optimized for accelerated model inference.
Approach: They propose to use open-source evaluation models to evaluate language model responses.
Outcome: The proposed model is lightweight, precise, efficient, and user-friendly, with an intuitive visualization interface for ease of deployment and use.
LUCE: A Dynamic Framework and Interactive Dashboard for Opinionated Text Analysis (2025.coling-demos)

Copied to clipboard

Challenge: LUCE is an advanced dynamic framework for analysing opinionated text . it features computational modules for different elements of opinions, e.g., sentiment/emotion, suggestion, figurative language, hate/toxic speech, and topics.
Approach: They introduce a dynamic framework with an interactive dashboard for analysing opinionated text . it features computational modules of text classification and extraction for different elements of opinions .
Outcome: The framework is validated in a relevant environment and its capabilities and performance demonstrated . it features trained models, python-based APIs, and a user-friendly dashboard .
RAGthoven: A Configurable Toolkit for RAG-enabled LLM Experimentation (2025.coling-demos)

Copied to clipboard

Challenge: Large Language Models (LLMs) have significantly altered the landscape of Natural Language Processing (NLP), but their use as a baseline method has not been extensive.
Approach: They propose a tool for automatic evaluation of RAG-based pipelines that provides a simple yet powerful abstraction.
Outcome: The proposed tool provides an automatic evaluation of RAG-based pipelines.
MuRAR: A Simple and Effective Multimodal Retrieval and Answer Refinement Framework for Multimodal Question Answering (2025.coling-demos)

Copied to clipboard

Challenge: Recent advances in retrieval-augmented generation have demonstrated impressive performance on the question-answering task.
Approach: They propose a retrieval-augmented generation framework that generates an initial text answer and retrieves multimodal data relevant to the snippets of the initial text.
Outcome: The proposed framework can be easily integrated into an enterprise chatbot to produce multimodal answers with minimal modifications.
Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference (2025.coling-demos)

Copied to clipboard

Challenge: Qualitative interviews are foundational to social science research, offering deep insights through open-ended conversations.
Approach: They introduce a human-like embodied AI interviewer which integrates android and humanoid robots equipped with advanced conversational capabilities.
Outcome: The proposed system performs well in a real-world case study at SIGDIAL 2024 with 42 participants, of whom 69% reported positive experiences.
CASE: Large Scale Topic Exploitation for Decision Support Systems (2025.coling-demos)

Copied to clipboard

Challenge: Topic models are still a major tool for information retrieval and summarization, but their integration into decision-making systems is limited.
Approach: They propose a tool for exploiting topic information for semantic analysis of large corpora using a Solr engine and a customized indexing strategy.
Outcome: The proposed approach can be used to analyze large corpora and perform thematic trend analysis, topic-based document retrieval, or similarity search.
GECTurk WEB: An Explainable Online Platform for Turkish Grammatical Error Detection and Correction (2025.coling-demos)

Copied to clipboard

Challenge: GECTurk WEB is an open-source, web-based system that can detect and correct most common forms of Turkish writing errors.
Approach: They propose a web-based system that detects and corrects most common errors in Turkish . it provides an easy-to-use tool for native speakers and second language learners .
Outcome: The proposed system achieves 88,3 system usability score and is shown to help learn/remember a grammatical rule.
GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek (2025.coling-demos)

Copied to clipboard

Challenge: GR-NLP-TOOLKIT is an open-source natural language processing toolkit for modern Greek.
Approach: They present GR-NLP-TOOLKIT, an open-source natural language processing toolkit for Greek.
Outcome: The toolkit provides state-of-the-art performance in five core NLP tasks . it can be easily installed in Python and is accessible through a demonstration platform on HuggingFace .
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization (2025.coling-demos)

Copied to clipboard

Challenge: ViSoLex is an open-source repository for Vietnamese lexical normalization . it provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization enabling users to retrieve standard forms of informal language and standardize text containing NSWs.
Approach: They propose to integrate pre-trained language models and weakly supervised learning techniques to ensure accurate and efficient normalization.
Outcome: The system provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization, enabling users to retrieve standard forms of informal language and standardize text containing NSWs.
CompUGE-Bench: Comparative Understanding and Generation Evaluation Benchmark for Comparative Question Answering (2025.coling-demos)

Copied to clipboard

Challenge: Comparative Question Answering systems help users make informed decisions by generating comparative information.
Approach: They propose a comprehensive benchmark designed to evaluate Comparative Question Answering systems.
Outcome: The proposed benchmark is available on HuggingFace Spaces . it unifies multiple datasets and provides a robust evaluation platform .
Autonomous Machine Learning-Based Peer Reviewer Selection System (2025.coling-demos)

Copied to clipboard

Challenge: Existing systems that match papers with experts are inefficient and often require long turnaround times.
Approach: They propose an autonomous peer reviewer selection system that employs the natural language processing model to match submitted papers with expert reviewers independently of traditional journals and conferences.
Outcome: The proposed system performs faster and smaller than current models while being more scalable.
CULTURALLY YOURS: A Reading Assistant for Cross-Cultural Content (2025.coling-demos)

Copied to clipboard

Challenge: Culturally Yours (CY) is a cultural reading assistant that helps users from diverse cultural backgrounds understand content from online sources that are written by people from a different culture.
Approach: They propose to use culturally sensitive language to personalize a cultural reading assistant tool that can identify cultural-specific items for users from varying cultural contexts.
Outcome: The tool personalizes to the user’s preferences based on the interaction of the user with the tool.
FEAT-writing: An Interactive Training System for Argumentative Writing (2025.coling-demos)

Copied to clipboard

Challenge: Argumentative writing is a critical skill for academic success, but many students struggle to develop these skills.
Approach: They developed an online system that provides students with automated feedback and exercises for argumentative writing.
Outcome: The proposed system improves argumentative writing quality among native English speakers and english-as-a-foreign-language learners.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations