Papers with easy-to-use
CogKTR: A Knowledge-Enhanced Text Representation Toolkit for Natural Language Understanding (2022.emnlp-demos)
Copied to clipboard
Zhuoran Jin, Tianyi Men, Hongbang Yuan, Yuyang Zhou, Pengfei Cao, Yubo Chen, Zhipeng Xue, Kang Liu, Jun Zhao
| Challenge: | Existing knowledge-enhanced methods are limited to knowledge-intensive tasks. |
| Approach: | They propose a knowledge-enhanced text representation toolkit for natural language understanding . it combines knowledge acquisition, knowledge representation, knowledge injection and knowledge application . |
| Outcome: | The proposed toolkit supports knowledge acquisition, knowledge representation, knowledge injection, and knowledge application. |
F-coref: Fast, Accurate and Easy to Use Coreference Resolution (2022.aacl-demo)
Copied to clipboard
| Challenge: | Existing models for coreference resolution are difficult to implement, consume a lot of GPU memory and take long to process each document. |
| Approach: | They propose a python package for fast, accurate, and easy-to-use English coreference resolution. |
| Outcome: | The proposed model can process 2.8K OntoNotes documents in 25 seconds on a V100 GPU, compared to 6 minutes for the LingMess model and 12 minutes of the popular AllenNLP coreference model. |
Text Characterization Toolkit (TCT) (2022.aacl-demo)
Copied to clipboard
Daniel Simig, Tianlu Wang, Verna Dankers, Peter Henderson, Khuyagbaatar Batsuren, Dieuwke Hupkes, Mona Diab
| Challenge: | Text Characterization Toolkit (TCT) is a tool that researchers can use to study characteristics of large datasets. |
| Approach: | They propose a text characterization toolkit that researchers can use to study characteristics of large datasets. |
| Outcome: | The proposed tool can be used to study characteristics of large datasets and to understand the influence of attributes on models’ behaviour. |
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors (2025.emnlp-main)
Copied to clipboard
| Challenge: | Evaluating the pedagogical capabilities of AI-based tutoring models is critical for guided progress in the field. |
| Approach: | They propose an open-source benchmark for holistic tutoring model evaluation. |
| Outcome: | The proposed model can discriminate between expert and novice teachers with high accuracy. |
SANTO: A Web-based Annotation Tool for Ontology-driven Slot Filling (P18-4)
Copied to clipboard
| Challenge: | SANTO is an annotation tool designed for complex relation extraction tasks . a subset of information extraction tasks can be typed n-ary relation extraction or slot filling . |
| Approach: | They propose a domain-adaptive annotation tool for complex slot filling tasks . SANTO enables fast and clearly structured annotation for multiple users in parallel . |
| Outcome: | The proposed tool can be used for slot filling tasks and import and export procedures of standard formats enable interoperability with external sources and tools. |
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Recent studies show that large language models can be used to construct complex multi-agent systems. |
| Approach: | They propose a toolkit for recursive multi-agent systems that supports custom tool-use, delegation schemes, event-based logging, and interactive replay. |
| Outcome: | The proposed tool achieves significant performance gains on agentic benchmarks and identify potential areas of improvement through visualization and debugging tools. |
SIMULEVAL: An Evaluation Toolkit for Simultaneous Translation (2020.emnlp-demos)
Copied to clipboard
| Challenge: | SimulEval is an evaluation toolkit for simultaneous text and speech translation. |
| Approach: | They propose a server-client scheme for simultaneous translation that uses server input and client policies to evaluate models. |
| Outcome: | The proposed evaluation toolkit is available for both text and speech translation. |
SEAGLE: A Platform for Comparative Evaluation of Semantic Encoders for Information Retrieval (D19-3)
Copied to clipboard
| Challenge: | Existing semantic text encoding models are limited in coverage and few attempts to empirically compare them on IR tasks have been made. |
| Approach: | They propose to implement word embedding aggregators and pretrained semantic encoders and to allow for their comparative evaluation on arbitrary IR collections. |
| Outcome: | The proposed model can be exploited via an easy-to-use web interface and its modular backend (micro-service architecture) can easily be extended with additional semantic search models. |
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents (2026.acl-demo)
Copied to clipboard
| Challenge: | Agentic systems are becoming more capable of defining strategies, taking actions, and solving complex, multi-step tasks. |
| Approach: | They propose an automatic, dynamic, and easy-to-use evaluation framework that provides textual insights into agent behavior on three levels of granularity: system, trace, and node. |
| Outcome: | The proposed framework produces high-quality, data-driven, insightful feedback on system, trace, and node. |
Statistical Uncertainty in Word Embeddings: GloVe-V (2024.emnlp-main)
Copied to clipboard
| Challenge: | Static word embeddings are ubiquitous in computational social science applications . however, assessing the statistical uncertainty in downstream conclusions remains challenging . |
| Approach: | They propose a method to obtain approximate, easy-to-use, and scalable reconstruction error variance estimates for one of the most widely used word embedding models. |
| Outcome: | The proposed method enables hypothesis testing in key word embedding tasks. |
Visualizing the “Dictionary of Regionalisms of France” (DRF) (L18-1)
Copied to clipboard
| Challenge: | a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France. |
| Approach: | They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France. |
| Outcome: | The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas . |
Fast Adaptation via Prompted Data: An Efficient Cross-Domain Fine-tuning Method for Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been successful in a variety of natural language understanding tasks, but domain discrepancies between the downstream task and the pre-training corpora may have hindered LLMs to excel further in the vertical applications. |
| Approach: | They propose a Fast Adaptation method for LLMs via Prompted Data that integrates downstream text corpora, gold labels and external knowledge sources into a highly controllable prompt. |
| Outcome: | The proposed method bridges the gap between the downstream task and the pre-training corpora and integrates downstream text corpors, gold labels and external knowledge sources into a highly controllable prompt. |
Increasing the Accessibility of Time-Aligned Speech Corpora with Spokes Mix (L18-1)
Copied to clipboard
| Challenge: | Spokes Mix is an online service providing access to spoken corpora of Polish . high-quality corporata of conversational language are expensive to acquire . |
| Approach: | a new online service provides access to spoken corpora of Polish . the service provides a centralized, easy-to-use corpus query engine with a responsive web interface . |
| Outcome: | the proposed service provides access to spoken corpora of Polish, including three newly released time-aligned collections of manually transcribed spoken-conversational data. |