Papers with CAT

29 papers
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation (2026.acl-demo)

Copied to clipboard

Challenge: Existing document translation pipelines face a tension between linguistic processing and layout preservation.
Approach: They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content.
Outcome: The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision.
BiSync: A Bilingual Editor for Synchronized Monolingual Texts (2023.acl-demo)

Copied to clipboard

Challenge: CAT systems often interfere with writing process by requiring users to access external resources.
Approach: They propose a bilingual writing assistant that allows users to freely compose text in two languages while maintaining the two monolingual texts synchronized.
Outcome: The proposed bilingual writing assistant can produce high accuracy with limited computational resources.
SmartMatch: Real-Time Semantic Retrieval for Translation Memory Systems (2026.eacl-demo)

Copied to clipboard

Challenge: Translation Memory (TM) systems are core components of computer-aided translation tools . however, they fail to retrieve semantically relevant content when surface similarity is low.
Approach: They propose an open-source demo and evaluation toolkit for TM retrieval that connects modern sentence encoders and strong lexical/fuzzy baselines with a vector database.
Outcome: The proposed toolkit exposes the end-to-end retrieval pipeline through a web-based UI for qualitative inspection and preference logging.
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks (2025.coling-industry)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) is a popular technique for parameter-efficient fine-tuning of Large Language Models.
Approach: They propose to combine LoRA modules to achieve skill composition . they propose to use concatenation of LoRAs to optimize weights for different LoRA training .
Outcome: The proposed model outperforms existing models and data- merging techniques on math-word problems and domain-specialized corpora.
Multi-Agent Orchestration for Terminology-Constrained Machine Translation in Industrial Localization (2026.acl-industry)

Copied to clipboard

Challenge: Accurate terminology is a non-negotiable requirement in industrial localization processes.
Approach: They propose a multi-agent LLM pipeline that orchestrates four specialized agents for terminology-constrained machine translation.
Outcome: The proposed system achieves 99.4% average accuracy while outperforming other systems on the WMT25 Terminology Translation benchmark.
A Compare Aggregate Transformer for Understanding Document-grounded Dialogue (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have focused on KS in unstructured documents, but dialogue history that is not related to the current dialogue may introduce noise in the KS processing.
Approach: They propose a Compare Aggregate Transformer to jointly denoise the dialogue context and aggregate the document information for response generation.
Outcome: The proposed model outperforms the state-of-the-art approach and strong baselines on a CMU_DoG dataset.
Towards Explainable Computerized Adaptive Testing with Large Language Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on minimizing the number of questions required to assess ability, lacking clear and reliable explanations for the question selection process.
Approach: They propose to use large language models to enhance computer adaptive testing (CAT) by providing human-like interpretability and explanations.
Outcome: The proposed agent-based CAT performs comparably or superior to traditional CAT methods in accuracy and significantly improves student trust and satisfaction.
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models (2026.acl-industry)

Copied to clipboard

Challenge: Existing compression methods for large reasoning models rely on uniform length reduction or coarse-grained difficulty estimation, often leading to performance degradation on difficult problems.
Approach: They propose a framework that incorporates model’s intrinsic self-certainty signals as confidence into the preference optimization process, which autonomously modulates reasoning lengths based on problem difficulty.
Outcome: The proposed framework outperforms state-of-the-art models on reasoning accuracy across multiple benchmarks on different base models.
Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have found that datasets with paired inputs are prone to spurious correlations, resulting in models trained only on those outperform chance.
Approach: They propose a counterfactual attentiveness test to measure reliance on spurious correlations by replacing part of the input with its counterpart from a different example.
Outcome: The proposed method improves models' attentiveness on ten datasets spanning four tasks: natural language inference, reading comprehension, paraphrase detection, and visual & language reasoning.
CATs are Fuzzy PETs: A Corpus and Analysis of Potentially Euphemistic Terms (2022.lrec-1)

Copied to clipboard

Challenge: Euphemisms are a difficult topic because they are subject to language change and humans may not agree on what is a euphemist.
Approach: They analyze a corpus of potentially euphemistic terms (PETs) and examples from the GloWbE corpus to examine their meanings.
Outcome: The proposed corpus of potentially euphemistic terms and examples from the GloWbE corpus show that PETs generally decrease negative and offensive sentiment.
A Knowledge Graph Reasoning-Based Model for Computerized Adaptive Testing (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have failed to account for the differences in concept relevance when a question involves multiple concepts .
Approach: They propose a Knowledge Graph Reasoning-Based Model for CAT that captures semantic and relational information between concepts and questions and incorporates multiple evaluation objectives.
Outcome: The proposed model outperforms existing methods on three authentic educational datasets.
GWLAN: General Word-Level AutocompletioN for Computer-Aided Translation (2021.acl-long)

Copied to clipboard

Challenge: Computer-aided translation (CAT) is a form of software that assists a human translator in the translation process.
Approach: They propose to use computer-aided translation (CAT) to assist a human translator in the translation process.
Outcome: The proposed method can give significantly more accurate predictions than baseline methods on CAT datasets.
Evaluating EcoLexiCAT: a Terminology-Enhanced CAT Tool (L18-1)

Copied to clipboard

Challenge: EcoLexiCAT is a web-based tool for terminology-enhanced translation of environmental texts . most terminological modules in CAT tools do not go beyond a simple glossary of source and target terms .
Approach: They propose to integrate terminology-enhanced translation into a web-based tool . EcoLexiCAT is a terminology-enriched CAT tool for the English-Spanish-English translation .
Outcome: The EcoLexiCAT tool is an open-source version of the CAT tool MateCat . it enriches a source text with information from a multimodal and multilingual terminological knowledge base on the environment .
Counterfactual Adversarial Learning with Representation Interpolation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models with statistical bias are prone to memorized correlations . large pre-trained models such as BERT have revolutionized the model development paradigm in natural language processing .
Approach: They propose a framework to tackle the problem from a causal perspective using a latent space interpolation approach.
Outcome: Extensive experiments show that CAT achieves substantial performance improvement over SOTA across different downstream tasks, including sentence classification, natural language inference and question answering.
Language Data Sharing in European Public Services – Overcoming Obstacles and Creating Sustainable Data Sharing Infrastructures (2020.lrec-1)

Copied to clipboard

Challenge: Data is key in training modern language technologies.
Approach: They summarise findings of first pan-European study on barriers to language data sharing . they identify structural challenges, disposition towards CAT tools and lack of digital skills . overcoming language barriers is one of the main challenges european citizens face .
Outcome: The paper summarises the findings of the first pan-European study on barriers to language data sharing . the findings highlight the barriers and recommend solutions to overcome them .
The FISKMÖ Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic Research (2020.lrec-1)

Copied to clipboard

Challenge: Finnish and Swedish are the two official languages of Finland.
Approach: They propose to compile a massive corpus of translated material between Finnish and Swedish . they also aim to develop open and freely accessible translation services for those two languages .
Outcome: The project aims to develop open and freely accessible translation services for Finnish and Swedish.
Let the CAT out of the bag: Contrastive Attributed explanations for Text (2022.emnlp-main)

Copied to clipboard

Challenge: XAI has seen an explosion of interest in explaining black box behavior . contrastive/counterfactual explanations have seen a surge of interest recently .
Approach: They propose a method which provides contrastive explanations for natural language text data with a novel twist by exploiting attribute classifiers.
Outcome: The proposed method outperforms state-of-the-art methods on four benchmark metrics.
CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing fine-tuning paradigms focus on aligning LLMs with task-specific objectives.
Approach: They propose a pipeline that leverages human priors to automatically generate token-level causal signals and introduce the Re-Attention mechanism to guide training.
Outcome: The proposed pipeline achieves an average improvement of 5.76% on the STG dataset and 1.56% on downstream tasks.
Learning to Decouple Relations: Few-Shot Relation Classification with Entity-Guided Attention and Confusion-Aware Training (2020.coling-main)

Copied to clipboard

Challenge: Existing few-shot relation classifiers struggle to distinguish them with few annotated instances due to high co-occurrence of some relations .
Approach: They propose a few-shot relation classification model with two mechanisms to decouple easily-confused relations.
Outcome: The proposed model achieves comparable and even better results to strong baselines in terms of accuracy.
Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation (2022.emnlp-main)

Copied to clipboard

Challenge: CAT tools based on translation memories (TMs) are limited in their use for a number of translation tasks due to the limited availability of in-domain TMs.
Approach: They propose a neural approach to exploit in-domain TMs and in-target-language (TL) monolingual corpora to exploit CAT tools.
Outcome: The proposed approach exploits in-domain TMs and in-target-language (TL) monolingual corpora and increases translation proposals on four language pairs.
Critical Learning Periods: Leveraging Early Training Dynamics for Efficient Data Pruning (2024.findings-acl)

Copied to clipboard

Challenge: Neural Machine Translation models are extremely data-hungry and require a large dataset to maintain data quality.
Approach: They propose a new data pruning technique that leverages early model training dynamics to identify the most relevant data points for model performance.
Outcome: The proposed technique outperforms the benchmarks on indo-European languages while pruning up to 50% of training data.
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede’s Cultural Dimensions (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) are deployed in many countries, but they fail to account for cultural variances among their potential users.
Approach: They propose to use Hofstede’s cultural dimension framework to quantify cultural alignment using latent variable analysis to evaluate large language models against cultural dimensions of regions like the United States, China, and Arab countries.
Outcome: The proposed model is compared against LLMs in the United States, China, and Arab countries and demonstrates that all models struggle to grasp cultural values, while GPT-4 shows a unique capability to adapt to cultural nuances, particularly in Chinese settings.
Supervised Adversarial Contrastive Learning for Emotion Recognition in Conversations (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to recognize emotions have limitations in discovering the intrinsic structure of data relevant to emotion labels, and struggle to extract generalized and robust representations.
Approach: They propose a supervised adversarial contrastive learning framework for learning class-spread structured representations in a controlled manner.
Outcome: The proposed framework can extract generalized and robust representations on three datasets and achieves state-of-the-art performance.
TallVocabL2Fi: A Tall Dataset of 15 Finnish L2 Learners’ Vocabulary (2022.lrec-1)

Copied to clipboard

Challenge: Existing work on second language knowledge has focused on the knowledge of small numbers of words, often geared towards measuring vocabulary size.
Approach: They propose a “tall” word knowledge response dataset containing information about a few learners’ knowledge of many words.
Outcome: The proposed dataset is based on a self-rating test and translation test and is compared with previous comparable datasets.
CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense Reasoning (2023.acl-long)

Copied to clipboard

Challenge: HKUST-KnowComp proposes a framework for commonsense reasoning that can be used to conceptualize commonsence knowledge bases at scale.
Approach: They propose a framework that integrates event conceptualization and instantiation to conceptualize commonsense knowledge bases at scale.
Outcome: The proposed framework achieves state-of-the-art on two conceptualization tasks and the acquired abstract commonsense knowledge significantly improves commonsence inference modeling.
INarIG: Iterative Non-autoregressive Instruct Generation Model For Word-Level Auto Completion (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models for word-level autocompletion (WLAC) only use human typed sequences as prefixes in decoding module.
Approach: They propose a novel iterative nonautoregressive instruct generation model for WLAC task . it uses human typed sequences and iterating decoding with subwords to fully utilize input information.
Outcome: The proposed model is more competent in dealing with low-frequency words, and achieves state-of-the-art results on the WMT22 and benchmark datasets.
How do LLMs’ Preferences Affect Event Argument Extraction? CAT: Addressing Preference Traps in Unsupervised EAE (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to supervised EAE suffer from preference traps due to misalignments between prior knowledge, instructions, or output constraints and LLMs’ preferences.
Approach: They propose an unsupervised EAE framework that handles LLMs' preference traps by targeting their prior knowledge and instructions.
Outcome: The proposed framework matches the best DeepSeek-R1 API model with a significantly lower time cost.
Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models excel at linear reasoning tasks but are underexplored on non-linear structures such as natural debates.
Approach: They evaluate whether Large Language Models can approximate structured reasoning from Computational Argumentation Theory.
Outcome: The proposed model performs well on dialogue-formatted debates without access to the underlying graph.
Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique (2025.emnlp-main)

Copied to clipboard

Challenge: Consensual Assessment Technique (CAT) for large language models is used to evaluate creativity, but is costly and time-consuming with non-experts.
Approach: They adapt the Consensual Assessment Technique (CAT) for Large Language Models to a 90-poem dataset with a ground truth based on publication venue.
Outcome: The proposed method outperforms the best human non-expert evaluations by significantly outperforming the best language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations