Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)

80 papers
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard (2025.naacl-industry)

Copied to clipboard

Challenge: Existing studies on the Open Ko-LLM Leaderboard have been limited to five months . this limited analysis of the Open LLM Leaderboard provides a more comprehensive understanding of the progress in developing large language models.
Approach: They conduct a longitudinal study over eleven months to address limitations of previous studies . they analyze 1,769 models over this period to provide a more comprehensive understanding .
Outcome: The study extends observation period of the Open Ko-LLM Leaderboard to eleven months . primary questions are: What are the specific challenges in improving LLM performance?
RTSM: Knowledge Distillation with Diverse Signals for Efficient Real-Time Semantic Matching in E-Commerce (2025.naacl-industry)

Copied to clipboard

Challenge: e-commerce product search is a key component of product discovery and sales in ecommerce . high computational demands of large transformer models pose challenges for their deployment in real-time scenarios.
Approach: They propose a framework for real-time semantic matching that leverages both soft labels from a teacher model and ground truth generated from pairwise query-product and query-query signals.
Outcome: The proposed framework outperforms teacher models and state-of-the-art models on e-commerce datasets.
WorkTeam: Constructing Workflows from Natural Language with Multi-Agents (2025.naacl-industry)

Copied to clipboard

Challenge: Existing workflow construction methods require specialized knowledge and task-switching skills.
Approach: They propose a multi-agent workflow framework that incorporates a supervisor, orchestrator, and filler agent.
Outcome: The proposed framework significantly increases the success rate of workflow construction . the proposed framework is based on a dataset of 3,695 real-world business samples .
How LLMs React to Industrial Spatio-Temporal Data? Assessing Hallucination with a Novel Traffic Incident Benchmark Dataset (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have a high potential to digitize and enhance the health & public services industry.
Approach: They propose to use a cross-lingual benchmark dataset to assess the robustness of state-of-the-art LLMs in the spatio vs temporal domain for traffic incident classification.
Outcome: The proposed model performs well in the spatio-temporal domain and in the non-English context.
Text2Sql: Pure Fine-Tuning and Pure Knowledge Distillation (2025.naacl-industry)

Copied to clipboard

Challenge: Text2Sql is a task that translates natural language questions and database schemas into SQL queries.
Approach: They employ pure fine-tuning strategy to reduce redundancy by using only 53% of the baseline prompt length to fine- tune the model.
Outcome: The model outperforms the baseline model by 8.2% and 8.6% in Test-suite accuracy (TS) and exact-set-match accuracy (EM) under the most refined Spider dev set of prompts, the model achieves 73.5% and 75.4%, respectively, approaching state-of-the-art (SOTA) levels.
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering (2025.naacl-industry)

Copied to clipboard

Challenge: Question Answering (QA) and Visual Question Answers (VQA) are well-studied problems in the language and vision domain.
Approach: They propose a question-answer generation framework that learns attention across multiple sources and decodes this information for robust and unbiased answer generation.
Outcome: The proposed framework can handle thousands of question types and scale to scale.
Finding-Centric Structuring of Japanese Radiology Reports and Analysis of Performance Gaps for Multiple Facilities (2025.naacl-industry)

Copied to clipboard

Challenge: Despite advances in IE, radiology reports are often recorded in free-text format, limiting their secondary application.
Approach: They propose a "Finding-Centric Structuring" approach which organizes reports around individual findings, facilitating secondary use.
Outcome: The proposed approach organizes radiology reports around individual findings, facilitating secondary use.
Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in handling complex dialogue tasks without requiring use case-specific fine-tuning.
Approach: They propose a framework that combines the scalability of LLM-generated labels with the precision of human annotations to achieve higher speed and accuracy comparable to larger models.
Outcome: The proposed framework significantly improves accuracy across utterance-level dialogue tasks, including sentiment detection (over 2%), dialogue act classification (over 1.5%), etc.
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have significantly advanced autonomous agents, particularly in zero-shot tool usage, also known as function calling.
Approach: They propose to integrate function descriptions into prompt formats and introduce a new Decision Token for conditional prompts.
Outcome: The proposed decision token improves function-calling accuracy and relevance detection and a translation pipeline overcomes multilingual limitations.
Exploring Straightforward Methods for Automatic Conversational Red-Teaming (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks.
Approach: They propose to use off-the-shelf large language models to create red-team attacks by eliciting undesired outputs from an attacker LLM.
Outcome: The proposed models can adapt their attack strategies based on prior attempts, but their effectiveness decreases as the alignment of the target model improves.
A Diverse and Effective Retrieval-Based Debt Collection System with Expert Knowledge (2025.naacl-industry)

Copied to clipboard

Challenge: Existing debt collection systems lack script diversity, contextual relevance and coherence due to their complexity.
Approach: They propose a script library based on real-world debt collection conversations and a retrieval based response system for contextual relevance.
Outcome: The proposed system improves script diversity and responds to debtor-collector conversations better through knowledge distillation.
Search Query Embeddings via User-behavior-driven Contrastive Learning (2025.naacl-industry)

Copied to clipboard

Challenge: Existing approaches to embed search queries are limited due to shortness and surface-level variations.
Approach: They propose a user-behavior-driven contrastive learning approach which directly aligns query embeddings according to user intent.
Outcome: The proposed model outperforms state-of-the-art text embedding models on real-world QU tasks while minimizing lexical similarities.
QSpell 250K: A Large-Scale, Practical Dataset for Chinese Search Query Spell Correction (2025.naacl-industry)

Copied to clipboard

Challenge: Chinese Search Query Spell Correction is a task designed to identify and correct typographical errors within queries.
Approach: They propose a large-scale benchmark specifically developed for Chinese Query Spell Correction.
Outcome: The proposed benchmark covers a broad range of topics, including formal entities, everyday colloquialisms and idiomatic expressions.
CONSTRUCTA: Automating Commercial Construction Schedules in Fabrication Facilities with Large Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Existing methods for automating construction scheduling are limited due to their domain knowledge and complexity.
Approach: They propose a framework leveraging LLMs to optimize construction schedules in commercial construction.
Outcome: The proposed framework improves missing value prediction, dependency analysis and planning performance compared to baseline methods.
Challenges and Remedies of Domain-Specific Classifiers as LLM Guardrails: Self-Harm as a Case Study (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have impressive capabilities in generating human-like text, but they pose significant risks in many domains and require guardrails throughout the lifecycle.
Approach: They propose to use a self-harm detector to test the performance of LLM guardrails in real-world environments.
Outcome: The proposed model performs poorly in open and closed domains and is almost unusable in the real world.
Mitigating Bias in Item Retrieval for Enhancing Exam Assembly in Vocational Education Services (2025.naacl-industry)

Copied to clipboard

Challenge: Despite the practical importance of exam assembly, few methods exist to support educators during manual item retrieval for exam assembly tasks.
Approach: They propose a mixed-integer programming re-ranking approach to improve relevance while mitigating bias on an industry-grade exam assembly platform.
Outcome: The proposed approach improves relevance and reduces bias by 17% when compared to other methods on a real-world exam assembly platform.
Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance (2025.naacl-industry)

Copied to clipboard

Challenge: Pretrained language models (PLMs) have revolutionized NLP but amplify linguistic inequities in multilingual applications.
Approach: They evaluate pretrained language models including Mistral, TowerInstruct, OpenHathi, Tamil-Llama, and Kan-Lama across eight languages spanning high-resource and low-resourced settings.
Outcome: The proposed models fail to bridge linguistic divides and are inefficient when compared to other models.
Towards Reliable and Practical Phishing Detection (2025.naacl-industry)

Copied to clipboard

Challenge: Existing datasets lack size and diversity, with only 609 voice phishing samples available in Korean and 638 smishing instances available in English.
Approach: They propose to use a Korean dataset to construct a reliable phishing detection system using language models to evaluate the model's in-domain and unseen attack detection performance.
Outcome: The proposed system performs reasonably well in voice and unseen attacks while smishing detection remains challenging.
Zero-Shot ATC Coding with Large Language Models for Clinical Assessments (2025.naacl-industry)

Copied to clipboard

Challenge: Manual assignment of ATC codes to prescription records is a significant bottleneck in healthcare operations.
Approach: They propose a method to automate the assignment of ATC codes to prescription records . they use locally deployable large language models to guide LLMs through the ontology .
Outcome: The proposed method achieves 78% exact match accuracy with GPT-4o and 60% with Llama 3.1 70B.
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have impacted the writing process, enhancing productivity by collaborating with humans in content creation platforms.
Approach: They propose a framework that uses explicit outlines to guide LLMs in generating goal-oriented, high-quality text.
Outcome: The proposed approach significantly improves text quality according to evaluations by LLMs and professional writers.
TaeBench: Improving Quality of Toxic Adversarial Examples (2025.naacl-industry)

Copied to clipboard

Challenge: Existing adversarial examples generate invalid or ambiguous examples that fool the systems into wrong detection.
Approach: They propose an annotation pipeline for quality control of generated toxic adversarial examples (TAE) they use model-based automated annotation and human-based quality verification to assess quality requirements of a TAE dataset.
Outcome: The proposed pipeline can transfer-attack SOTA toxicity content moderation models and services with adversarial training.
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs (2025.naacl-industry)

Copied to clipboard

Challenge: Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models . however, the leaderboard has faced significant limitations over time due to its academic nature .
Approach: They propose an improved version of the Open Ko-LLM Leaderboard to improve benchmarking . original benchmarks replaced with new tasks that align with real-world capabilities . four new native Korean benchmarks are introduced to better reflect distinct characteristics of Korean language .
Outcome: The proposed framework improves the Open Ko-LLM Leaderboard2 benchmark suite.
CuriousLLM: Elevating Multi-Document Question Answering with LLM-Enhanced Knowledge Graph Reasoning (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved significant success in open-domain question answering, however, they continue to face challenges such as knowledge cutoffs and hallucinations.
Approach: They propose a new mechanism that integrates a curiosity-driven reasoning mechanism into an LLM agent to generate relevant follow-up questions.
Outcome: The proposed enhancement integrates a curiosity-driven reasoning mechanism into an LLM agent, enabling it to generate relevant follow-up questions, thereby guiding the information retrieval process more efficiently.
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents (2025.naacl-industry)

Copied to clipboard

Challenge: Maintaining consistent character personas remains a significant challenge due to variability in information extraction.
Approach: They propose a framework to dynamically reconstruct character personas through Character Persona Training.
Outcome: The proposed framework is evaluated through Big Five personality evaluations and creative tasks, in which characters generate original narratives.
Efficient Continual Pre-training of LLMs for Low-resource Languages (2025.naacl-industry)

Copied to clipboard

Challenge: Open-source large language models (LLMs) are a promising tool for low-resource languages . however, there is still a substantial performance gap between high-resourced languages and LRLs .
Approach: They develop an algorithm to select a subset of texts from a larger corpus and use it to select tokens for LLMs.
Outcome: The proposed algorithm reduces the cost of continual pre-training (CPT) with large amounts of language-specific data.
DSRAG: A Double-Stream Retrieval-Augmented Generation Framework for Countless Intent Detection (2025.naacl-industry)

Copied to clipboard

Challenge: Current intent detection work experiments with minor intent categories.
Approach: They propose a retrieval-augmented generation framework that uses query-to-query and query- to-metadata approaches to retrieve intents from metadata.
Outcome: The proposed framework improves on query-to-query (Q2Q) and query- to-metadata (Q 2M) approaches.
Octopus: On-device language model for function calling of software APIs (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are pivotal for advanced text processing and generation.
Approach: They propose a framework to train on-device Large Language Models optimized for invoking software APIs.
Outcome: The proposed model outperforms GPT-4 in API calling tasks while maintaining inference speed.
MoFE: Mixture of Frozen Experts Architecture (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are characterized by their immense size, often consisting of at least one billion parameters.
Approach: They propose a mixture of Frozen Experts architecture that integrates PEFT and MoE to enhance both training efficiency and model scalability.
Outcome: The proposed architecture outperforms other methods while achieving the highest efficiency.
FinLLM-B: When Large Language Models Meet Financial Breakout Trading (2025.naacl-industry)

Copied to clipboard

Challenge: Existing methods for financial breakout detection are subpar, despite large data and knowledge required.
Approach: They propose a financial breakout dataset and introduce FinLLM-B, a large language model for financial breakout detection.
Outcome: The proposed model outperforms GPT-3.5 in the field of financial breakout detection.
QueryShield: A Platform to Mitigate Enterprise Data Leakage in Queries to External LLMs (2025.naacl-industry)

Copied to clipboard

Challenge: Unrestricted access to external Large Language Models (LLMs) based services like ChatGPT and Gemini can lead to data leakages, especially for large enterprises providing products and services that require confidentiality guarantees.
Approach: They propose a platform that enterprises can use to query external Large Language Models without leaking confidential internal and client information.
Outcome: The proposed solution minimizes data leakage while limiting impact to semantics while preserving the accuracy of the model candidates.
SwissADT: An Audio Description Translation System for Swiss Languages (2025.naacl-industry)

Copied to clipboard

Challenge: despite advances in multilingual machine translation, lack of well-crafted AD data impedes development of audio description translation systems.
Approach: They propose an audio description translation system for three main Swiss languages and English . they combine human expertise with the power of Large Language Models to improve quality .
Outcome: The proposed system is designed to enhance accessibility for multilingual populations in Switzerland.
Chinese Morph Resolution in E-commerce Live Streaming Scenarios (2025.naacl-industry)

Copied to clipboard

Challenge: Live morph resolution task is used to detect e-commerce live streaming violations . morphs are used to evade scrutiny and engage in false advertising .
Approach: They propose a task to detect morph violations in live streaming scenarios . they use large language models to generate additional training data .
Outcome: The proposed method improves performance and improves live streaming regulation.
MonoTODia: Translating Monologue Requests to Task-Oriented Dialogues (2025.naacl-industry)

Copied to clipboard

Challenge: Data scarcity is one of the main problems when it comes to real-world applications of transformer-based models.
Approach: They propose a method to source annotated German monologues from existing monologue material to train TOD systems.
Outcome: The proposed model can be used to train TOD systems on a real-world example of a travel booking service.
MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have been used in clinical decision support, medical education and patient communication.
Approach: They propose a benchmark to evaluate large language models in the domain of medical ethics and assess their grasp of medical ethical principles and their application across diverse scenarios.
Outcome: The proposed framework assesses the models’ grasp of medical ethics principles and their ability to apply them across diverse scenarios.
Predicting ICU Length of Stay for Patients using Latent Categorization of Health Conditions (2025.naacl-industry)

Copied to clipboard

Challenge: Traditional approaches to predicting the duration of a patient's stay in an Intensive Care Unit (ICU) rely on structured clinical data, but recent advances in language models offer significant potential to utilize unstructured text data for ICU length-of-stay (LoS) predictions.
Approach: They propose a method for analyzing nursing notes to predict ICU length-of-stay of patients.
Outcome: The proposed model outperforms baseline models on the MIMIC-III dataset and shows that it significantly outperformed existing models.
RevieWeaver: Weaving Together Review Insights by Leveraging LLMs and Semantic Similarity (2025.naacl-industry)

Copied to clipboard

Challenge: RevieWeaver extracts key product features and provides concise review summaries . a condensed list of key features, pros, and cons, along with a brief summary of customer opinions can help mitigate this issue.
Approach: They propose a framework that extracts key product features and provides concise review summaries.
Outcome: The proposed framework scales efficiently to 30 million reviews and ensures reproducibility and controllability.
MedCodER: A Generative AI Assistant for Medical Coding (2025.naacl-industry)

Copied to clipboard

Challenge: Medical coding is time-consuming and error-prone due to large label space, lengthy text inputs, and the absence of supporting evidence annotations.
Approach: They propose a Generative AI framework for automatic medical coding that leverages extraction, retrieval, and re-ranking techniques as core components.
Outcome: The proposed framework outperforms existing methods on the International Classification of Diseases (ICD) code prediction scale.
Visual Zero-Shot E-Commerce Product Attribute Value Extraction (2025.naacl-industry)

Copied to clipboard

Challenge: Existing zero-shot product attribute value extraction approaches require sellers to manually provide product descriptions.
Approach: They propose a cross-modal zero-shot attribute value generation framework based on CLIP that uses product images as inputs for zero- shot inference.
Outcome: The proposed framework significantly outperforms other vision-language models for zero-shot attribute value extraction.
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Typical evaluations of Large Language Models (LLMs) report a single accuracy metric per dataset, often derived from an optimized setup.
Approach: They propose a framework for non-adversarial evaluation of large language models that evaluates models by repeatedly testing them on the same benchmarks in various setups.
Outcome: The proposed framework evaluates models by repeatedly testing them on the same benchmarks in various setups to give a realistic estimate of their accuracy and consistency.
Evaluating Large Language Models with Enterprise Benchmarks (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks lack domain-specific datasets for evaluating large language models . existing benchmarks often lack domain specific datasets, which can be difficult to convert to standardized metrics or regulatory issues.
Approach: They propose to use 25 publicly available domain-specific English benchmarks from diverse domains . they propose to combine a wide range of natural language processing tasks for holistic evaluation .
Outcome: The proposed framework includes 25 publicly available domain-specific English benchmarks from diverse enterprise domains like financial services, legal, climate, cyber security, and 2 public Japanese finance benchmarks.
Can Post-Training Quantization Benefit from an Additional QLoRA Integration? (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models require considerable computing resources, which can be costly and often unavailable.
Approach: They propose to integrate 4-bit Post-training Quantization with QLoRA to address these issues . they demonstrate that the integration outperforms standard quantization and fine-tuning .
Outcome: The proposed integration outperforms standard PTQ and 16-bit full-parameter fine-tuning on LLMs.
From Generating Answers to Building Explanations: Integrating Multi-Round RAG and Causal Modeling for Scientific QA (2025.naacl-industry)

Copied to clipboard

Challenge: Application of Large Language Models to complex causal question answering can be stymied by their opacity and propensity for hallucination.
Approach: They propose a causal QA approach that combines iterative RAG with a formal model of causation.
Outcome: The proposed approach is implemented into a Collaborative Research Assistant (Cora) and evaluated in the life sciences domain.
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice (2025.naacl-industry)

Copied to clipboard

Challenge: Existing methods for jailbreaking large-language models are limited by their limitations . authors present a mutation-based fuzzing technique that generates effective jailbreaking templates .
Approach: They propose a mutation-based fuzzing technique for efficiently finding effective jailbreaking templates that combine with harmful questions to generate harmful responses.
Outcome: The proposed technique achieves 95% attack success rates on public datasets for leading LLMs . it also shows impressive generalizability to unseen harmful questions and improves model defenses to prompt attacks.
Does Self-Attention Need Separate Weights in Transformers? (2025.naacl-industry)

Copied to clipboard

Challenge: Experimental results show a 66.53% reduction in parameter size within the attention block and competitive accuracy improvements of 3.55% and 0.89% over symmetric and pairwise attention-based models, respectively.
Approach: They propose a simplified approach where a single weight matrix is used for Keys, Queries, and Values instead of separate matrices for each.
Outcome: The proposed approach outperforms the BERT baseline on GLUE tasks even outperforming the standard BERT model in handling noisy and out-of-domain data.
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling (2025.naacl-industry)

Copied to clipboard

Challenge: Existing methods that only deal with flat text chunks use a graph structure to handle complex questions.
Approach: They propose layout-aware graph modeling for multimodal RAG using document layout parsing to take into account relationship of multimodalities.
Outcome: The proposed method can handle complex questions that require information from multimodalities.
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being used for communication tasks across different regions.
Approach: They propose a benchmark to evaluate whether Large Language Models are ethically aligned and can be used in real-world situations.
Outcome: The proposed benchmark evaluates whether LLMs comply with or resist swearing instructions and assesses their alignment with ethical frameworks, cultural nuances, and language comprehension capabilities.
Natural Language Processing for Human Resources: A Survey (2025.naacl-industry)

Copied to clipboard

Challenge: Recent advances in NLP have the potential to transform HR processes, from recruitment to employee management.
Approach: They analyze key tasks such as information extraction and text classification and their roles in downstream applications like recommendation and language generation while discussing ethical concerns.
Outcome: The proposed frameworks can be applied to HR tasks and to recommendation, language generation, and interaction.
Implementing Retrieval Augmented Generation Technique on Unstructured and Structured Data Sources in a Call Center of a Large Financial Institution (2025.naacl-industry)

Copied to clipboard

Challenge: generative AI models can extract accurate facts from external unstructureddata sources.
Approach: They propose to use data embeddings to enhance RAG for retrieving facts from structured data sources, achieving low latency and high reliability.
Outcome: The proposed model achieves low latency and high reliability and average response time of 7.33 seconds.
Granite Guardian: Comprehensive LLM Safeguarding (2025.naacl-industry)

Copied to clipboard

Challenge: a suite of advanced models is designed to detect and mitigate risks associated with prompts and responses.
Approach: a team of researchers develop a model family to detect and mitigate risks associated with prompts and responses. the model family is based on the Granite 3.0 language models.
Outcome: a new model family is designed to detect and mitigate risks associated with prompts and responses.
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions (2025.naacl-industry)

Copied to clipboard

Challenge: Streaming automatic speech recognition models use high power consumption to improve usability and accuracy.
Approach: They propose to optimize on-device speech recognition models by adjusting component energy sensitivities based on their specific energy sensitities to reduce power consumption.
Outcome: The proposed approach achieves up to 47% lower energy usage while preserving comparable model accuracy and improving real-time performance compared to leading methods.
Break-Ideate-Generate (BrIdGe): Moving beyond Translations for Localization using LLMs (2025.naacl-industry)

Copied to clipboard

Challenge: Traditional methods of localization focus on linguistic conversion, but content needs to align with the target audience’s cultural norms, nuances, and technical requirements.
Approach: They propose a new approach for Large Languages Models (LLMs) called Break-Ideate-Generate (BrIdGe) that breaks the source content into granular facts, organizes the granules and executes the plan to ‘generate’ localized content.
Outcome: The proposed approach 'breaks' the source content into granular facts, ‘ideates’ an action plan for content creation in the target language by organizing the granules, and finally executes the plan to ‘generate’ localized content.
Concept Distillation from Strong to Weak Models via Hypotheses-to-Theories Prompting (2025.naacl-industry)

Copied to clipboard

Challenge: Concept Distillation (CD) is an automated prompt optimization technique for enhancing weaker models on complex tasks.
Approach: They propose an automatic prompt optimization technique for enhancing weaker models on complex tasks using a base prompt and a strong model to generate reasons for these mistakes.
Outcome: The proposed technique improves weaker models on NL2Code and mathematical reasoning tasks, while preserving performance.
Towards Reliable Agents: Benchmarking Customized LLM-Based Retrieval-Augmented Generation Frameworks with Deployment Validation (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks for general-purpose RAG systems, such as CRAG, RGB, MultiHop-RAG, and CRUD-RAGG, are limited and lack a benchmark specifically tailored to evaluate frameworks.
Approach: They evaluated OpenAI’s Assistants API versus a RAG assistant built with Langchain and deployed a system based on benchmark insights as a course assistant over a two-year span.
Outcome: The proposed benchmarks show that domain-specific retrieval impacts response accuracy and highlight key challenges in real-world deployment.
Query Variant Detection Using Retriever as Environment (2025.naacl-industry)

Copied to clipboard

Challenge: Identifying query variants is nontrivial as highly similar query pairs may fail to differ in word form, order, or phrasing despite sharing the same intent.
Approach: They propose to use retrieval as an environment feedback to capture semantic equivalence . experimental results demonstrate the efficacy of the proposed method .
Outcome: The proposed method improves query variant detection across diverse scenarios.
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have potential to automate hiring but inherent biases may lead to unfair hiring practices.
Approach: They evaluate how factors such as gender, race, and educational background influence model decisions.
Outcome: The proposed model reduces biases related to gender and race, but implicit biase concerning educational background remains significant.
Goal-Driven Data Story, Narrations and Explanations (2025.naacl-industry)

Copied to clipboard

Challenge: Unlike existing tools, our system addresses the ambiguity of vague, multi-line queries, setting a new benchmark in data storytelling by tackling complexities no existing system comprehensively handles.
Approach: They propose a system that processes and interprets vague, open-ended, and multi-line complex queries, transforming them into coherent, actionable data stories.
Outcome: The proposed system processes and interprets vague, open-ended, and multi-line complex queries, transforming them into coherent, actionable data stories.
VIT-Pro: Visual Instruction Tuning for Product Images (2025.naacl-industry)

Copied to clipboard

Challenge: general-purpose vision-language models struggle to understand and converse about real-world e-commerce product images.
Approach: a new approach is proposed to use large-scale image-text pairs to train a generative VLM for e-commerce product images.
Outcome: The proposed model outperforms general-purpose VLMs on multiple vision tasks in the e-commerce domain.
AutoKB: Automated Creation of Structured Knowledge Bases for Domain-Specific Support (2025.naacl-industry)

Copied to clipboard

Challenge: Effective customer support requires domain-specific solutions tailored to users’ issues.
Approach: They propose an automated pipeline for building a domain-specific KB with a hierarchical tree structure that maps user issues to precise and domain-compliant solutions.
Outcome: Experiments in troubleshooting and medical domains show that the proposed pipeline outperforms LLMs and unstructured knowledge bases and is 75 times more cost-effective than manual methods.
Medical Spoken Named Entity Recognition (2025.naacl-industry)

Copied to clipboard

Challenge: Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc.
Approach: They present a spoken NER dataset in the medical domain using pre-trained models that are encoder-only and sequence-to-sequence.
Outcome: The dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.
PLEX: Adaptive Parameter-Efficient Fine-Tuning for Code LLMs using Lottery-Tickets (2025.naacl-industry)

Copied to clipboard

Challenge: PLEX is a lottery-ticket based parameter-efficient fine-tuning method that adapts large language models to well-supported and underrepresented programming languages (PLs) in pretraining.
Approach: They propose a lottery-ticket based parameter-efficient fine-tuning method that adapts large language models to well-supported and underrepresented programming languages (PLs)
Outcome: The proposed method achieves state-of-the-art performance among PEFT methods while maintaining competitive results with reduced computational overhead.
Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain (2025.naacl-industry)

Copied to clipboard

Challenge: a conversational AI system that answers standard airport queries and resolves airport terminology is ideal for airports from the top 20 in terms of annual passenger numbers.
Approach: They propose a Conversational AI system that enables staff to communicate with flight systems . the system answers standard airport queries and resolves airport terminology .
Outcome: The proposed system answers standard airport queries and resolves airport terminology, jargon, abbreviations and dynamic questions involving reasoning.
LLM Safety for Children (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly impacting children through education, toys, and therapy, offering benefits like improved mental health and parental controls.
Approach: They propose a comprehensive approach to evaluating LLM safety specifically for children by listing potential risks that children may encounter when using LLM-powered applications.
Outcome: The proposed model bridges the gap in child safety literature across various fields.
RxLens: Multi-Agent LLM-powered Scan and Order for Pharmacy (2025.naacl-industry)

Copied to clipboard

Challenge: paper prescriptions are difficult for customers to interpret and are often unstructured, handwritten, and illegible.
Approach: They propose a multi-step Large Language Model-based solution for automated pharmacy cart construction.
Outcome: The proposed solution can yield up to 19% - 40% and 11% - 26% increase in Recall@3 relative to SOTA methods.
Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs (2025.naacl-industry)

Copied to clipboard

Challenge: Domain- and customer-specific requirements complicate the problem of NL2SQL customization.
Approach: They propose a distilled customization framework tailored for NL2SQL tasks.
Outcome: The proposed framework outperforms teacher models on three benchmarks and achieves an average improvement of 36% in execution accuracy.
eC-Tab2Text: Aspect-Based Text Generation from e-Commerce Product Tables (2025.naacl-industry)

Copied to clipboard

Challenge: eC-Tab2Text dataset is designed to capture product attributes and user-specific queries.
Approach: They propose a novel dataset to capture the intricacies of e-commerce including detailed product attributes and user-specific queries.
Outcome: The proposed dataset outperforms existing generalpurpose LLMs in generating accurate product reviews.
RAD-Bench: Evaluating Large Language Models’ Capabilities in Retrieval Augmented Dialogues (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks assess LLMs' chat abilities in multi-turn dialogues or their use of retrieval for augmented responses in limited tasks such as knowledge QA or numeric reasoning.
Approach: They propose a benchmark to evaluate LLMs' capabilities in multi-turn dialogues following retrievals.
Outcome: The proposed benchmark evaluates LLMs' ability to perform in multi-turn dialogues following retrievals over 6 representative scenarios.
Conflict and Overlap Classification in Construction Standards Using a Large Language Model (2025.naacl-industry)

Copied to clipboard

Challenge: Current manual approaches to analyzing overlapping or conflicting content are time-consuming, costly, and error-prone.
Approach: They propose a large language model that uses a construction domain-adapted large language for the semantic comparison of sentences in construction standards.
Outcome: The proposed framework achieves 97.9% accuracy and 0.907 macro F1-score in classifying sentences from Korean construction standards as overlapping, conflicting, or neutral.
Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text (2025.naacl-industry)

Copied to clipboard

Challenge: Proteins play critical roles in biological systems, yet 99.7% of 227 million known protein sequences remain uncharacterized due to the limitations of experimental methods.
Approach: They propose a multimodal large language model that interprets protein sequences and generates informative text to address open-ended questions about protein functions and attributes.
Outcome: The proposed model outperforms existing models in open-ended question-answering tasks.
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia (2025.naacl-industry)

Copied to clipboard

Challenge: Using the entire dataset, shuffling answer options introduces instability in the insurance and finance sectors.
Approach: They propose a dataset for evaluation of performance in vocational and professional certification exams in Indonesia.
Outcome: The proposed dataset includes 8,834 multiple-choice questions from 27 large language models across six key sectors.
CodeGenWrangler: Data Wrangling task automation using Code-Generating Models (2025.naacl-industry)

Copied to clipboard

Challenge: Tabular datasets in industrial settings often encompass extensive data with numerous rows and columns.
Approach: They propose a system that leverages large language models to generate executable code for data wrangling tasks . they identify inherent patterns in the data while leveraging external knowledge .
Outcome: The proposed system detects patterns in the data while leveraging external knowledge . it generates executable code for data-wrangling tasks like missing value imputation and error correction .
Dialogue Language Model with Large-Scale Persona Data Engineering (2025.naacl-industry)

Copied to clipboard

Challenge: Existing persona-consistent dialogue models lack robustness due to limited scale and diversity of datasets.
Approach: They propose an open-domain persona dialogue system that employs extensive generative pre-training on a persona dialog dataset to enhance persona consistency.
Outcome: The proposed model generates vast persona dialogue datasets and addresses invalid persona bias.
Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service (2025.naacl-industry)

Copied to clipboard

Challenge: Hallucination is a problem in large language models that produce incorrect output . authors propose a reliable and high-speed production system to detect and rectify hallucinations .
Approach: They propose a high-speed production system that detects hallucinations in LLMs . they propose NER, natural language inference, span-based detection and a rewriting mechanism .
Outcome: The proposed system detects a wide range of hallucinations in LLM responses.
Improved Near-Duplicate Detection for Aggregated and Paywalled News-Feeds (2025.naacl-industry)

Copied to clipboard

Challenge: News aggregators provide comprehensive and timely news stories that are sourced from diverse sources but differ in phrasing, formatting or supplemented with additional details.
Approach: They propose a method that combines embeddings from pretrained language model and latent metadata of a news article followed by community detection to identify clusters of near-duplicates.
Outcome: The proposed approach can detect nuanced similarities and differences in news snippets using pretrained language model and latent metadata of a news article followed by community detection.
Pisets: A Robust Speech Recognition System for Lectures and Interviews (2025.naacl-industry)

Copied to clipboard

Challenge: Sustainable speech recognition systems are essential for scientists, journalists, and anyone processing audio recordings of interviews and meetings.
Approach: They propose a speech-to-text system "Pisets" which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model.
Outcome: The proposed system ensures robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model.
CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search (2025.naacl-industry)

Copied to clipboard

Challenge: Relevance modeling between queries and items is a key component of commercial search engines.
Approach: They propose a framework for continual pre-training of LLMs to enhance domain knowledge . they employ queries and multi-field item to jointly pre-train for enhancing domain knowledge.
Outcome: The proposed model achieves convincing performance compared to strong baselines.
Schema and Natural Language Aware In-Context Learning for Improved GraphQL Query Generation (2025.naacl-industry)

Copied to clipboard

Challenge: GraphQL is a flexible alternative to REST APIs, but generating complex queries remains challenging.
Approach: They propose a framework that integrates GraphQL schemas with natural language inputs to improve query generation accuracy.
Outcome: The proposed framework improves performance on a publicly available complex GraphQL dataset.
Chatbot Arena Estimate: towards a generalized performance benchmark for LLM capabilities (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmark aggregation methods, such as Elo-based systems, can be resource-intensive, public facing, and time-consuming.
Approach: They propose a framework for aggregating performance across diverse benchmarks that generates a “Goodness” and a ‘Fastness” score.
Outcome: The proposed framework achieves higher Pearson correlation with Chatbot Arena Elo scores than MMLU’s correlation with chatbot Arena scores, validating its reliability for real-world LLM evaluation.
Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Recent literature focuses on constructing large audio language models (LALMs) but they are limited in temporal reasoning, which may hinder commercial applications .
Approach: They propose a data augmentation technique for generating reliable audio temporal questions and answers using an LLM.
Outcome: The proposed model performs well on public audio benchmark datasets and is optimized for edge applications.
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) face limitations due to outdated knowledge, hallucinations, and poor reasoning in complex contexts.
Approach: They propose a Hybrid Parameter-Adaptive RAG system for the AI legal domain with NYC Local Law 144 as the test case.
Outcome: The proposed system improves retrieval accuracy, response fidelity, and contextual precision on NYC Local Law 144 . Empirical evidence indicates that many AI tools overstate their ability to prevent hallucinations in legal and policy contexts.
An Efficient Context-Dependent Memory Framework for LLM-Centric Agents (2025.naacl-industry)

Copied to clipboard

Challenge: a recent study has demonstrated that context-dependent memory encoding can help to retrieve key memory cues essential for problem-solving.
Approach: They propose an efficient architecture miming human memory processes through multistage encoding, context-aware storage, and retrieval strategies for LLM-centric agents.
Outcome: The proposed architecture surpasses state-of-the-art online LLM-centric approaches on two interactive decision-making benchmarks in the navigation and manipulation domain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations