Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)
Copied to clipboard
| Challenge: | Existing studies on the Open Ko-LLM Leaderboard have been limited to five months . this limited analysis of the Open LLM Leaderboard provides a more comprehensive understanding of the progress in developing large language models. |
| Approach: | They conduct a longitudinal study over eleven months to address limitations of previous studies . they analyze 1,769 models over this period to provide a more comprehensive understanding . |
| Outcome: | The study extends observation period of the Open Ko-LLM Leaderboard to eleven months . primary questions are: What are the specific challenges in improving LLM performance? |
Copied to clipboard
| Challenge: | e-commerce product search is a key component of product discovery and sales in ecommerce . high computational demands of large transformer models pose challenges for their deployment in real-time scenarios. |
| Approach: | They propose a framework for real-time semantic matching that leverages both soft labels from a teacher model and ground truth generated from pairwise query-product and query-query signals. |
| Outcome: | The proposed framework outperforms teacher models and state-of-the-art models on e-commerce datasets. |
Copied to clipboard
| Challenge: | Existing workflow construction methods require specialized knowledge and task-switching skills. |
| Approach: | They propose a multi-agent workflow framework that incorporates a supervisor, orchestrator, and filler agent. |
| Outcome: | The proposed framework significantly increases the success rate of workflow construction . the proposed framework is based on a dataset of 3,695 real-world business samples . |
Copied to clipboard
| Challenge: | Large language models (LLMs) have a high potential to digitize and enhance the health & public services industry. |
| Approach: | They propose to use a cross-lingual benchmark dataset to assess the robustness of state-of-the-art LLMs in the spatio vs temporal domain for traffic incident classification. |
| Outcome: | The proposed model performs well in the spatio-temporal domain and in the non-English context. |
Copied to clipboard
| Challenge: | Text2Sql is a task that translates natural language questions and database schemas into SQL queries. |
| Approach: | They employ pure fine-tuning strategy to reduce redundancy by using only 53% of the baseline prompt length to fine- tune the model. |
| Outcome: | The model outperforms the baseline model by 8.2% and 8.6% in Test-suite accuracy (TS) and exact-set-match accuracy (EM) under the most refined Spider dev set of prompts, the model achieves 73.5% and 75.4%, respectively, approaching state-of-the-art (SOTA) levels. |
Copied to clipboard
| Challenge: | Question Answering (QA) and Visual Question Answers (VQA) are well-studied problems in the language and vision domain. |
| Approach: | They propose a question-answer generation framework that learns attention across multiple sources and decodes this information for robust and unbiased answer generation. |
| Outcome: | The proposed framework can handle thousands of question types and scale to scale. |
Copied to clipboard
| Challenge: | Despite advances in IE, radiology reports are often recorded in free-text format, limiting their secondary application. |
| Approach: | They propose a "Finding-Centric Structuring" approach which organizes reports around individual findings, facilitating secondary use. |
| Outcome: | The proposed approach organizes radiology reports around individual findings, facilitating secondary use. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable capabilities in handling complex dialogue tasks without requiring use case-specific fine-tuning. |
| Approach: | They propose a framework that combines the scalability of LLM-generated labels with the precision of human annotations to achieve higher speed and accuracy comparable to larger models. |
| Outcome: | The proposed framework significantly improves accuracy across utterance-level dialogue tasks, including sentiment detection (over 2%), dialogue act classification (over 1.5%), etc. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have significantly advanced autonomous agents, particularly in zero-shot tool usage, also known as function calling. |
| Approach: | They propose to integrate function descriptions into prompt formats and introduce a new Decision Token for conditional prompts. |
| Outcome: | The proposed decision token improves function-calling accuracy and relevance detection and a translation pipeline overcomes multilingual limitations. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks. |
| Approach: | They propose to use off-the-shelf large language models to create red-team attacks by eliciting undesired outputs from an attacker LLM. |
| Outcome: | The proposed models can adapt their attack strategies based on prior attempts, but their effectiveness decreases as the alignment of the target model improves. |
Copied to clipboard
| Challenge: | Existing debt collection systems lack script diversity, contextual relevance and coherence due to their complexity. |
| Approach: | They propose a script library based on real-world debt collection conversations and a retrieval based response system for contextual relevance. |
| Outcome: | The proposed system improves script diversity and responds to debtor-collector conversations better through knowledge distillation. |
Copied to clipboard
| Challenge: | Existing approaches to embed search queries are limited due to shortness and surface-level variations. |
| Approach: | They propose a user-behavior-driven contrastive learning approach which directly aligns query embeddings according to user intent. |
| Outcome: | The proposed model outperforms state-of-the-art text embedding models on real-world QU tasks while minimizing lexical similarities. |
Copied to clipboard
| Challenge: | Chinese Search Query Spell Correction is a task designed to identify and correct typographical errors within queries. |
| Approach: | They propose a large-scale benchmark specifically developed for Chinese Query Spell Correction. |
| Outcome: | The proposed benchmark covers a broad range of topics, including formal entities, everyday colloquialisms and idiomatic expressions. |
Copied to clipboard
| Challenge: | Existing methods for automating construction scheduling are limited due to their domain knowledge and complexity. |
| Approach: | They propose a framework leveraging LLMs to optimize construction schedules in commercial construction. |
| Outcome: | The proposed framework improves missing value prediction, dependency analysis and planning performance compared to baseline methods. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have impressive capabilities in generating human-like text, but they pose significant risks in many domains and require guardrails throughout the lifecycle. |
| Approach: | They propose to use a self-harm detector to test the performance of LLM guardrails in real-world environments. |
| Outcome: | The proposed model performs poorly in open and closed domains and is almost unusable in the real world. |
Copied to clipboard
| Challenge: | Despite the practical importance of exam assembly, few methods exist to support educators during manual item retrieval for exam assembly tasks. |
| Approach: | They propose a mixed-integer programming re-ranking approach to improve relevance while mitigating bias on an industry-grade exam assembly platform. |
| Outcome: | The proposed approach improves relevance and reduces bias by 17% when compared to other methods on a real-world exam assembly platform. |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have revolutionized NLP but amplify linguistic inequities in multilingual applications. |
| Approach: | They evaluate pretrained language models including Mistral, TowerInstruct, OpenHathi, Tamil-Llama, and Kan-Lama across eight languages spanning high-resource and low-resourced settings. |
| Outcome: | The proposed models fail to bridge linguistic divides and are inefficient when compared to other models. |
Copied to clipboard
| Challenge: | Existing datasets lack size and diversity, with only 609 voice phishing samples available in Korean and 638 smishing instances available in English. |
| Approach: | They propose to use a Korean dataset to construct a reliable phishing detection system using language models to evaluate the model's in-domain and unseen attack detection performance. |
| Outcome: | The proposed system performs reasonably well in voice and unseen attacks while smishing detection remains challenging. |
Copied to clipboard
| Challenge: | Manual assignment of ATC codes to prescription records is a significant bottleneck in healthcare operations. |
| Approach: | They propose a method to automate the assignment of ATC codes to prescription records . they use locally deployable large language models to guide LLMs through the ontology . |
| Outcome: | The proposed method achieves 78% exact match accuracy with GPT-4o and 60% with Llama 3.1 70B. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have impacted the writing process, enhancing productivity by collaborating with humans in content creation platforms. |
| Approach: | They propose a framework that uses explicit outlines to guide LLMs in generating goal-oriented, high-quality text. |
| Outcome: | The proposed approach significantly improves text quality according to evaluations by LLMs and professional writers. |
Copied to clipboard
| Challenge: | Existing adversarial examples generate invalid or ambiguous examples that fool the systems into wrong detection. |
| Approach: | They propose an annotation pipeline for quality control of generated toxic adversarial examples (TAE) they use model-based automated annotation and human-based quality verification to assess quality requirements of a TAE dataset. |
| Outcome: | The proposed pipeline can transfer-attack SOTA toxicity content moderation models and services with adversarial training. |
Copied to clipboard
| Challenge: | Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models . however, the leaderboard has faced significant limitations over time due to its academic nature . |
| Approach: | They propose an improved version of the Open Ko-LLM Leaderboard to improve benchmarking . original benchmarks replaced with new tasks that align with real-world capabilities . four new native Korean benchmarks are introduced to better reflect distinct characteristics of Korean language . |
| Outcome: | The proposed framework improves the Open Ko-LLM Leaderboard2 benchmark suite. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved significant success in open-domain question answering, however, they continue to face challenges such as knowledge cutoffs and hallucinations. |
| Approach: | They propose a new mechanism that integrates a curiosity-driven reasoning mechanism into an LLM agent to generate relevant follow-up questions. |
| Outcome: | The proposed enhancement integrates a curiosity-driven reasoning mechanism into an LLM agent, enabling it to generate relevant follow-up questions, thereby guiding the information retrieval process more efficiently. |
Copied to clipboard
| Challenge: | Maintaining consistent character personas remains a significant challenge due to variability in information extraction. |
| Approach: | They propose a framework to dynamically reconstruct character personas through Character Persona Training. |
| Outcome: | The proposed framework is evaluated through Big Five personality evaluations and creative tasks, in which characters generate original narratives. |
Copied to clipboard
| Challenge: | Open-source large language models (LLMs) are a promising tool for low-resource languages . however, there is still a substantial performance gap between high-resourced languages and LRLs . |
| Approach: | They develop an algorithm to select a subset of texts from a larger corpus and use it to select tokens for LLMs. |
| Outcome: | The proposed algorithm reduces the cost of continual pre-training (CPT) with large amounts of language-specific data. |
Copied to clipboard
| Challenge: | Current intent detection work experiments with minor intent categories. |
| Approach: | They propose a retrieval-augmented generation framework that uses query-to-query and query- to-metadata approaches to retrieve intents from metadata. |
| Outcome: | The proposed framework improves on query-to-query (Q2Q) and query- to-metadata (Q 2M) approaches. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are pivotal for advanced text processing and generation. |
| Approach: | They propose a framework to train on-device Large Language Models optimized for invoking software APIs. |
| Outcome: | The proposed model outperforms GPT-4 in API calling tasks while maintaining inference speed. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are characterized by their immense size, often consisting of at least one billion parameters. |
| Approach: | They propose a mixture of Frozen Experts architecture that integrates PEFT and MoE to enhance both training efficiency and model scalability. |
| Outcome: | The proposed architecture outperforms other methods while achieving the highest efficiency. |
Copied to clipboard
| Challenge: | Existing methods for financial breakout detection are subpar, despite large data and knowledge required. |
| Approach: | They propose a financial breakout dataset and introduce FinLLM-B, a large language model for financial breakout detection. |
| Outcome: | The proposed model outperforms GPT-3.5 in the field of financial breakout detection. |
Copied to clipboard
| Challenge: | Unrestricted access to external Large Language Models (LLMs) based services like ChatGPT and Gemini can lead to data leakages, especially for large enterprises providing products and services that require confidentiality guarantees. |
| Approach: | They propose a platform that enterprises can use to query external Large Language Models without leaking confidential internal and client information. |
| Outcome: | The proposed solution minimizes data leakage while limiting impact to semantics while preserving the accuracy of the model candidates. |
Copied to clipboard
| Challenge: | despite advances in multilingual machine translation, lack of well-crafted AD data impedes development of audio description translation systems. |
| Approach: | They propose an audio description translation system for three main Swiss languages and English . they combine human expertise with the power of Large Language Models to improve quality . |
| Outcome: | The proposed system is designed to enhance accessibility for multilingual populations in Switzerland. |
Copied to clipboard
| Challenge: | Live morph resolution task is used to detect e-commerce live streaming violations . morphs are used to evade scrutiny and engage in false advertising . |
| Approach: | They propose a task to detect morph violations in live streaming scenarios . they use large language models to generate additional training data . |
| Outcome: | The proposed method improves performance and improves live streaming regulation. |
Copied to clipboard
| Challenge: | Data scarcity is one of the main problems when it comes to real-world applications of transformer-based models. |
| Approach: | They propose a method to source annotated German monologues from existing monologue material to train TOD systems. |
| Outcome: | The proposed model can be used to train TOD systems on a real-world example of a travel booking service. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been used in clinical decision support, medical education and patient communication. |
| Approach: | They propose a benchmark to evaluate large language models in the domain of medical ethics and assess their grasp of medical ethical principles and their application across diverse scenarios. |
| Outcome: | The proposed framework assesses the models’ grasp of medical ethics principles and their ability to apply them across diverse scenarios. |
Copied to clipboard
| Challenge: | Traditional approaches to predicting the duration of a patient's stay in an Intensive Care Unit (ICU) rely on structured clinical data, but recent advances in language models offer significant potential to utilize unstructured text data for ICU length-of-stay (LoS) predictions. |
| Approach: | They propose a method for analyzing nursing notes to predict ICU length-of-stay of patients. |
| Outcome: | The proposed model outperforms baseline models on the MIMIC-III dataset and shows that it significantly outperformed existing models. |
Copied to clipboard
| Challenge: | RevieWeaver extracts key product features and provides concise review summaries . a condensed list of key features, pros, and cons, along with a brief summary of customer opinions can help mitigate this issue. |
| Approach: | They propose a framework that extracts key product features and provides concise review summaries. |
| Outcome: | The proposed framework scales efficiently to 30 million reviews and ensures reproducibility and controllability. |
Copied to clipboard
| Challenge: | Medical coding is time-consuming and error-prone due to large label space, lengthy text inputs, and the absence of supporting evidence annotations. |
| Approach: | They propose a Generative AI framework for automatic medical coding that leverages extraction, retrieval, and re-ranking techniques as core components. |
| Outcome: | The proposed framework outperforms existing methods on the International Classification of Diseases (ICD) code prediction scale. |
Copied to clipboard
| Challenge: | Existing zero-shot product attribute value extraction approaches require sellers to manually provide product descriptions. |
| Approach: | They propose a cross-modal zero-shot attribute value generation framework based on CLIP that uses product images as inputs for zero- shot inference. |
| Outcome: | The proposed framework significantly outperforms other vision-language models for zero-shot attribute value extraction. |
Copied to clipboard
| Challenge: | Typical evaluations of Large Language Models (LLMs) report a single accuracy metric per dataset, often derived from an optimized setup. |
| Approach: | They propose a framework for non-adversarial evaluation of large language models that evaluates models by repeatedly testing them on the same benchmarks in various setups. |
| Outcome: | The proposed framework evaluates models by repeatedly testing them on the same benchmarks in various setups to give a realistic estimate of their accuracy and consistency. |
Copied to clipboard
| Challenge: | Existing benchmarks lack domain-specific datasets for evaluating large language models . existing benchmarks often lack domain specific datasets, which can be difficult to convert to standardized metrics or regulatory issues. |
| Approach: | They propose to use 25 publicly available domain-specific English benchmarks from diverse domains . they propose to combine a wide range of natural language processing tasks for holistic evaluation . |
| Outcome: | The proposed framework includes 25 publicly available domain-specific English benchmarks from diverse enterprise domains like financial services, legal, climate, cyber security, and 2 public Japanese finance benchmarks. |
Copied to clipboard
| Challenge: | Large language models require considerable computing resources, which can be costly and often unavailable. |
| Approach: | They propose to integrate 4-bit Post-training Quantization with QLoRA to address these issues . they demonstrate that the integration outperforms standard quantization and fine-tuning . |
| Outcome: | The proposed integration outperforms standard PTQ and 16-bit full-parameter fine-tuning on LLMs. |
Copied to clipboard
| Challenge: | Application of Large Language Models to complex causal question answering can be stymied by their opacity and propensity for hallucination. |
| Approach: | They propose a causal QA approach that combines iterative RAG with a formal model of causation. |
| Outcome: | The proposed approach is implemented into a Collaborative Research Assistant (Cora) and evaluated in the life sciences domain. |
Copied to clipboard
| Challenge: | Existing methods for jailbreaking large-language models are limited by their limitations . authors present a mutation-based fuzzing technique that generates effective jailbreaking templates . |
| Approach: | They propose a mutation-based fuzzing technique for efficiently finding effective jailbreaking templates that combine with harmful questions to generate harmful responses. |
| Outcome: | The proposed technique achieves 95% attack success rates on public datasets for leading LLMs . it also shows impressive generalizability to unseen harmful questions and improves model defenses to prompt attacks. |
Copied to clipboard
| Challenge: | Experimental results show a 66.53% reduction in parameter size within the attention block and competitive accuracy improvements of 3.55% and 0.89% over symmetric and pairwise attention-based models, respectively. |
| Approach: | They propose a simplified approach where a single weight matrix is used for Keys, Queries, and Values instead of separate matrices for each. |
| Outcome: | The proposed approach outperforms the BERT baseline on GLUE tasks even outperforming the standard BERT model in handling noisy and out-of-domain data. |
Copied to clipboard
| Challenge: | Existing methods that only deal with flat text chunks use a graph structure to handle complex questions. |
| Approach: | They propose layout-aware graph modeling for multimodal RAG using document layout parsing to take into account relationship of multimodalities. |
| Outcome: | The proposed method can handle complex questions that require information from multimodalities. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being used for communication tasks across different regions. |
| Approach: | They propose a benchmark to evaluate whether Large Language Models are ethically aligned and can be used in real-world situations. |
| Outcome: | The proposed benchmark evaluates whether LLMs comply with or resist swearing instructions and assesses their alignment with ethical frameworks, cultural nuances, and language comprehension capabilities. |
Copied to clipboard
| Challenge: | Recent advances in NLP have the potential to transform HR processes, from recruitment to employee management. |
| Approach: | They analyze key tasks such as information extraction and text classification and their roles in downstream applications like recommendation and language generation while discussing ethical concerns. |
| Outcome: | The proposed frameworks can be applied to HR tasks and to recommendation, language generation, and interaction. |
Copied to clipboard
| Challenge: | generative AI models can extract accurate facts from external unstructureddata sources. |
| Approach: | They propose to use data embeddings to enhance RAG for retrieving facts from structured data sources, achieving low latency and high reliability. |
| Outcome: | The proposed model achieves low latency and high reliability and average response time of 7.33 seconds. |
Copied to clipboard
| Challenge: | a suite of advanced models is designed to detect and mitigate risks associated with prompts and responses. |
| Approach: | a team of researchers develop a model family to detect and mitigate risks associated with prompts and responses. the model family is based on the Granite 3.0 language models. |
| Outcome: | a new model family is designed to detect and mitigate risks associated with prompts and responses. |
Copied to clipboard
| Challenge: | Streaming automatic speech recognition models use high power consumption to improve usability and accuracy. |
| Approach: | They propose to optimize on-device speech recognition models by adjusting component energy sensitivities based on their specific energy sensitities to reduce power consumption. |
| Outcome: | The proposed approach achieves up to 47% lower energy usage while preserving comparable model accuracy and improving real-time performance compared to leading methods. |
Copied to clipboard
| Challenge: | Traditional methods of localization focus on linguistic conversion, but content needs to align with the target audience’s cultural norms, nuances, and technical requirements. |
| Approach: | They propose a new approach for Large Languages Models (LLMs) called Break-Ideate-Generate (BrIdGe) that breaks the source content into granular facts, organizes the granules and executes the plan to ‘generate’ localized content. |
| Outcome: | The proposed approach 'breaks' the source content into granular facts, ‘ideates’ an action plan for content creation in the target language by organizing the granules, and finally executes the plan to ‘generate’ localized content. |
Copied to clipboard
| Challenge: | Concept Distillation (CD) is an automated prompt optimization technique for enhancing weaker models on complex tasks. |
| Approach: | They propose an automatic prompt optimization technique for enhancing weaker models on complex tasks using a base prompt and a strong model to generate reasons for these mistakes. |
| Outcome: | The proposed technique improves weaker models on NL2Code and mathematical reasoning tasks, while preserving performance. |
Copied to clipboard
| Challenge: | Existing benchmarks for general-purpose RAG systems, such as CRAG, RGB, MultiHop-RAG, and CRUD-RAGG, are limited and lack a benchmark specifically tailored to evaluate frameworks. |
| Approach: | They evaluated OpenAI’s Assistants API versus a RAG assistant built with Langchain and deployed a system based on benchmark insights as a course assistant over a two-year span. |
| Outcome: | The proposed benchmarks show that domain-specific retrieval impacts response accuracy and highlight key challenges in real-world deployment. |
Copied to clipboard
| Challenge: | Identifying query variants is nontrivial as highly similar query pairs may fail to differ in word form, order, or phrasing despite sharing the same intent. |
| Approach: | They propose to use retrieval as an environment feedback to capture semantic equivalence . experimental results demonstrate the efficacy of the proposed method . |
| Outcome: | The proposed method improves query variant detection across diverse scenarios. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have potential to automate hiring but inherent biases may lead to unfair hiring practices. |
| Approach: | They evaluate how factors such as gender, race, and educational background influence model decisions. |
| Outcome: | The proposed model reduces biases related to gender and race, but implicit biase concerning educational background remains significant. |
Copied to clipboard
| Challenge: | Unlike existing tools, our system addresses the ambiguity of vague, multi-line queries, setting a new benchmark in data storytelling by tackling complexities no existing system comprehensively handles. |
| Approach: | They propose a system that processes and interprets vague, open-ended, and multi-line complex queries, transforming them into coherent, actionable data stories. |
| Outcome: | The proposed system processes and interprets vague, open-ended, and multi-line complex queries, transforming them into coherent, actionable data stories. |
Copied to clipboard
| Challenge: | general-purpose vision-language models struggle to understand and converse about real-world e-commerce product images. |
| Approach: | a new approach is proposed to use large-scale image-text pairs to train a generative VLM for e-commerce product images. |
| Outcome: | The proposed model outperforms general-purpose VLMs on multiple vision tasks in the e-commerce domain. |
Copied to clipboard
| Challenge: | Effective customer support requires domain-specific solutions tailored to users’ issues. |
| Approach: | They propose an automated pipeline for building a domain-specific KB with a hierarchical tree structure that maps user issues to precise and domain-compliant solutions. |
| Outcome: | Experiments in troubleshooting and medical domains show that the proposed pipeline outperforms LLMs and unstructured knowledge bases and is 75 times more cost-effective than manual methods. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. |
| Approach: | They present a spoken NER dataset in the medical domain using pre-trained models that are encoder-only and sequence-to-sequence. |
| Outcome: | The dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types. |
Copied to clipboard
| Challenge: | PLEX is a lottery-ticket based parameter-efficient fine-tuning method that adapts large language models to well-supported and underrepresented programming languages (PLs) in pretraining. |
| Approach: | They propose a lottery-ticket based parameter-efficient fine-tuning method that adapts large language models to well-supported and underrepresented programming languages (PLs) |
| Outcome: | The proposed method achieves state-of-the-art performance among PEFT methods while maintaining competitive results with reduced computational overhead. |
Copied to clipboard
| Challenge: | a conversational AI system that answers standard airport queries and resolves airport terminology is ideal for airports from the top 20 in terms of annual passenger numbers. |
| Approach: | They propose a Conversational AI system that enables staff to communicate with flight systems . the system answers standard airport queries and resolves airport terminology . |
| Outcome: | The proposed system answers standard airport queries and resolves airport terminology, jargon, abbreviations and dynamic questions involving reasoning. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly impacting children through education, toys, and therapy, offering benefits like improved mental health and parental controls. |
| Approach: | They propose a comprehensive approach to evaluating LLM safety specifically for children by listing potential risks that children may encounter when using LLM-powered applications. |
| Outcome: | The proposed model bridges the gap in child safety literature across various fields. |
Copied to clipboard
| Challenge: | paper prescriptions are difficult for customers to interpret and are often unstructured, handwritten, and illegible. |
| Approach: | They propose a multi-step Large Language Model-based solution for automated pharmacy cart construction. |
| Outcome: | The proposed solution can yield up to 19% - 40% and 11% - 26% increase in Recall@3 relative to SOTA methods. |
Copied to clipboard
| Challenge: | Domain- and customer-specific requirements complicate the problem of NL2SQL customization. |
| Approach: | They propose a distilled customization framework tailored for NL2SQL tasks. |
| Outcome: | The proposed framework outperforms teacher models on three benchmarks and achieves an average improvement of 36% in execution accuracy. |
Copied to clipboard
| Challenge: | eC-Tab2Text dataset is designed to capture product attributes and user-specific queries. |
| Approach: | They propose a novel dataset to capture the intricacies of e-commerce including detailed product attributes and user-specific queries. |
| Outcome: | The proposed dataset outperforms existing generalpurpose LLMs in generating accurate product reviews. |
Copied to clipboard
| Challenge: | Existing benchmarks assess LLMs' chat abilities in multi-turn dialogues or their use of retrieval for augmented responses in limited tasks such as knowledge QA or numeric reasoning. |
| Approach: | They propose a benchmark to evaluate LLMs' capabilities in multi-turn dialogues following retrievals. |
| Outcome: | The proposed benchmark evaluates LLMs' ability to perform in multi-turn dialogues following retrievals over 6 representative scenarios. |
Copied to clipboard
| Challenge: | Current manual approaches to analyzing overlapping or conflicting content are time-consuming, costly, and error-prone. |
| Approach: | They propose a large language model that uses a construction domain-adapted large language for the semantic comparison of sentences in construction standards. |
| Outcome: | The proposed framework achieves 97.9% accuracy and 0.907 macro F1-score in classifying sentences from Korean construction standards as overlapping, conflicting, or neutral. |
Copied to clipboard
| Challenge: | Proteins play critical roles in biological systems, yet 99.7% of 227 million known protein sequences remain uncharacterized due to the limitations of experimental methods. |
| Approach: | They propose a multimodal large language model that interprets protein sequences and generates informative text to address open-ended questions about protein functions and attributes. |
| Outcome: | The proposed model outperforms existing models in open-ended question-answering tasks. |
Copied to clipboard
| Challenge: | Using the entire dataset, shuffling answer options introduces instability in the insurance and finance sectors. |
| Approach: | They propose a dataset for evaluation of performance in vocational and professional certification exams in Indonesia. |
| Outcome: | The proposed dataset includes 8,834 multiple-choice questions from 27 large language models across six key sectors. |
Copied to clipboard
| Challenge: | Tabular datasets in industrial settings often encompass extensive data with numerous rows and columns. |
| Approach: | They propose a system that leverages large language models to generate executable code for data wrangling tasks . they identify inherent patterns in the data while leveraging external knowledge . |
| Outcome: | The proposed system detects patterns in the data while leveraging external knowledge . it generates executable code for data-wrangling tasks like missing value imputation and error correction . |
Copied to clipboard
| Challenge: | Existing persona-consistent dialogue models lack robustness due to limited scale and diversity of datasets. |
| Approach: | They propose an open-domain persona dialogue system that employs extensive generative pre-training on a persona dialog dataset to enhance persona consistency. |
| Outcome: | The proposed model generates vast persona dialogue datasets and addresses invalid persona bias. |
Copied to clipboard
| Challenge: | Hallucination is a problem in large language models that produce incorrect output . authors propose a reliable and high-speed production system to detect and rectify hallucinations . |
| Approach: | They propose a high-speed production system that detects hallucinations in LLMs . they propose NER, natural language inference, span-based detection and a rewriting mechanism . |
| Outcome: | The proposed system detects a wide range of hallucinations in LLM responses. |
Copied to clipboard
| Challenge: | News aggregators provide comprehensive and timely news stories that are sourced from diverse sources but differ in phrasing, formatting or supplemented with additional details. |
| Approach: | They propose a method that combines embeddings from pretrained language model and latent metadata of a news article followed by community detection to identify clusters of near-duplicates. |
| Outcome: | The proposed approach can detect nuanced similarities and differences in news snippets using pretrained language model and latent metadata of a news article followed by community detection. |
Copied to clipboard
| Challenge: | Sustainable speech recognition systems are essential for scientists, journalists, and anyone processing audio recordings of interviews and meetings. |
| Approach: | They propose a speech-to-text system "Pisets" which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. |
| Outcome: | The proposed system ensures robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. |
Copied to clipboard
| Challenge: | Relevance modeling between queries and items is a key component of commercial search engines. |
| Approach: | They propose a framework for continual pre-training of LLMs to enhance domain knowledge . they employ queries and multi-field item to jointly pre-train for enhancing domain knowledge. |
| Outcome: | The proposed model achieves convincing performance compared to strong baselines. |
Copied to clipboard
| Challenge: | GraphQL is a flexible alternative to REST APIs, but generating complex queries remains challenging. |
| Approach: | They propose a framework that integrates GraphQL schemas with natural language inputs to improve query generation accuracy. |
| Outcome: | The proposed framework improves performance on a publicly available complex GraphQL dataset. |
Copied to clipboard
| Challenge: | Existing benchmark aggregation methods, such as Elo-based systems, can be resource-intensive, public facing, and time-consuming. |
| Approach: | They propose a framework for aggregating performance across diverse benchmarks that generates a “Goodness” and a ‘Fastness” score. |
| Outcome: | The proposed framework achieves higher Pearson correlation with Chatbot Arena Elo scores than MMLU’s correlation with chatbot Arena scores, validating its reliability for real-world LLM evaluation. |
Copied to clipboard
| Challenge: | Recent literature focuses on constructing large audio language models (LALMs) but they are limited in temporal reasoning, which may hinder commercial applications . |
| Approach: | They propose a data augmentation technique for generating reliable audio temporal questions and answers using an LLM. |
| Outcome: | The proposed model performs well on public audio benchmark datasets and is optimized for edge applications. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face limitations due to outdated knowledge, hallucinations, and poor reasoning in complex contexts. |
| Approach: | They propose a Hybrid Parameter-Adaptive RAG system for the AI legal domain with NYC Local Law 144 as the test case. |
| Outcome: | The proposed system improves retrieval accuracy, response fidelity, and contextual precision on NYC Local Law 144 . Empirical evidence indicates that many AI tools overstate their ability to prevent hallucinations in legal and policy contexts. |
Copied to clipboard
| Challenge: | a recent study has demonstrated that context-dependent memory encoding can help to retrieve key memory cues essential for problem-solving. |
| Approach: | They propose an efficient architecture miming human memory processes through multistage encoding, context-aware storage, and retrieval strategies for LLM-centric agents. |
| Outcome: | The proposed architecture surpasses state-of-the-art online LLM-centric approaches on two interactive decision-making benchmarks in the navigation and manipulation domain. |