Papers with GPT-4o-mini

27 papers
Copyright Detective: A Forensic System to Evidence LLMs Flickering Copyright Leakage Risks (2026.acl-demo)

Copied to clipboard

Challenge: **Copyright Detective** is the first interactive forensic system for detecting, analyzing, and visualizing potential copyright risks in LLM outputs.
Approach: They propose a system that detects copyright infringements and visualizes them . they use content recall testing, paraphrase-level similarity analysis and persuasive jailbreak probing .
Outcome: The proposed system detects, analyzes, and visualizes potential copyright risks in LLM outputs.
Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence (2026.eacl-industry)

Copied to clipboard

Challenge: Existing benchmarks focus on tool usage or task completion, overlooking an agent’s capacity to adhere to multi-step policies, navigate task dependencies, and remain robust to unpredictable user or environment behavior.
Approach: They propose a benchmark to assess policy-aware agents in customer support using a dynamic-prompt agent and a static-promped agent that explicitly models policy control.
Outcome: The proposed benchmark assesses agent's ability to adhere to multi-step policies, navigate task dependencies, and remain robust to unpredictable user or environment behavior.
Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource Scenarios (2026.eacl-short)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems for low-resource languages produce erroneous transcripts due to limited annotated data and linguistic complexity.
Approach: They compare language models and large language models for post-ASR correction in Hindi . they observe a scaling trend under zero-shot ICL where mid-sized LLMs degrade performance before marginal recovery at extreme scales.
Outcome: The proposed model outperforms larger models in both fine-tuning and in-context learning settings.
An Efficient Context-Dependent Memory Framework for LLM-Centric Agents (2025.naacl-industry)

Copied to clipboard

Challenge: a recent study has demonstrated that context-dependent memory encoding can help to retrieve key memory cues essential for problem-solving.
Approach: They propose an efficient architecture miming human memory processes through multistage encoding, context-aware storage, and retrieval strategies for LLM-centric agents.
Outcome: The proposed architecture surpasses state-of-the-art online LLM-centric approaches on two interactive decision-making benchmarks in the navigation and manipulation domain.
COMPKE: Complex Question Answering under Knowledge Editing (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for knowledge editing do not accurately evaluate how well models apply knowledge in real-life situations.
Approach: They propose a benchmark to evaluate how well updated models apply new knowledge in real-life situations.
Outcome: The proposed method achieves 39.47 accuracy on GPT-4o-mini but drops significantly to 3.83 on Qwen2.5-3B.
Semi-supervised Fine-tuning for Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing LLMs require labeled data, which can be costly in real-world applications.
Approach: They propose a framework that can fully exploit labeled and unlabeled data for LLM fine-tuning . they conducted experiments using GPT-4o-mini and Llama-3.1 on seven general or domain-specific datasets .
Outcome: The proposed framework can fully exploit labeled and unlabeled data for LLM alignment from a propagate-and-select manner.
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing instruction-tuning datasets lack execution-grounded supervision and offer limited support for iterative code correction.
Approach: They propose a large-scale instruction tuning dataset for Python-based visualization and self-correction.
Outcome: The proposed dataset outperforms strong open-source baselines and proprietary models like GPT-4o-mini.
Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy (2025.findings-emnlp)

Copied to clipboard

Challenge: Strong attribute control can distort meaning, while prioritizing semantic preservation may weaken attribute alignment.
Approach: They propose a method that restricts accepted samples to text meeting a minimum BERTScore threshold and applies gradient-assisted proposal generation to improve attribute alignment.
Outcome: a new method for counterfactual text generation improves attribute alignment and semantic preservation . the proposed method achieved the best macro F1-score in two of three test sets .
Par-ITA: Benchmarking Seq2Seq and LLMs on a Human-Supervised Parallel Corpus for Italian Hyperpartisan Neutralization (2026.acl-long)

Copied to clipboard

Challenge: a new study examines the role of hyperpartisan content in online polarization in the social web.
Approach: They propose a human-supervised parallel corpus for italian hyperpartisan neutralization of 2,475 paragraph pairs.
Outcome: The proposed dataset is the first human-supervised parallel corpus for italian hyperpartisan neutralization of 2,475 paragraph pairs.
MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods to align large language models with multilingual encoders raise accuracy for low-resource languages (LRLs) but performance of LLMs in low- and high-resourced languages remains a problem.
Approach: They propose a model-stacking framework that iteratively refines in 2-stages based on a curriculum strategy and adapts only a small set of DoRA weights.
Outcome: The proposed framework improves exact-match accuracy by +12.9 pp over MindMerger and outperforms GPT-4o-mini by 15.2 pp on the AfriMGSM benchmark.
Stronger Universal and Transferable Attacks by Suppressing Refusals (2025.naacl-long)

Copied to clipboard

Challenge: Efforts have focused on aligning models to human preferences (RLHF) . yet, it is believed that such optimization-based attacks are sample-specific.
Approach: They propose an algorithm to embed a "safety feature" into models to make them safe for mass deployment.
Outcome: The proposed attack achieves 25% success rate against the state-of-the-art Circuit Breaker defense, compared to 2.5% by white-box GCG.
FinMaster: A Holistic Benchmark for Full-Pipeline Financial Management with Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks lack domain-specific data, realistic workflow-level task design, and standardized workflow- level evaluation.
Approach: a new benchmark evaluates large language models on financial management workflows . the global financial services market is projected to grow to $37 trillion by 2027 .
Outcome: a new benchmark for large language models on financial management workflows reveals critical capability gaps . accuracy drops from 90% on basic tasks to 40% on complex scenarios requiring multi-step reasoning . the global financial services market reached $25.8 trillion in 2022 and is projected to grow to $37 trillion by 2027 .
SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are increasingly deployed in multi-turn settings such as tutoring, support, and counseling where reliability depends on preserving consistent roles, personas, and goals across long horizons.
Approach: They propose a framework that decomposes LLM–LLM conversations into a modular, stability-first framework that allows for a stable persona-driven agent simulation for multi-turn dialogue generation.
Outcome: The proposed framework decomposes the LLM-based model into four main components: persona creation, plausibility validation, and natural-language persona crafting.
LITERA: An LLM Based Approach to Latin-to-English Translation (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown promise in addressing these challenges across languages like Latin.
Approach: They propose a Latin-to-English translation platform based on GPT-4o and GPT4o that combines a fine-tuned version of GPT-3o and a sophisticated algorithm to produce literal translations.
Outcome: The model is based on two languages: Latin Interpretation and Translations into English for Research Assistance and GPT-4o.
LCFO: Long Context and Long Form Output Dataset and Benchmarking (2025.findings-acl)

Copied to clipboard

Challenge: Using long text outputs to evaluate progress in summarization and summary expansion tasks is challenging.
Approach: They propose a framework for assessing gradual summarization and summary expansion capabilities across diverse domains.
Outcome: The proposed framework provides alignments between specific QA pairs and corresponding summaries in 7 domains.
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: We argue that knowledge-retrieval and reasoning tasks are not ideal for measuring generalization, as LLMs are not trained for specific tasks.
Approach: They propose a statistically motivated framework using personalization to assess generalization in Large Language Models.
Outcome: The proposed framework outperforms existing models on movie and music recommendation datasets, but all models have room for improvement, especially Llama.
Automatic Mathematic In-Context Example Generation for LLM Using Multi-Modal Consistency (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for in-context learning require annotated datasets, resulting in higher computational costs and lower quality examples.
Approach: They propose a framework that automatically generates high-quality in-context examples to enhance LLMs’ mathematical reasoning.
Outcome: Evaluated on four math problem datasets, the proposed framework outperforms baseline methods with LLM accuracy ranging from 87.0% to 99.3%.
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond.
Approach: They propose a synthetic dataset in the financial domain that integrates Chain-of-Thought reasoning into LLMs in a supervised manner to facilitate effective long-context understanding.
Outcome: The proposed model outperforms standard GPT-4o-mini on the Loong benchmark and fine tunes LLaMA-3.1-8B-Instruct on the model, achieving a 28.0% gain on the financial subset.
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly deployed via third-party system prompts downloaded from public marketplaces.
Approach: They propose a framework that optimizes system prompts to trigger LLMs to output compromised responses only for specific queries.
Outcome: The proposed framework achieves up to 70% F1 reduction on targeted queries with minimal degradation to general capabilities.
Profiling LLM’s Copyright Infringement Risks under Adversarial Persuasive Prompting (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have demonstrated impressive capabilities in text generation but raise concerns regarding potential copyright infringement.
Approach: They propose a structured persuasion workflow to analyze the influence of persuasive prompts on LLM outputs.
Outcome: The proposed method analyzes the influence of persuasive prompts on LLM outputs.
RJE: A Retrieval-Judgment-Exploration Framework for Efficient Knowledge Graph Question Answering with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Knowledge graph question answering (KGQA) aims to answer natural language questions using knowledge graphs.
Approach: They propose a framework that retrieves refined reasoning paths and evaluates their sufficiency.
Outcome: The proposed framework outperforms existing baselines while enabling small open-source LLMs to achieve competitive results without fine-tuning LLM.
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks assess LLM performance in single-course settings and lack systematic evaluation in multi-course scenarios, where a patient’s condition evolves over time.
Approach: They propose to use large language models to assess their performance in multi-course clinical decision-making scenarios where a patient’s condition evolves over time.
Outcome: The proposed model includes 1,275 Chinese and 5,804 English samples across four stages from admission to discharge.
Tree-of-Prompts: Abstracting Control-Flow for Prompt Optimization (2025.findings-acl)

Copied to clipboard

Challenge: Existing prompt optimization methods struggle with disjoint cases in complex tasks.
Approach: They propose a tree-of-prompts structure which expands child prompts from parent prompts . they propose to use a nested if-else structure to address varying similarities and complexities .
Outcome: The proposed tree-of-prompts outperforms PromptAgent and MoP on Gorilla, MATH and subset of BBH benchmarks.
Boosting LLM’s Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Molecular structure elucidation involves deducing a molecule’s structure from various types of spectral data, which is crucial in chemical experimental analysis.
Approach: They propose a Knowledge-enhanced reasoning framework for Molecular Structure Elucidation that leverages Monte Carlo Tree Search for test-time scaling as a plugin to extend the LLMs’ coverage of the chemical structure space.
Outcome: The proposed framework significantly improves on both GPT-4o-mini and GPT4o, and a specialized molecule-spectrum scorer improves performance.
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text (2025.emnlp-main)

Copied to clipboard

Challenge: Existing VLMs produce more positive descriptions for high-income countries compared to middle- or low-income nations, even when country attribution is the only variable changed.
Approach: They propose to automate the process by generating textual summaries of charts using vision-language models to understand how a country’s economic status influences the sentiment of generated summary.
Outcome: The proposed model amplifys geo-economic biases in 6,000 chart-country pairs from six widely used vision-language models to understand how a country’s economic status influences the sentiment of generated summaries.
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing Privacy Policy Question Answering systems exhibit performance disparities across English dialects, disadvantaging speakers of non-standard varieties.
Approach: They propose a framework that integrates a Dialect Agent and a Privacy Policy Agent to mitigate dialectal biases.
Outcome: The proposed framework improves GPT-4o-mini’s zero-shot accuracy from 0.394 to 0.601 on PrivacyQA and 0.352 to 0.464 on PolicyQA.
Vulnerability of LLMs’ Stated Belief? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly employed in question-answering tasks.
Approach: They analyze how different persuasive strategies influence stated belief stability . they also examine whether verbalized confidence prompting increases vulnerability .
Outcome: The proposed model exhibits extreme compliance, with 82.5% of belief changes occurring at the first persuasive turn.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations