Papers with GPT-3.5-turbo
Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing work on LLMs that only enhance reasoning abilities, but which lack factual hallucination and slow-thinking capabilities, argues that SPP is a cognitive synergist. |
| Approach: | They propose a Solo Performance Prompting (SPP) that transforms a single LLM into a cognitive synergist by engaging in multi-turn self-collaboration with multiple personas. |
| Outcome: | The proposed model reduces factual hallucination and maintains strong reasoning abilities on three challenging tasks . |
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)
Copied to clipboard
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. |
| Approach: | They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets. |
| Outcome: | The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions. |
Towards Better Graph-based Cross-document Relation Extraction via Non-bridge Entity Enhancement and Prediction Debiasing (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on relation extraction ignore non-bridge entities, leading to bias during inference. |
| Approach: | They propose a graph-based cross-document Relation Extraction model with non-bridge entity enhancement and prediction debiasing that integrates non-cross entities with target entities and bridge entities. |
| Outcome: | The proposed model outperforms baseline models on open and closed datasets. |
Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models (2024.naacl-long)
Copied to clipboard
Yi Luo, Zhenghao Lin, YuHao Zhang, Jiashuo Sun, Chen Lin, Chengjin Xu, Xiangdong Su, Yelong Shen, Jian Guo, Yeyun Gong
| Challenge: | Guide-Align is a guideline-oriented approach to augment the safety and quality of Large Language Models. |
| Approach: | They propose a guideline-oriented method to augment the safety and quality of large language models. |
| Outcome: | The proposed method outperforms existing methods on three benchmarks and shows significant improvements in security and quality. |
Selene: Pioneering Automated Proof in Software Verification (2024.acl-long)
Copied to clipboard
| Challenge: | Currently, software verification is resource-intensive and manpower-consuming. |
| Approach: | They propose a project-level automated proof benchmark based on the seL4 operating system . they propose augmentations to enhance the flexibility of the framework and lightweight verification environment . |
| Outcome: | The proposed framework provides a comprehensive framework for end-to-end proof generation and a lightweight verification environment. |
R3 Prompting: Review, Rephrase and Resolve for Chain-of-Thought Reasoning in Large Language Models under Noisy Context (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have evaluated LLMs under noise-free context but the dilemma for LLM to produce inaccurate results under noisy context has not been fully investigated. |
| Approach: | They propose a new method for CoT reasoning using Chain-of-Thought prompting that interacts with LLMs to perform key sentence extraction, variable declaration and answer prediction. |
| Outcome: | The proposed method outperforms existing CoT prompting methods on five reasoning tasks under noisy context. |
Training Language Models to Generate Text with Citations via Fine-grained Rewards (2024.acl-long)
Copied to clipboard
| Challenge: | Recent Large Language Models (LLMs) are prone to hallucination and their outputs often contain incorrect or unverifiable claims. |
| Approach: | They propose a training framework using fine-grained rewards to teach LLMs to generate highly supportive and relevant citations while ensuring the correctness of their responses. |
| Outcome: | The proposed training framework outperforms existing methods on QA datasets and surpasses GPT-3.5-turbo on LLaMA-2-7B. |
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT (2024.lrec-main)
Copied to clipboard
Amirhossein Abaskohi, Sara Baruni, Mostafa Masoudi, Nesa Abbasi, Mohammad Hadi Babalou, Ali Edalat, Sepehr Kamahi, Samin Mahdizadeh Sani, Nikoo Naghavian, Danial Namazifard, Pouya Sadeghi, Yadollah Yaghoobzadeh
| Challenge: | a new study examines the efficacy of large language models (LLMs) for Persian . ChatGPT and LLMs have shown remarkable performance in English, but their efficiency for low-resource languages remains an open question. |
| Approach: | They present a benchmarking study of large language models (LLMs) for Persian . they focus on GPT-3.5-turbo, but also GPT-4 and OpenChat-3.5 . |
| Outcome: | The proposed model performs better in Persian than other low-resource languages . the study is the first comprehensive benchmarking of large language models . |
Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data Curation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for instruction tuning do not include associating instructions with existing datasets. |
| Approach: | They propose a dynamic growth paradigm for the automatic curation of instruction-tuning data . they use existing datasets to automatically construct instruction-uning datasets . |
| Outcome: | The proposed model reduces the API cost for generating instructions and provides high-quality data. |
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization (2025.findings-naacl)
Copied to clipboard
| Challenge: | a recent study investigated hallucinations in multi-document summarization tasks . but, it is unclear how challenges arising from handling multiple documents affect outputs . |
| Approach: | They investigate how hallucinations manifest in large language models when summarizing topic-specific information from a set of documents. |
| Outcome: | The proposed benchmarks show that the models generate more hallucinations than baselines . the results highlight the need for more effective approaches to mitigate hallucinosity in MDS . |
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes (2024.findings-acl)
Copied to clipboard
Sunjun Kweon, Junu Kim, Jiyoun Kim, Sujeong Im, Eunbyeol Cho, Seongsu Bae, Jungwoo Oh, Gyubok Lee, Jong Hak Moon, Seng Chan You, Seungjin Baek, Chang Hoon Han, Yoon Bin Jung, Yohan Jo, Edward Choi
| Challenge: | Clinical notes are an extensive repository of information specific to individual patients. |
| Approach: | They create synthetic large-scale clinical notes using publicly available case reports extracted from biomedical literature and train a clinical large language model, Asclepius. |
| Outcome: | The proposed model outperforms several other models and is supported by detailed evaluations conducted by GPT-4 and medical professionals. |
FinTextQA: A Dataset for Long-form Financial Question Answering (2024.acl-long)
Copied to clipboard
| Challenge: | Existing financial question answering datasets lack scope diversity and question complexity. |
| Approach: | They propose to use a dataset for long-form question answering in finance to evaluate QA systems. |
| Outcome: | The proposed dataset includes 1,262 high-quality, source-attributed QA pairs extracted and selected from finance textbooks and government agency websites. |
Speech-based Slot Filling using Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have shown an unprecedented ability across various language tasks. |
| Approach: | They propose to use prompts and LoRA fine-tuning to improve slot filling robustness . they propose a linearised knowledge injection scheme to integrate dynamic external knowledge into LLMs. |
| Outcome: | The proposed model improves slot filling with noisy ASR transcriptions with 6.7% and 17.6% absolute SLU-F1 improvements compared to a fully fine-tuned Flan-T5-XL model. |
Mitigating Demonstration Bias through Global Coevolutionary Reasoning (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for chain-of-thought prompting rely on manual demonstrations . experimental results show that GCR outperforms baseline methods without performance degradation . |
| Approach: | They propose a method that uses random samples to generate demonstrations in zero-shot settings. |
| Outcome: | The proposed method outperforms baseline methods on ten datasets without demonstration bias. |
Hire Me or Not? Examining Language Model’s Behavior with Occupation Attributes (2025.coling-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been widely integrated into production pipelines due to their impressive performance across multiple tasks. |
| Approach: | They construct a dataset using a standard occupation classification knowledge base and tested it on three families of LLMs. |
| Outcome: | The proposed framework analyzes LLMs’ behavior with respect to gender stereotypes in the context of occupation decision making. |
Can Large Language Models perform Relation-based Argument Mining? (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for RbAM fail to perform satisfactorily across different datasets. |
| Approach: | They propose to use relation-based argument mining to determine agreement (support) and disagreement (attack) relations amongst textual arguments in binary and ternary settings. |
| Outcome: | The proposed method outperforms the best performing (RoBERTa-based) baseline on two open-source LLMs and with GPT-3.5-turbo on several datasets for (binary and ternary) RbAM. |
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods to improve output quality without aggregating input tokens are limited by the complexity of aggregation of responses. |
| Approach: | They propose to extract and integrate segment-level commonalities from candidate samples to enhance performance of LLMs in open-ended and reasoning tasks. |
| Outcome: | The proposed method improves performance on reasoning, code generation and mathematical reasoning tasks without requiring additional models and overlooking the knowledge present among the candidates. |
Zero-shot and Few-shot Learning with Instruction-following LLMs for Claim Matching in Automated Fact-checking (2025.coling-main)
Copied to clipboard
| Challenge: | Claim matching (CM) is a binary classification task that can be used to determine if two claims can be verified using the same piece of evidence or fact-check. |
| Approach: | They propose a claim matching task that uses binary classification and large language models to test out learning approaches to the task. |
| Outcome: | The proposed task can be tackled by leveraging mature tasks such as natural language inference or paraphrase detection. |
Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent ubiquity and disruptive impacts of large language models have raised concerns about their potential to be misused. |
| Approach: | They propose a strategy that leverages LLMs' generative and emergent reasoning capabilities to counter human-written and LLM-generated disinformation. |
| Outcome: | The proposed strategy synthesizes authentic and deceptive LLM-generated content through paraphrase-based and perturbation-based prefix-style prompts, respectively. |
Evaluating Gender Bias of LLMs in Making Morality Judgements (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities in a multitude of NLP tasks, but are still not immune to limitations such as gender bias. |
| Approach: | They propose to use a dataset to examine whether LLMs possess gender bias when asked to give moral opinions. |
| Outcome: | The proposed models show that they are biased when asked to give moral opinions. |
CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions (2024.findings-acl)
Copied to clipboard
| Challenge: | a benchmark is designed to evaluate the capabilities of large language models (LLMs) as agents in decision making and operational tasks. |
| Approach: | They propose a benchmark to evaluate LLMs in the context of Chinese societal applications . they propose he benchmark will evaluate tool invocation ability of LLM and task completion ability . |
| Outcome: | The proposed benchmark features 398 APIs across 27 widely-used Apps across 14 domains. |
Motion Generation from Fine-grained Textual Descriptions (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing models for motion generation from textual descriptions are limited to coarse-grained descriptions. |
| Approach: | They build a large-scale language-motion dataset specializing in fine-grained textual descriptions . they feed it with step-by-step instructions with pseudo-code compulsory checks . quantitative evaluation shows that the model outperforms MotionDiffuse in generating spatially or chronologically composite motions . |
| Outcome: | The proposed model outperforms existing models in generating human motion sequences from textual descriptions by a large margin. |