Papers with Claude-2
Over-Reasoning and Redundant Calculation of Large Language Models (2024.eacl-short)
Copied to clipboard
| Challenge: | Large language models (LLMs) can solve problems step-by-step, but it is unclear whether they know when to use CoT and whether they are always necessary. |
| Approach: | They propose to use LLMs to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero. |
| Outcome: | The proposed model generates redundant calculations and reasoning on a manually constructed math QA dataset, but it is unclear whether it is necessary to use CoT reasoning. |
BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP (2024.lrec-main)
Copied to clipboard
Mohsinul Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, Enamul Hoque
| Challenge: | Large Language Models (LLMs) have emerged as one of the most important breakthroughs in natural language processing. |
| Approach: | They propose to evaluate LLMs in Bengali to benchmark their performance . they select Bangla NLP tasks such as text summarization, question answering, paraphrasing . |
| Outcome: | The proposed model performs better in some tasks than current models, but in most tasks, it is poor . |
Beyond Binary: Towards Embracing Complexities in Cyberbullying Detection and Intervention - a Position Paper (2024.lrec-main)
Copied to clipboard
Kanishk Verma, Kolawole John Adebayo, Joachim Wagner, Megan Reynolds, Rebecca Umbach, Tijana Milosevic, Brian Davis
| Challenge: | Existing methods for CB detection oversimplify the problem of CB as a binary classification task. |
| Approach: | They propose to use large language models to generate CB-related datasets . they propose to combine cognitive and linguistic models to help identify CB incidents . |
| Outcome: | The proposed approach aims to help researchers and policymakers make informed decisions . it uses large language models such as Claude-2 and Llama2-Chat to generate CB-related datasets . |
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models exhibit remarkable generative capabilities but can be misused for harmful purposes. |
| Approach: | They propose a framework that transforms natural language inputs into code inputs. |
| Outcome: | The proposed framework bypasses the safety guardrails of all models more than 80% of the time. |
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency (2024.lrec-main)
Copied to clipboard
| Challenge: | Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination. |
| Approach: | They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully. |
| Outcome: | The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness. |
What Factors Influence LLMs’ Judgments? A Case Study on Question Answering (2024.lrec-main)
Copied to clipboard
Lei Chen, Bobo Li, Li Zheng, Haining Wang, Zixiang Meng, Runfeng Shi, Hao Fei, Jun Zhou, Fei Li, Chong Teng, Donghong Ji
| Challenge: | Existing studies indicate that Large Language Models perform at a level comparable to humans with advantages of speed and cost-effectiveness in different fields. |
| Approach: | They propose to introduce four unexplored factors and a new dimension of question difficulty to provide a more comprehensive understanding of LLMs’ judgments across varying question intricacies. |
| Outcome: | The proposed dimensions of question difficulty and answer quantity provide valuable insights into optimizing LLMs’ performance as judges. |
Can Large Language Models Unlock Novel Scientific Research Ideas? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and ChatGPT have marked a turning point in the integration of Artificial Intelligence (AI) into people’s everyday lives. |
| Approach: | They conduct a human evaluation of the novelty, relevancy, and feasibility of the generated future research ideas. |
| Outcome: | The proposed models generate more diverse ideas than GPT-4, GPT-3.5, and Gemini 1.0. |