Challenge: OOMB is a novel benchmark designed to assess the ability of large language models (LLMs) to extract and analyze opinions from diverse and complex online environments.
Approach: They propose an online opinion mining benchmark to assess the ability of large language models to extract and analyze opinions from diverse online environments.
Outcome: The proposed benchmark assesses the ability of large language models to mine opinions effectively from diverse and complex online environments.

Similar Papers

Evaluating the Performance of Large Language Models via Debates (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are evolving and impacting various fields . current methods for evaluation are based on fixed, domain-specific questions or rely on human input, making them unscalable.
Approach: They propose a benchmarking framework based on debates between LLMs, judged by another LLM.
Outcome: The proposed framework achieves rankings that align closely with popular rankings based on human input eliminating the need for costly crowdsourcing.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
Large Language Models for Generative Recommendation: A Survey and Visionary Discussions (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the field of natural language processing but are not fully able to leverage the generative power of LLM.
Approach: They examine the progress, methods, and future directions of large language models . they examine what generative recommendation is, why RS should advance to generative recommendations .
Outcome: The proposed approach can be simplified to generate recommendations from the entire pool of items.
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)

Copied to clipboard

Challenge: introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.
Approach: They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them.
Outcome: The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods.
Argument Mining with Fine-Tuned Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Argument Mining (AM) pipelines use fine-tuned large language models (LLMs) . initial approaches employ supervised machine learning algorithms, such as Maximum Entropy classifiers and Logistic Regressions.
Approach: They propose to model the three main AM sub-tasks as text generation tasks and fine-tune eight popular quantized and non-quantized large language models (LLMs) on the benchmark PE, AbstRCT, and CDCP datasets.
Outcome: The proposed pipeline achieves state-of-the-art across all AM sub-tasks and datasets, showing significant improvements over previous benchmarks.
From Annotation to Adaptation: Metrics, Synthetic Data, and Aspect Extraction for Aspect-Based Sentiment Analysis with Large Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Using a synthetic sports feedback dataset, we evaluate open-weight LLMs’ ability to extract aspect-polarity pairs.
Approach: They propose a metric to facilitate the evaluation of aspect extraction with generative models.
Outcome: The proposed metric improves the performance of open-weight LLMs in the Aspect-Based Sentiment Analysis task.
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations.
Approach: They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Outcome: The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Exploring the Potential of Large Language Models in Computational Argumentation (2024.acl-long)

Copied to clipboard

Challenge: Argumentation is an essential tool in various domains, including law, public policy, and artificial intelligence.
Approach: They propose to evaluate LLMs on various computational argumentation tasks . they organize existing tasks into six main categories and standardize the format of 14 datasets .
Outcome: The proposed model performs well on argument mining and argument generation tasks.
Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews (2025.findings-acl)

Copied to clipboard

Challenge: Existing large language models (LLMs) are difficult to evaluate and often lack the ability to capture user opinions.
Approach: They propose an LLM-powered interviewer that conducts in-the-moment user experience interviews right after users interact with LLMs and automatically gathers insights about user opinions from massive interview logs.
Outcome: The proposed interviewer captures interesting user opinions, e.g., bipolar views on the displayed reasoning process of DeepSeek-R1 and demands for information freshness and multi-modality.
Can Language Models Recognize Convincing Arguments? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have found that large language models can generate persuasive content without engaging in human experimentation.
Approach: They extend a dataset with debates, votes, and user traits to measure LLMs' ability to distinguish between strong and weak arguments, predict stances based on beliefs and demographic characteristics, and determine appeal of argument to individual based upon their traits.
Outcome: The proposed tasks outperform human predictions in detecting convincing arguments in debates, votes, and user traits.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations