Challenge: Recent advances in measuring hardness-wise properties of data guide language models in sample selection within low-resource scenarios.
Approach: They propose to use class-wise hardness to measure class-specific properties of data in the semantic embedding space by modeling class geometry in the . semantic embeddining space.
Outcome: The proposed method surpasses instance-level metrics by over 59 percent on Pearson‘s correlation on measuring class-wise hardness.

Similar Papers

The Unreasonable Effectiveness of Easy Training Data for Hard Tasks (2024.acl-long)

Copied to clipboard

Challenge: Existing pretrained language models perform well on hard data, but hard data is noisier and costlier to collect.
Approach: They propose to use in-context learning, linear classifier heads, and QLoRA to show that pretrained language models generalize relatively well from easy to hard data.
Outcome: The proposed model generalizes well from easy to hard data even better than oracle models finetuned on hard data.
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes (2024.acl-long)

Copied to clipboard

Challenge: Complex reasoning ability is one of the most important features of Large Language Models.
Approach: They propose a new benchmark that measures the reasoning ability of Large Language Models . it contains 900 algorithmic questions belonging to the NP-Hard complexity class .
Outcome: The proposed benchmark contains 900 questions belonging to the NP-Hard complexity class and is updated on a monthly basis.
DirectProbe: Studying Representations without Classifiers (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches for probing opaque representations often use training classifiers and use the accuracy, mutual information, or complexity as a proxy for the representation’s goodness.
Approach: They propose a heuristic that directly studies the geometry of a representation by building upon the notion of 'version space' they argue that doing so can be unreliable because different representations may need different classifiers .
Outcome: Experiments with linguistic tasks and contextualized embeddings show that even without training classifiers, DirectProbe can shine lights on how an embeddable space represents labels and anticipate the classifier performance for the representation.
Semantic Accuracy in Natural Language Generation: A Thesis Proposal (2023.acl-srw)

Copied to clipboard

Challenge: Using large pre-trained language models, it is essential to research their reliability . if a human does not know the answer to a question, the socially acceptable behavior is to say 'I do not know' failing to fulfill this expectation can lead to distrust, or spread of misinformation.
Approach: They propose a method for evaluating semantic accuracy and a benchmark for NLG metrics.
Outcome: The proposed method evaluates semantic accuracy and provides a benchmark for NLG metrics.
Measuring Fine-Grained Domain Relevance of Terms: A Hierarchical Core-Fringe Approach (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to measure fine-grained domain relevance are needed for downstream tasks in natural language processing.
Approach: They propose to measure fine-grained domain relevance, defined as the degree that a term is relevant to a given domain.
Outcome: The proposed method outperforms baselines and surpasses professional human performance.
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios (2024.findings-emnlp)

Copied to clipboard

Challenge: Using large language models, we evaluated their robustness on multiple datasets.
Approach: They propose a new metric for assessing model robustness by empirical evaluation of several models on multiple datasets.
Outcome: The proposed metric is based on a set of datasets that are constructed by introducing naturally-occurring, non-malicious perturbations or by generating semantically equivalent paraphrases of input questions or statements.
Measure and Improve Robustness in NLP Models: A Survey (2022.naacl-main)

Copied to clipboard

Challenge: Despite the performance gains, NLP models are still fragile and brittle to out-of-domain data, adversarial attacks, or small perturbation to the input.
Approach: They propose a survey of how to define, measure and improve robustness in NLP by connecting multiple definitions of robustness and identifying failures.
Outcome: The proposed models are robust against unseen or challenging scenarios, but are still fragile and brittle to out-of-domain data and adversarial attacks.
Calibrated Interpretation: Confidence Estimation in Semantic Parsing (2023.tacl-1)

Copied to clipboard

Challenge: Sequence generation models are increasingly being used to translate natural language into programs . calibration of such models is a key component of safety, says aaron sagar .
Approach: They investigate whether calibration of popular generation models varies across models and datasets . they find that calibration varies among models and data sets, and that it is important to include it in evaluations if it is included .
Outcome: The calibration of popular generation models varies across models and datasets . the authors find that the accuracy of models is dependent on confidence .
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)

Copied to clipboard

Challenge: Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities.
Approach: They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility.
Outcome: The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions.
Quantifying training challenges of dependency parsers (C18-1)

Copied to clipboard

Challenge: a new metric is introduced to evaluate the difficulty to learn a given class of dependencies . a series of systematic computations using that metric have revealed interesting properties of the 3 considered parsing algorithms .
Approach: They introduce a new metric to evaluate the difficulty to learn a given class of dependencies . they use it to characterize the information conveyed by cross-lingual parsers .
Outcome: The proposed metric reveals the kind of dependencies that require high effort during training . it also shows that cross-lingual parsers can provide better quality information .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations