Papers with F1-scores
BiQuAD: Towards QA based on deeper text understanding (2021.starsem-1)
Copied to clipboard
| Challenge: | Recent question answering and machine reading benchmarks require systems to pinpoint the span of the answer to a given text. |
| Approach: | They propose a dataset that requires deeper comprehension to answer questions extractively and deductively. |
| Outcome: | The proposed dataset outperforms existing benchmarks on extractive and deductive questions. |
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection (2024.eacl-srw)
Copied to clipboard
| Challenge: | Existing approaches to multimodal hateful content detection focus on detecting hate speech from text-based content, but they fail to address modality-specific features. |
| Approach: | They propose a context-aware attention framework for multimodal hateful content detection that integrates an attention layer to meaningfully align the visual and textual features. |
| Outcome: | The proposed framework achieves F1-scores of 69.7% and 70.3% on two hateful meme datasets and shows 2.5% and 3.2% performance improvement over the state-of-the-art systems. |
Does Character-level Information Always Improve DRS-based Semantic Parsing? (2023.starsem-1)
Copied to clipboard
| Challenge: | incorporating character-level information does not improve the performance in English and German, and is not sensitive to correct character order in Dutch. |
| Approach: | They propose to incorporate character-level representations into a neural semantic parser for Discourse Representation Structures and to test their performance using order of character sequences. |
| Outcome: | The proposed parser improves in English, German, Dutch, and Italian in four languages. |
Prediction for the Newsroom: Which Articles Will Get the Most Comments? (N18-3)
Copied to clipboard
| Challenge: | a new method to support manual moderation of discussion sections is proposed. |
| Approach: | They propose to support manual moderation by proactively drawing attention of moderators to articles that most likely need their intervention. |
| Outcome: | The proposed method outperforms the current state-of-the-art methods on a 7-million-comment dataset. |
Using Subtext to Enhance Generative IDRR (2025.acl-short)
Copied to clipboard
| Challenge: | Arguments contain subtexts, but they are connotative and need prompts to be recognized . a lightweight subtext generator is helpful when the prompt doesn't raise a complex CoT. |
| Approach: | They leverage LLaMA to generate subtexts for argument pairs and verify their effectiveness . they construct a baseline IDRR using the decoder-only backbone LLama . |
| Outcome: | The proposed approach achieves higher F1 scores on two benchmarks than previous models. |
Nested Named Entity Recognition via Second-best Sequence Learning and Decoding (2020.tacl-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is the task of identifying text spans associated with proper names and classifying them according to their semantic class such as person or organization. |
| Approach: | They propose a method that treats the tag sequence for nested entities as the second best path within the span of their parent entity. |
| Outcome: | The proposed method achieves F1-scores of 85.82%, 84.34%, and 77.36% on ACE-2004, ACE 2005, and GENIA datasets. |
Fact from Fiction: Finding Serialized Novels in Newspapers (2025.acl-srw)
Copied to clipboard
| Challenge: | Among underrepresented but widely read forms are serialized fiction and feuilleton novels embedded in newspapers rather than published as standalone volumes. |
| Approach: | They propose to annotate 1,394 articles and evaluate classification pipelines using both selected linguistic features and embeddings to identify serialized fiction and feuilleton fiction. |
| Outcome: | The proposed methods achieve F1-scores of 0.91 in an annotated dataset of 1,394 articles and support the construction of alternative literary corpora and contribute to work on modeling the fiction–nonfiction boundary at scale. |
ViHOS: Hate Speech Spans Detection for Vietnamese (2023.eacl-main)
Copied to clipboard
| Challenge: | Increasing use of social networking sites can cause problems for human moderators to review tagged comments. |
| Approach: | They present a dataset that contains 26k spans on 11k comments and detailed annotation guidelines . they also provide definitions of hateful and offensive spans in Vietnamese comments . |
| Outcome: | The proposed dataset shows that it is difficult to detect specific types of spans in the dataset . the dataset is the first human-annotated corpus containing 26k spans on 11k comments . |
Inverse is Better! Fast and Accurate Prompt for Few-shot Slot Tagging (2022.findings-acl)
Copied to clipboard
| Challenge: | Recent results show that prompting methods are inefficient for slot tagging tasks . inverse prompting only requires a one-turn prediction for each slot type . |
| Approach: | They propose an inverse prompting paradigm that reversely predicts slot values given slot types . the method is faster and significantly improves the effect on 10-shot setting . |
| Outcome: | The proposed method improves over 6.1 F1-scores on 10-shot setting and achieves new state-of-the-art performance. |
Retrieve and Copy: Scaling ASR Personalization to Large Catalogs (2023.emnlp-industry)
Copied to clipboard
| Challenge: | End-to-end ASR models struggle to recognize uncommon domain-specific words due to limited audio context. |
| Approach: | They propose a "Retrieve and Copy" mechanism to improve latency while retaining the accuracy even when scaled to a large catalog. |
| Outcome: | The proposed method achieves 6% more word error rate reduction and 3.6% improvement in F1 when scaled to a large catalog size while retaining the accuracy. |
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey (2026.findings-eacl)
Copied to clipboard
| Challenge: | Social media platforms such as X (formerly Twitter), Facebook, and Reddit generate user-generated content. |
| Approach: | They propose a framework to assess privacy risks in social media by evaluating vulnerabilities across six dimensions: data collection, preprocessing, visibility, fairness, computational risk, and regulatory compliance. |
| Outcome: | The proposed framework assesses privacy risks across six dimensions . it achieves F1-scores of 0.58–0.84, but incurs 1% - 23% drop under fine-tuning . |
To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource Settings (2021.acl-long)
Copied to clipboard
| Challenge: | Part-of-Speech (POS) tags are routinely included in many NLP tasks. |
| Approach: | They propose to use POS tags to examine morphological learning in low-resource languages . they find that POS tagging improves joint segmentation and glossing . |
| Outcome: | The proposed task is tested on two identical datasets with the Transformer architecture. |
Efficient, Uncertainty-based Moderation of Neural Networks Text Classifiers (2022.findings-acl)
Copied to clipboard
| Challenge: | A series of benchmarking experiments based on three different datasets and three state-of-the-art classifiers show that our framework can improve the classification F1-scores by 5.1 to 11.2% (up to approx. 98 to 99%) |
| Approach: | They propose a semi-automated approach that passes unconfident, probably incorrect classifications to human moderators to minimize the workload. |
| Outcome: | The proposed approach can improve the classification F1-scores by 5.1 to 11.2% (up to approx. 98 to 99%) while reducing the moderation load up to 73.3% compared to a random moderation. |
GLGR: Question-aware Global-to-Local Graph Reasoning for Multi-party Dialogue Reading Comprehension (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches for multi-hop reasoning are lacking for local graph reasoning . existing approaches neglect local semantic structures in utterances . |
| Approach: | They propose a question-aware global-to-local graph reasoning approach that expands the canonical Interlocutor-Utterance graph by introducing a query node. |
| Outcome: | The proposed approach outperforms existing methods on Molweni and FriendsQA. |
Contextual String Embeddings for Sequence Labeling (C18-1)
Copied to clipboard
| Challenge: | Recent advances in language modeling have made it viable to model language as distributions over characters. |
| Approach: | They propose to leverage internal states of a trained character language model to produce a new type of word embeddings. |
| Outcome: | The proposed embeddings outperform the state-of-the-art on four classic sequence labeling tasks. |
Smart “Chef”: Verifying the Effect of Role-based Paraphrasing for Aspect Term Extraction (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Aspect Term Extraction (ATE) is a task of automatically extracting aspect terms from sentences. |
| Approach: | They propose to automatically rewrite sentences from virtual experts with different roles . they leverage ChatGPT to determine virtual experts in the considered domains . |
| Outcome: | The proposed method can be used to expand the predictions obtained on the original sentences without retraining or fine-tuning the baseline extractors. |
Fine-Grained Error Analysis and Fair Evaluation of Labeled Spans (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotations with incorrect label or boundaries count as two errors instead of one, despite being closer to the target annotation than false positives or false negatives. |
| Approach: | They propose an algorithm for error identification in flat and multi-level annotations and propose a procedure for calculating meaningful precision, recall, and F1-scores based on the more fine-grained error types. |
| Outcome: | The proposed procedure prevents double penalties and allows for a more detailed error analysis, providing more insight into the actual weaknesses of a system. |
Online Near-Duplicate Detection of News Articles (2020.lrec-1)
Copied to clipboard
| Challenge: | Near-duplicate documents are prevalent in news corpora and cost significant . bloating corporata with redundant information and computational costs are among the costs . |
| Approach: | They propose an online system which flags a near-duplicate document by finding its most likely original. |
| Outcome: | The proposed system can be used in many real-world applications. |
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) do not perform well on the datasets. |
| Approach: | They propose to use a Spatial Reasoning Characterization framework and a spatial reasoning path framework to study spatial reasoning. |
| Outcome: | The proposed framework and datasets outperform state-of-the-art models in spatial reasoning. |
Identifying the Human Values behind Arguments (2022.acl-long)
Copied to clipboard
| Challenge: | et al., 2003) examines human values in natural language arguments . authors provide a dataset of 5270 arguments from four geographical cultures . |
| Approach: | They propose a multi-level taxonomy of human values with 54 values and a dataset of 5270 arguments from four geographical cultures, manually annotated for human values. |
| Outcome: | The proposed model shows that human values are more diverse than previously thought . it shows that people disagree on the best course forward on controversial issues . |
A Generalized Approach to Protest Event Detection in German Local News (2022.lrec-1)
Copied to clipboard
| Challenge: | Social scientists conduct protest event analysis to learn about developments and trends of the forms, scale and hot topics of political protests. |
| Approach: | They propose to use a German language resource to analyze newspaper articles on protest events . they train and evaluate transformer-based text classifiers to automatically detect relevant newspaper articles . |
| Outcome: | The proposed method achieves a binary F1-score of 93.3 %, but does not generalize well to other datasets. |
CleanCoNLL: A Nearly Noise-Free Named Entity Recognition Dataset (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing models achieve F1-scores comparable to or exceed noise level in CoNLL-03 . current models have significant annotation errors, incompleteness, and inconsistencies in the data . |
| Approach: | They propose to add a layer of entity linking annotation to the CoNLL-03 corpus to correct 7.0% of all labels. |
| Outcome: | The proposed approach corrects 7.0% of all labels in the English CoNLL-03 dataset. |
Aspect-based Document Similarity for Research Papers (2020.coling-main)
Copied to clipboard
| Challenge: | Traditional document similarity measures do not consider in what aspects two documents are similar. |
| Approach: | They extend document similarity with aspect information by performing a pairwise document classification task. |
| Outcome: | The proposed approach is best performing on 172,073 research paper pairs from the ACL Anthology and CORD-19 corpus. |
Exploring the Potential of Large Language Models (LLMs) for Low-resource Languages: A Study on Named-Entity Recognition (NER) and Part-Of-Speech (POS) Tagging for Nepali Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Language Models excel in various tasks like Named Entity Recognition and Part-of-Speech tagging. |
| Approach: | They propose to use large language models to perform NLP tasks such as Named Entity Recognition and Part-of-Speech tagging in Nepali. |
| Outcome: | The proposed models perform better than other approaches for Nepali NER and POS tagging tasks. |
Incorporating Zoning Information into Argument Mining from Biomedical Literature (2022.lrec-1)
Copied to clipboard
| Challenge: | Argumentative zoning is a text zonation scheme that is used to segment text into zones that serve distinct functions. |
| Approach: | They propose to use zoning information to incorporate into argument mining tasks . they add zonation labels predicted by an off-the-shelf model to the beginning of each sentence . |
| Outcome: | The proposed models improve argument mining models without additional annotation cost. |
German SRL: Corpus Construction and Model Training (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing semantic role annotation resources are lacking for German. |
| Approach: | They propose a translation-based approach to train German semantic role models using semantic annotations and alignment models. |
| Outcome: | The proposed method achieves competitive evaluation scores, but avoids limitations of previous approaches. |
Reproducing a Morphosyntactic Tagger with a Meta-BiLSTM Model over Context Sensitive Token Encodings (2020.lrec-1)
Copied to clipboard
| Challenge: | Reproducibility of research results is only recently beginning to be practiced and acknowledged . a research culture that focuses on beating previous benchmarks while disregarding the need to contribute to scientific knowledge and understanding is a problem, says a researcher. |
| Approach: | They reproduced work on morphosyntactic tagging using a meta-model . they did not contact the original authors for reproduction . |
| Outcome: | The proposed model outperforms previous models on morphological tagging tasks but fails to match the F1-scores reported for the meta-BiLSTM model. |
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It’s Best to Relate Perspectives! (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to subjectivity in natural language processing are subjective . authors argue that disagreement should not be regarded as a problem . |
| Approach: | They propose to account for subjective perspectives of individuals and objective concepts that build a common ground between annotators. |
| Outcome: | The proposed architectures increase the averaged annotator-individual F1-scores up to 43% over a majority-label model. |
LLMSegm: Surface-level Morphological Segmentation Using Large Language Model (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to morphological segmentation split word into its morphemes . LLMSegm is applicable in low-data settings and low-resourced languages . |
| Approach: | They propose a novel approach to surface-level morphological segmentation leveraging large language models. |
| Outcome: | The proposed method is applicable in low-data settings and low-resource languages. |
Resolving Legalese: A Multilingual Exploration of Negation Scope Resolution in Legal Documents (2024.lrec-main)
Copied to clipboard
| Challenge: | Negation scope resolution is a challenging task for NLP because of the complexity of legal texts and lack of annotated in-domain negation corpora. |
| Approach: | They propose to use annotated court decisions to improve negation scope resolution . they release annotations in german, french, and italian to train models without legal data . |
| Outcome: | The proposed models achieve token-level F1-scores of up to 86.7% in zero-shot and multilingual settings. |
SGCM: Salience-Guided Context Modeling for Question Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Identifying relevant sentences to answers is crucial for reasoning the possible questions before generation. |
| Approach: | They propose a salience-guided approach to enhance Paragraph-level Question Generation by identifying salient sentences that manifest relevance. |
| Outcome: | The proposed approach achieves Rouge-L, BLEU4, BERTScore, Q-BLUE-3 and F1-scores compared to baseline on FairytaleQA. |
LLM-FK: Multi-Agent LLM Reasoning for Foreign Key Detection in Large-Scale Complex Databases (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for detecting missing foreign keys are limited in capturing semantic dependencies across schemas. |
| Approach: | They propose a framework that integrates four agents to detect missing foreign keys . they propose combinatorial search space explosion, ambiguous inference and global inconsistency . |
| Outcome: | The proposed framework achieves F1-scores above 93% on large-scale MusicBrainz database . it reduces candidate search space by two to three orders of magnitude without losing true FKs . |