Papers with fine-grained
Agentic AI for Human Resources: LLM-Driven Candidate Assessment (2026.eacl-demo)
Copied to clipboard
Kamer Ali Yuksel, Abdul Basit Anees, Ashraf Hatim Elneima, Sanjika Hewavitharana, Mohamed Al-Badrashiny, Hassan Sawaf
| Challenge: | Current systems rely on keyword matching and shallow keyword-based screening, leading to missed opportunities and inconsistent evaluations. |
| Approach: | They propose a framework that uses Large Language Models to automate candidate assessment in recruitment. |
| Outcome: | The proposed framework outputs detailed assessment reports, candidate comparisons, and ranked recommendations that are transparent, auditable, and suitable for real-world hiring workflows. |
An Empirical Study on Fine-Grained Named Entity Recognition (C18-1)
Copied to clipboard
Khai Mai, Thai-Hoang Pham, Minh Trung Nguyen, Tuan Duc Nguyen, Danushka Bollegala, Ryohei Sasano, Satoshi Sekine
| Challenge: | Named entity recognition (NER) is a well studied topic in natural language processing. |
| Approach: | They propose to remove the CNN layer and use dictionary and category embeddings to improve Japanese FG-NER performance. |
| Outcome: | The proposed method improves Japanese FG-NER F-score from 66.76% to 75.18%. |
STAMP: Selective Task-Aware Mechanism for Text Privacy (2026.eacl-long)
Copied to clipboard
| Challenge: | Experimental evaluations on SQuAD, Yelp, and AG News datasets demonstrate that STAMP achieves superior privacy–utility trade-offs across varying per-token privacy budgets. |
| Approach: | They propose a new framework for task-aware text privatization that selectively allocates privacy budgets across tokens by jointly considering (i) each token’s importance to the downstream task and (ii) its privacy sensitivity. |
| Outcome: | The proposed framework achieves superior privacy–utility trade-offs on SQuAD, Yelp, and AG News datasets. |
Leveraging Only the Category Name for Aspect Detection through Prompt-based Constrained Clustering (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Aspect category detection (ACD) aims to automatically identify user-concerned aspects from online reviews. |
| Approach: | They propose a method that relies on the category name of each aspect and a pretrained language model to generate constraints for clustering. |
| Outcome: | The proposed framework performs better than existing weakly supervised methods on nine benchmark datasets. |
End-to-end Neural Information Status Classification (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on information status classification and bridging anaphora recognition assume that gold mention or syntactic tree information is given. |
| Approach: | They propose an end-to-end neural approach for information status classification using a mention extraction component and an information status assignment component. |
| Outcome: | The proposed system achieves state-of-the-art on fine-grained IS classification based on gold mentions and better than baselines on ISNotes and SciCorp. |
FIDELITY: Fine-grained Interpretable Distillation for Effective Language Insights and Topic Yielding (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for topic modeling generate contextually specific and semantically intuitive topics, especially in dynamic environments and low-resource languages. |
| Approach: | They propose a hybrid method that combines topic modeling and text summarization to produce fine-grained, semantically rich, and contextually relevant output. |
| Outcome: | FIDELITY outperforms traditional models in topic diversity, similarity, and ability to process new, unseen documents. |
Enhancing Low-resource Fine-grained Named Entity Recognition by Leveraging Coarse-grained Datasets (2023.emnlp-main)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) often suffers from insufficient labeled data when the number of annotations exceeds several tens of labels. |
| Approach: | They propose a model with a fine-to- coarse mapping matrix to leverage hierarchical structure explicitly. |
| Outcome: | The proposed model outperforms both K-shot learning and supervised learning methods when dealing with a small number of fine-grained annotations. |
Adversarial Learning of Poisson Factorisation Model for Gauging Brand Sentiment in User Reviews (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing models for sentiment-topic extraction assume topics are grouped under discrete sentiment categories such as ‘positive’, ‘negative’ and ‘neural’. |
| Approach: | They propose a Brand-Topic Model which aims to detect brand-associated polarity-bearing topics from product reviews. |
| Outcome: | The proposed model outperforms existing models on Amazon reviews and shows that it is more coherent and unique than existing models. |
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization (2025.findings-emnlp)
Copied to clipboard
Fangxin Liu, Zongwu Wang, Jinhong Xia, Junping Zhao, Shouren Zhao, Jinjin Li, Jian Liu, Li Jiang, Haibing Guan
| Challenge: | Existing methods for quantization of large language models struggle to adapt to dynamic workloads. |
| Approach: | a new framework optimizes the trade-off between inference speed and accuracy . FlexQuant enables fine-grained, layer-wise mixed-precision quantization . |
| Outcome: | a new framework optimizes the trade-off between inference speed and accuracy . it achieves a 1.3 speedup across diverse language tasks with negligible accuracy loss . |
pFedGPT: Hierarchically Optimizing LoRA Aggregation Weights for Personalized Federated GPT Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning Large Language Models (LLMs) struggle with data heterogeneity and adapt shared global knowledge to individual client needs. |
| Approach: | They propose a framework that leverages Hierarchical Bayesian Optimization (HBO) for fine-grained, personalized LoRA aggregation. |
| Outcome: | The proposed framework achieves state-of-the-art (SOTA) performance on personalized FL benchmarks while introducing only minimal (approx. 4%) additional optimization overhead. |
Fine-grained Information Status Classification Using Discourse Context-Aware BERT (2020.coling-main)
Copied to clipboard
| Challenge: | Existing work on fine-grained information status (IS) relies on many hand-crafted linguistic features. |
| Approach: | They propose a discourse context-aware BERT model for fine-grained IS classification . they show an improvement of 10.5 F1 points for bridging anaphora recognition . |
| Outcome: | The proposed model achieves 4.8 absolute accuracy improvement on ISNotes corpus compared to previous work on bridging anaphora recognition . |
SWiPE: A Dataset for Document-Level Simplification of Wikipedia Pages (2023.acl-long)
Copied to clipboard
| Challenge: | Prior work on document-level simplification has focused on sentence-level edits, while many desirable edits require document- level context. |
| Approach: | They propose a dataset that reconstructs the document-level editing process from English Wikipedia to paired Simple Wikipedia articles. |
| Outcome: | The proposed dataset reconstructs the document-level editing process from English Wikipedia (EW) articles to paired Simple Wikipedia (SEW) pages. |
CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing LLMs rely on surface-level similarity or manual heuristics to evaluate personalization . Existing evaluation protocols for personalization are lacking sufficient data-driven validation. |
| Approach: | They propose a benchmark to assess personalization by mining CIPDs to quantify individual preferences. |
| Outcome: | The proposed benchmark provides a more comprehensive and discriminative standard than generic metrics. |
RanLoRA: Residual-aware Nonlinear Low-Rank Adaptation (2026.findings-acl)
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) relying on linear low-rank projections restricts adaptation to linear subspaces, limiting flexibility on complex downstream tasks. |
| Approach: | They propose a nonlinear low-rank Adaptation approach that leverages pretrained weights to decompose them into principal components that are kept frozen and residual components that can be used for task-specific adaptation. |
| Outcome: | The proposed approach outperforms vanilla LoRA and representative variants on commonsense reasoning, image classification, and mathematical reasoning tasks. |
HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring (2025.acl-long)
Copied to clipboard
| Challenge: | Existing literature focuses on binary, document-level detection, neglecting texts composed jointly by human and LLM contributions. |
| Approach: | They propose to use a dataset to generate human-AI coauthored texts via an automatic pipeline with word-level attribution labels. |
| Outcome: | The proposed method can detect human-AI coauthored texts with a numeric AI ratio. |
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation (2026.acl-long)
Copied to clipboard
Bosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke, Xiaoying Ling, Ying Zhang, Aohan Zeng, Hongning Wang, Minlie Huang
| Challenge: | Existing evaluation models for instruction-following have many shortcomings, such as substantial costs and unreliable assessments. |
| Approach: | They propose an LLM critic for fine-grained instruction-following evaluation using a checklist generator and a constraint-level preference optimization method. |
| Outcome: | The proposed model beats strong LLM-as-a-Judge baselines in evaluations under lower computational overhead compared to baselines. |
Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science (2026.findings-acl)
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) outperform existing benchmarks in both natural language and coding domains. |
| Approach: | They propose a scalable benchmark that integrates vision and language modalities to address this gap by eliminating textual shortcuts. |
| Outcome: | The new benchmark outperforms existing benchmarks in both natural language and coding domains. |
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Social reasoning is a core competency of social intelligence and requires specialized neural and cognitive systems to be able to interpret multimodal interactions. |
| Approach: | They propose to use social reasoning traces to generate fine-grained explanations using external knowledge. |
| Outcome: | The proposed model is based on 272 videos of human interactions and 1,486 human-annotated reasoning traces related to inferences about these interactions. |
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | SteerVLM is a lightweight steering module designed to guide Vision-Language Models (VLMs) towards outputs that better adhere to desired instructions. |
| Approach: | They propose a lightweight steering module that learns from latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the language modality with image context. |
| Outcome: | The proposed steering module outperforms existing intervention techniques on steering and hallucination mitigation benchmarks for VLMs. |
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Infodemics and health misinformation have significant negative impact on individuals and society . generative AI has significantly accelerated the spread and expanded the reach of health misinfo . |
| Approach: | MM-Health is a large scale multimodal misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated multiplemodal information . |
| Outcome: | MM-Health is a large scale misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated content . |