Challenge: Academic citations are widely used for evaluating research and tracing knowledge flows.
Approach: They propose a computational pipeline to quantify citation fidelity at the sentence level by identifying citations in citing papers and corresponding claims in cited papers.
Outcome: The proposed pipeline identifies citations in citing papers and the corresponding claims in cited papers and applies supervised models to measure fidelity at the sentence level.

Similar Papers

Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields (2025.coling-main)

Copied to clipboard

Challenge: citation age is a key factor in determining whether older works are cited in scientific journals or not.
Approach: They examine the tendency of NLP to cite older work across 20 fields of study over 43 years (1980–2023) . they put NLP’s propensity to citation older work in the context of these 20 other fields to see whether differences can be observed .
Outcome: The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks)
Forgotten Knowledge: Examining the Citational Amnesia in NLP (2023.acl-long)

Copied to clipboard

Challenge: a recent study examines how far back in time we tend to cite papers . citation patterns are correlated with age, age, and other factors .
Approach: They analyze citation patterns across time and examine temporal changes . they find that 62% of cited papers are from the immediate five years prior to publication .
Outcome: The authors show that citing papers is the primary method of scientific writing . they show that the trend has reversed and current papers have low temporal diversity .
On Forgetting to Cite Older Papers: An Analysis of the ACL Anthology (2020.acl-main)

Copied to clipboard

Challenge: a growing number of published papers are citing older work, but the rate of citations is stable . a recent paper cited work from recent years, whereas papers published 15 or more years ago are cited at a stable rate.
Approach: They analyze citations in papers published at selected ACL venues between 2010 and 2019 . they find that recent papers are cited significantly more often in recent years .
Outcome: The authors analyze citations in journals and conferences between 2010 and 2019 . they find that recent papers cite more recent work, but papers published 15 or more years ago are cited at a stable rate.
In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis (2026.acl-long)

Copied to clipboard

Challenge: citation counts are a shallow view that fails to capture how a paper has influenced subsequent work.
Approach: They propose a task to generate nuanced, expressive, and time-aware impact summaries . they analyze fine-grained confirmatory and correction citation intents to generate summary .
Outcome: The proposed task shows moderate to strong human correlation on subjective metrics such as insightfulness.
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other)
Approach: They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 .
Outcome: The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show .
L-CiteEval: A Suite for Evaluating Fidelity of Long-context Models (2025.acl-long)

Copied to clipboard

Challenge: Long-context models (LCMs) have seen remarkable advancements in recent years, facilitating tasks like long-document QA.
Approach: They propose an out-of-the-box suite that can assess both generation quality and fidelity in long-context understanding tasks.
Outcome: The proposed suite can assess both generation quality and fidelity in long-context understanding tasks.
Examining Citations of Natural Language Processing Literature (2020.acl-main)

Copied to clipboard

Challenge: citations of NLP papers have decreased in recent years, but long papers get three times as many citation as short papers . citation data from the ACL Anthology and Google Scholar can be used to understand the field and quantify the impact of different types of papers.
Approach: They extract data from the ACL Anthology and Google Scholar to examine trends in citations of NLP papers.
Outcome: The results show that only about 56% of the papers in AA are cited ten or more times . CL Journal has the most cited papers, but its citation dominance has lessened .
CiteBench: A Benchmark for Scientific Citation Text Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on citation text generation are based upon widely diverging task definitions, making it hard to study this task systematically.
Approach: They propose a benchmark for citation text generation that unifies multiple datasets and enables standardized evaluation of citation texts across task designs and domains.
Outcome: The proposed benchmark examines the performance of multiple strong baselines and enables standardized evaluation of citation text generation models across task designs and domains.
CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have emerged as powerful assistants for scientific writing, but reliability of LLM alone is in doubt.
Approach: They propose a retrieval-aware agent framework to provide more faithful grounding for citation validation.
Outcome: The proposed framework improves over the baseline and achieves 68.1% accuracy on the CiteME benchmark, approaching human performance.
CiteEval: Principle-Driven Citation Evaluation for Source Attribution (2025.acl-long)

Copied to clipboard

Challenge: Current evaluation frameworks rely on NLI to assess binary or ternary support from cited sources, which is suboptimal for citation evaluation.
Approach: They propose a citation evaluation framework based on fine-grained citation ratings within a broad context and construct a multi-domain benchmark with high-quality human annotations.
Outcome: The proposed framework provides a high-quality human annotation benchmark and a suite of model-based metrics that exhibit strong correlation with human judgments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations