Geographic Citation Gaps in NLP Research (2022.emnlp-main)

Copied to clipboard

Challenge: a vast number of papers accepted at top NLP venues come from a handful of western countries and (lately) China.
Approach: They ask researchers to examine the relationship between geographical location and publication success . they use a dataset of 70,000 papers from the ACL Anthology to examine their citation network .
Outcome: The proposed dataset of 70,000 papers from the ACL Anthology shows that there are substantial geographical disparities in paper acceptance and citations .

Similar Papers

Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World (2022.aacl-main)

Copied to clipboard

Challenge: Linguistic disparity in the NLP world is widely acknowledged, but the reasons behind it are rarely discussed within the field.
Approach: They propose to categorise languages based on speaker population and vitality . they also analyse the distribution of language data resources and amount of NLP/CL research .
Outcome: The proposed model identifies the reasons for the disparity and suggests ways to overcome it.
Examining Citations of Natural Language Processing Literature (2020.acl-main)

Copied to clipboard

Challenge: citations of NLP papers have decreased in recent years, but long papers get three times as many citation as short papers . citation data from the ACL Anthology and Google Scholar can be used to understand the field and quantify the impact of different types of papers.
Approach: They extract data from the ACL Anthology and Google Scholar to examine trends in citations of NLP papers.
Outcome: The results show that only about 56% of the papers in AA are cited ten or more times . CL Journal has the most cited papers, but its citation dominance has lessened .
NLP Needs Diversity outside of ‘Diversity’ (2025.findings-emnlp)

Copied to clipboard

Challenge: a new position paper argues that diversity in NLP is concentrated on a small number of areas surrounding fairness .
Approach: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Outcome: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Forgotten Knowledge: Examining the Citational Amnesia in NLP (2023.acl-long)

Copied to clipboard

Challenge: a recent study examines how far back in time we tend to cite papers . citation patterns are correlated with age, age, and other factors .
Approach: They analyze citation patterns across time and examine temporal changes . they find that 62% of cited papers are from the immediate five years prior to publication .
Outcome: The authors show that citing papers is the primary method of scientific writing . they show that the trend has reversed and current papers have low temporal diversity .
Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields (2025.coling-main)

Copied to clipboard

Challenge: citation age is a key factor in determining whether older works are cited in scientific journals or not.
Approach: They examine the tendency of NLP to cite older work across 20 fields of study over 43 years (1980–2023) . they put NLP’s propensity to citation older work in the context of these 20 other fields to see whether differences can be observed .
Outcome: The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks)
Shoulders of Giants: A Look at the Degree and Utility of Openness in NLP Research (2024.acl-short)

Copied to clipboard

Challenge: We analysed a sample of NLP research papers published in ACL Anthology . we observe a wide language-wise disparity in publicly available NLP-related artefacts .
Approach: They analysed NLP research papers archived in ACL Anthology to quantify degree of openness and benefit of such an open culture in the NLP community.
Outcome: The results show that more than 30% of the papers published in ACL Anthology do not release their artefacts publicly.
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)

Copied to clipboard

Challenge: linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages .
Approach: They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions .
Outcome: The proposed model is representative of the language diversity and coverage of natural language processing systems.
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)

Copied to clipboard

Challenge: African languages are often left behind in state-of-the-art natural language processing systems and large language models.
Approach: They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions .
Outcome: The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years .
NLP Scholar: A Dataset for Examining the State of NLP Research (2020.lrec-1)

Copied to clipboard

Challenge: Google Scholar is the largest web search engine for academic literature and provides access to rich metadata associated with the papers.
Approach: They extracted citation information from the ACL Anthology (AA) for about 44 thousand NLP papers and identified authors who published at least three papers there.
Outcome: The ACL Anthology (AA) is the largest repository of articles on Natural Language Processing (NLP).
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other)
Approach: They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 .
Outcome: The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations