Collaboration or Corporate Capture? Quantifying NLP’s Reliance on Industry Artifacts and Contributions (2024.acl-long)
Copied to clipboard
| Challenge: | EMNLP 2022 citations are three times greater than expected for pre-trained models . industry participation in the Association of Computational Linguistics (ACL) anthology has increased 180% from 2017 to 2022. |
| Approach: | They surveyed 100 papers published at EMNLP 2022 to determine the ratio of their citations to industry models. |
| Outcome: | a new study shows that industry citations are three times greater than expected . the study aims to better understand whether industry collaboration is still collaboration . industry participation in the 2023 AI index report is the top takeaway . |
Similar Papers
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)
Copied to clipboard
| Challenge: | In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other) |
| Approach: | They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 . |
| Outcome: | The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show . |
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research (2023.acl-long)
Copied to clipboard
Mohamed Abdalla, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort
| Challenge: | Recent advances in deep learning methods for natural language processing (NLP) have created new business opportunities and made NLP research critical for industry development. |
| Approach: | They examine industry presence in the field since the early 90s and characterize it using a corpus of 78,187 NLP publications and 701 resumes of NLP publication authors. |
| Outcome: | The authors find that industry presence among NLP authors has been steady before a steep increase over the past five years (180% growth from 2017 to 2022). |
Is NLP Ready for Standardization? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | a number of scientific fields, including telecommunications, networks and multimedia, lack standards in the field of NLP. |
| Approach: | They propose to examine how NLP lacks standards and how that can impact society, industry and regulations. |
| Outcome: | The proposed standards examine the needs of NLP researchers and industry . they argue that the lack of standards can impact the field, society and industry. |
Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty . |
| Approach: | They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes . |
| Outcome: | The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies. |
Shoulders of Giants: A Look at the Degree and Utility of Openness in NLP Research (2024.acl-short)
Copied to clipboard
| Challenge: | We analysed a sample of NLP research papers published in ACL Anthology . we observe a wide language-wise disparity in publicly available NLP-related artefacts . |
| Approach: | They analysed NLP research papers archived in ACL Anthology to quantify degree of openness and benefit of such an open culture in the NLP community. |
| Outcome: | The results show that more than 30% of the papers published in ACL Anthology do not release their artefacts publicly. |
Understanding the Gap: an Analysis of Research Collaborations in NLP and Language Documentation (2025.findings-acl)
Copied to clipboard
| Challenge: | despite 20 years of NLP work, practical use of this work remains vanishingly scarce. |
| Approach: | They propose to use interviews and surveys to examine the lack of NLP adoption in LD . they find that linguists and language communities have little or no use of Nlp in their work . |
| Outcome: | a new study shows that linguists and language researchers are not using NLP in LD . the findings highlight the importance of misaligned professional incentives and LD software . |
On “Scientific Debt” in NLP: A Case for More Rigour in Language Model Pre-Training Research (2023.acl-long)
Copied to clipboard
Made Nindyatama Nityasya, Haryo Wibowo, Alham Fikri Aji, Genta Winata, Radityo Eko Prasojo, Phil Blunsom, Adhiguna Kuncoro
| Challenge: | Despite rapid recent progress, current research practices conflate different sources of model improvement without conducting proper ablation studies and principled comparisons . authors conclude with recommendations for how to encourage and incentivize this line of work . |
| Approach: | They critique current research practices in the field of language model pre-training . they examine the success of language models pre-trained on large amounts of data . |
| Outcome: | The proposed models can achieve competitive or better performance than BERT under comparable conditions. |
The Nature of NLP: Analyzing Contributions in NLP Papers (2025.acl-long)
Copied to clipboard
| Challenge: | despite this, what constitutes NLP research remains debated . |
| Approach: | They propose a taxonomy of research contributions and introduce a task of automatically identifying contribution statements and classifying their types from NLP research papers. |
| Outcome: | The proposed model analyzes 29k NLP research papers to understand their contributions . |
Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess Hypotheses (2020.acl-main)
Copied to clipboard
| Challenge: | Empirical research in natural language processing has adopted a narrow set of principles for assessing hypotheses . alternative approaches to assess hypothese rely on p-value computation, which suffers from several known issues. |
| Approach: | They propose to compare different methods for assessing hypotheses . they argue that practitioners should first decide their target hypothesis before choosing a method . |
| Outcome: | The proposed method differs from other methods, but is not widely used in NLP . the proposed method is based on a p-value computation, but has a small gap in accuracy . |
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)
Copied to clipboard
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos, Vivek Srikumar, Nicholas Rizzolo, Lev Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew, Zhili Feng, John Wieting, Xiaodong Yu, Yangqiu Song, Shashank Gupta, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth
| Challenge: | a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks. |
| Approach: | They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community . |
| Outcome: | The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges. |