| Challenge: | despite its central role, the notion of text in natural language processing is vague, authors argue . a conceptual framework for capturing text differences is lacking, authors say . authors propose a two-tier taxonomy of linguistic and non-linguistic elements available in textual sources . |
| Approach: | They propose a taxonomy of linguistic and non-linguistic elements available in textual sources and can be used in NLP modeling. |
| Outcome: | The proposed taxonomy examines the production and transformation of textual data . it outlines key desiderata and challenges of the emerging inclusive approach to text in NLP . |
Similar Papers
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
Translational NLP: A New Paradigm and General Principles for Natural Language Processing Research (2021.naacl-main)
Copied to clipboard
| Challenge: | Natural language processing research is often assumed to emerge naturally . many innovations go unapplied and important questions remain unstudied . |
| Approach: | They propose a new paradigm to structure and facilitate the processes by which basic and applied NLP research inform one another. |
| Outcome: | The proposed framework provides a roadmap for developing Translational NLP as a dedicated research area. |
Putting Natural in Natural Language Processing (2023.findings-acl)
Copied to clipboard
| Challenge: | human language is firstly spoken and only secondarily written. |
| Approach: | aaron carroll: human language is firstly spoken and only secondarily written . carroll says the field of NLP has overwhelmingly focused on processing written language . he says the focus is on a subset of human language which is convenient to work with . |
| Outcome: | the ACL 2023 theme track urges the community to check the reality of the progress in NLP . |
Returning the N to NLP: Towards Contextually Personalized Classification Models (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study shows that NLP models treat language as universal, but that it is based on sociolinguistic research. |
| Approach: | They propose to incorporate user-dependent, contextual personal and social aspects into neural NLP models by means of socially contextual personalization. |
| Outcome: | The proposed approach could be adapted to better personalize the language of users . it outlines a possible direction to incorporate these aspects into neural NLP models . |
A Systematic Review of Reproducibility Research in Natural Language Processing (2021.eacl-main)
Copied to clipboard
| Challenge: | Despite the recent progress in reproducibility, the field is far from reaching a consensus on how reproducibility should be defined, measured and addressed. |
| Approach: | They propose to provide a wide-angle snapshot of current work on reproducibility in NLP. |
| Outcome: | The proposed work will provide a wide-angle snapshot of current work on reproducibility in NLP. |
Learning with Limited Text Data (2022.acl-tutorials)
Copied to clipboard
| Challenge: | Natural Language Processing (NLP) relies on labeled data to perform state-of-the-art performance . labeles are often required to label large amounts of textual data . this tutorial will provide an overview of labeleing in NLP . |
| Approach: | This tutorial will provide a systematic overview of methods for learning from limited labeled data. |
| Outcome: | This tutorial will provide a systematic and up-to-date overview of the proposed methods . it will highlight current challenges and future directions . |
Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender (2022.coling-1)
Copied to clipboard
| Challenge: | Current modeling of 3rd person pronouns ignores neopronoun phenomena like naive pronounes, which are not (yet) widely established. |
| Approach: | They propose to validate existing and novel approaches for modeling 3rd person pronouns in language technology and validate them through a survey. |
| Outcome: | The proposed model excludes non-binary users, while ignoring gender-specific phenomena. |
From Text to Context: Contextualizing Language with Humans, Groups, and Communities for Socially Aware NLP (2024.naacl-tutorials)
Copied to clipboard
Adithya V Ganesan, Siddharth Mangalik, Vasudha Varadarajan, Nikita Soni, Swanie Juhng, João Sedoc, H. Andrew Schwartz, Salvatore Giorgi, Ryan L Boyd
| Challenge: | This tutorial will cover the latest techniques and libraries for doing so at each level of analysis. |
| Approach: | This tutorial will cover the latest techniques and libraries for doing so at each level of analysis. |
| Outcome: | The tutorial covers human-centered techniques that provide benefit to traditional document- or word-level NLP tasks. |
The Importance of Modeling Social Factors of Language: Theory and Practice (2021.naacl-main)
Copied to clipboard
| Challenge: | Current NLP models focus on information content while ignoring language’s social factors. |
| Approach: | They propose that NLP systems focus on information content while ignoring language’s social factors to improve performance. |
| Outcome: | The proposed approach improves the performance of existing systems, open up new applications, and increase fairness and usability for all users. |
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging. |
| Approach: | They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned . |
| Outcome: | The proposed methods are compared with existing models and compare them with existing ones. |