Papers with categorization
SanskritShala: A Neural Sanskrit NLP Toolkit with Web-Based Interface for Pedagogical and Annotation Purposes (2023.acl-demo)
Copied to clipboard
| Challenge: | SanskritShala is a neural-based Sanskrit NLP toolkit that is available as a web-based application . |
| Approach: | They propose a neural Sanskrit NLP toolkit that facilitates linguistic analyses for word segmentation, morphological tagging, dependency parsing, and compound type identification. |
| Outcome: | The proposed toolkit reports state-of-the-art performance on benchmark datasets . it is built with easy-to-use interactive data annotation features . |
Assessing How Users Display Self-Disclosure and Authenticity in Conversation with Human-Like Agents: A Case Study of Luda Lee (2022.findings-aacl)
Copied to clipboard
| Challenge: | Existing studies on how people interact with conversational agents have not investigated the interaction authenticity of human-like agents. |
| Approach: | They construct a taxonomy to discern the users’ self-disclosure in the dialogue and the communication authenticity displayed in the user posting. |
| Outcome: | The proposed taxonomy can be used for future research and industrial development. |
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing approaches to extract actionable suggestions from customer reviews are often mixed-intent, unstructured text. |
| Approach: | They propose a hybrid pipeline that uses a RoBERTa classifier and a precision–recall surrogate to extract actionable suggestions from customer reviews. |
| Outcome: | The proposed pipeline outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence. |
HFT-CNN: Learning Hierarchical Category Structure for Multi-label Short Text Categorization (D18-1)
Copied to clipboard
| Challenge: | Existing methods for categorization of short texts use non-hierarchical flat model, but they are limited by domain-independent knowledge distribution. |
| Approach: | They propose a method which leverages hierarchical relationships between pre-defined categories to tackle the data sparsity problem. |
| Outcome: | The proposed method is competitive with the state-of-the-art methods on a multi-label categorization task for short texts using two benchmark datasets. |
Casting Light on Invisible Cities: Computationally Engaging with Literary Criticism (N19-1)
Copied to clipboard
| Challenge: | Literary critics often attempt to uncover meaning in a single work of literature through careful reading and analysis. |
| Approach: | They propose to use a literary theory to analyze Italo Calvino's novel Invisible Cities to leverage contextualized representations to embed each city's description and use unsupervised methods to cluster embeddings. |
| Outcome: | The proposed method can be applied to Italo Calvino’s novel Invisible Cities . authors compare results to similarity judgments generated by human readers . |
Life is a Circus and We are the Clowns: Automatically Finding Analogies between Situations and Processes (2022.emnlp-main)
Copied to clipboard
| Challenge: | Analogy-making gives rise to reasoning, abstraction, flexible categorization and counterfactual inference – abilities that current AI systems lack. |
| Approach: | They propose an interpretable, scalable algorithm that extracts analogies from a pair of natural language procedural texts and finds a mapping between the different domains based on relational similarity. |
| Outcome: | The proposed algorithm can extract analogies from a large dataset and achieve 79% precision. |
Love Me, Love Me, Say (and Write!) that You Love Me: Enriching the WASABI Song Corpus with Lyrics Annotations (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of songs enriched with metadata extracted from music databases on the Web contains 1.73M songs with lyrics (1.41M unique lyrics) a researcher proposes methods to extract relevant information from lyrics, including their structure segmentation, topic, explicitness of lyrics content, salient passages of a song and emotions conveyed. |
| Approach: | They propose to extract relevant information from lyrics by using music databases . they propose to use metadata extracted from music databases to analyze lyrics . |
| Outcome: | The proposed methods can be exploited by music search engines and music professionals to better handle large collections of lyrics. |
Disentangling Categorization in Multi-agent Emergent Communication (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent work on the emergence of language between artificial agents has not isolated the effect of categorization power on inter-communication ability. |
| Approach: | They propose to use disentangled representations to quantify categorization power of agents to enable differential analysis between combinations of heterogeneous systems. |
| Outcome: | The proposed method reduces signaling accuracy by 40% despite encouraging compositionality in the artificial language. |
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | This work presents a corpus manually annotated with named entities for six Slavic languages . |
| Approach: | They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models . |
| Outcome: | The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier. |
A Survey on Patent Analysis: From NLP to Multimodal AI (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in pretrained language models and large language models have demonstrated transformative capabilities across diverse domains. |
| Approach: | They propose a taxonomy for categorization based on tasks in the patent life cycle . they introduce a novel taxonomies for categorizing based upon tasks in patent life cycles . |
| Outcome: | The proposed method is based on tasks in the patent life cycle and provides a taxonomy for categorization based upon tasks in patent life cycles. |
RoBERT2VecTM: A Novel Approach for Topic Extraction in Islamic Studies (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a new approach to investigate “Hadith” texts presents challenges due to the complexity of Arabic . a novel neural-based approach to analyze “Matn” topics outperforms traditional NLP models . |
| Approach: | They propose a novel approach to analyze Arabic “Hadith” texts using the Contextualized Topic Model. |
| Outcome: | The proposed approach outperforms state-of-the-art models by generating more coherent topics in Arabic. |
Pseudonymization Categories across Domain Boundaries (2024.lrec-main)
Copied to clipboard
Maria Irena Szawerna, Simon Dobnik, Therese Lindström Tiedemann, Ricardo Muñoz Sánchez, Xuan-Son Vu, Elena Volodina
| Challenge: | Linguistic data can contain personal information, which is limited in accessibility . a universal system of tags for categorizing PIIs could be developed to replace them . |
| Approach: | They analyze tagsets used for anonymization and pseudonymization to find out what kinds of PII appear in different domains. |
| Outcome: | The proposed system would allow for dynamic pseudonymization while keeping the data readable and useful for future research. |