Papers with categorization

12 papers
SanskritShala: A Neural Sanskrit NLP Toolkit with Web-Based Interface for Pedagogical and Annotation Purposes (2023.acl-demo)

Copied to clipboard

Challenge: SanskritShala is a neural-based Sanskrit NLP toolkit that is available as a web-based application .
Approach: They propose a neural Sanskrit NLP toolkit that facilitates linguistic analyses for word segmentation, morphological tagging, dependency parsing, and compound type identification.
Outcome: The proposed toolkit reports state-of-the-art performance on benchmark datasets . it is built with easy-to-use interactive data annotation features .
Assessing How Users Display Self-Disclosure and Authenticity in Conversation with Human-Like Agents: A Case Study of Luda Lee (2022.findings-aacl)

Copied to clipboard

Challenge: Existing studies on how people interact with conversational agents have not investigated the interaction authenticity of human-like agents.
Approach: They construct a taxonomy to discern the users’ self-disclosure in the dialogue and the communication authenticity displayed in the user posting.
Outcome: The proposed taxonomy can be used for future research and industrial development.
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews (2026.eacl-industry)

Copied to clipboard

Challenge: Existing approaches to extract actionable suggestions from customer reviews are often mixed-intent, unstructured text.
Approach: They propose a hybrid pipeline that uses a RoBERTa classifier and a precision–recall surrogate to extract actionable suggestions from customer reviews.
Outcome: The proposed pipeline outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence.
HFT-CNN: Learning Hierarchical Category Structure for Multi-label Short Text Categorization (D18-1)

Copied to clipboard

Challenge: Existing methods for categorization of short texts use non-hierarchical flat model, but they are limited by domain-independent knowledge distribution.
Approach: They propose a method which leverages hierarchical relationships between pre-defined categories to tackle the data sparsity problem.
Outcome: The proposed method is competitive with the state-of-the-art methods on a multi-label categorization task for short texts using two benchmark datasets.
Casting Light on Invisible Cities: Computationally Engaging with Literary Criticism (N19-1)

Copied to clipboard

Challenge: Literary critics often attempt to uncover meaning in a single work of literature through careful reading and analysis.
Approach: They propose to use a literary theory to analyze Italo Calvino's novel Invisible Cities to leverage contextualized representations to embed each city's description and use unsupervised methods to cluster embeddings.
Outcome: The proposed method can be applied to Italo Calvino’s novel Invisible Cities . authors compare results to similarity judgments generated by human readers .
Life is a Circus and We are the Clowns: Automatically Finding Analogies between Situations and Processes (2022.emnlp-main)

Copied to clipboard

Challenge: Analogy-making gives rise to reasoning, abstraction, flexible categorization and counterfactual inference – abilities that current AI systems lack.
Approach: They propose an interpretable, scalable algorithm that extracts analogies from a pair of natural language procedural texts and finds a mapping between the different domains based on relational similarity.
Outcome: The proposed algorithm can extract analogies from a large dataset and achieve 79% precision.
Love Me, Love Me, Say (and Write!) that You Love Me: Enriching the WASABI Song Corpus with Lyrics Annotations (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of songs enriched with metadata extracted from music databases on the Web contains 1.73M songs with lyrics (1.41M unique lyrics) a researcher proposes methods to extract relevant information from lyrics, including their structure segmentation, topic, explicitness of lyrics content, salient passages of a song and emotions conveyed.
Approach: They propose to extract relevant information from lyrics by using music databases . they propose to use metadata extracted from music databases to analyze lyrics .
Outcome: The proposed methods can be exploited by music search engines and music professionals to better handle large collections of lyrics.
Disentangling Categorization in Multi-agent Emergent Communication (2022.naacl-main)

Copied to clipboard

Challenge: Recent work on the emergence of language between artificial agents has not isolated the effect of categorization power on inter-communication ability.
Approach: They propose to use disentangled representations to quantify categorization power of agents to enable differential analysis between combinations of heterogeneous systems.
Outcome: The proposed method reduces signaling accuracy by 40% despite encouraging compositionality in the artificial language.
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)

Copied to clipboard

Challenge: This work presents a corpus manually annotated with named entities for six Slavic languages .
Approach: They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models .
Outcome: The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier.
A Survey on Patent Analysis: From NLP to Multimodal AI (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in pretrained language models and large language models have demonstrated transformative capabilities across diverse domains.
Approach: They propose a taxonomy for categorization based on tasks in the patent life cycle . they introduce a novel taxonomies for categorizing based upon tasks in patent life cycles .
Outcome: The proposed method is based on tasks in the patent life cycle and provides a taxonomy for categorization based upon tasks in patent life cycles.
RoBERT2VecTM: A Novel Approach for Topic Extraction in Islamic Studies (2024.findings-emnlp)

Copied to clipboard

Challenge: a new approach to investigate “Hadith” texts presents challenges due to the complexity of Arabic . a novel neural-based approach to analyze “Matn” topics outperforms traditional NLP models .
Approach: They propose a novel approach to analyze Arabic “Hadith” texts using the Contextualized Topic Model.
Outcome: The proposed approach outperforms state-of-the-art models by generating more coherent topics in Arabic.
Pseudonymization Categories across Domain Boundaries (2024.lrec-main)

Copied to clipboard

Challenge: Linguistic data can contain personal information, which is limited in accessibility . a universal system of tags for categorizing PIIs could be developed to replace them .
Approach: They analyze tagsets used for anonymization and pseudonymization to find out what kinds of PII appear in different domains.
Outcome: The proposed system would allow for dynamic pseudonymization while keeping the data readable and useful for future research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations