Papers with linguistics
LingConv: An Interactive Toolkit for Controlled Paraphrase Generation with Linguistic Attribute Control (2025.emnlp-demos)
Copied to clipboard
| Challenge: | LINGCONV is an interactive toolkit for controllable text generation . it allows fine-grained control over 40 specific linguistic attributes spanning lexical, syntactic, and discourse dimensions. |
| Approach: | They propose a toolkit for paraphrase generation that allows finegrained control over 40 specific linguistic attributes. |
| Outcome: | The toolkit is available at https://mohdelgaar-lingconv.hf.space, with a demo video at https:youtu.be/wRBJEJ6EALQ. |
Inducing Grammar from Long Short-Term Memory Networks by Shapley Decomposition (2020.acl-srw)
Copied to clipboard
| Challenge: | a recent study shows that modern neural networks understand sentences implicitly by inducing recursive structures. |
| Approach: | They propose to explicitly induce grammar by tracing the computational process of a long short-term memory network. |
| Outcome: | The proposed model can explicitly induce grammar without external knowledge . tracing the computational process of a long short-term memory network is shown to be effective . |
Lost and Found: Computational Quality Assurance of Crowdsourced Knowledge on Morphological Defectivity in Wiktionary (2025.acl-srw)
Copied to clipboard
| Challenge: | a recent study shows that wikis are not reliable for linguistic knowledge of defects in understudied languages. |
| Approach: | They customize a neural morphological analyzer to annotate Latin and Italian corpora . they validated morphology using crowd-sourced data from Wiktionary to find defects . |
| Outcome: | The proposed algorithm annotates Latin and Italian corpora using crowd-sourced data . results show that 7% of Latin lemmata listed as defective show strong corpus evidence of being non-defective. |
LiViTo: Linguistic and Visual Features Tool for Assisted Analysis of Historic Manuscripts (2020.lrec-1)
Copied to clipboard
| Challenge: | a mixed methods approach is feasible for the identification of scribes and authors in handwritten documents. |
| Approach: | They propose a mixed methods approach to the identification of scribes and authors in handwritten documents . they use a software tool which combines linguistic insights and computer vision techniques . |
| Outcome: | The proposed tool can be used to identify scribes and authors in handwritten documents. |
Syntax-guided Contrastive Learning for Pre-trained Language Model (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing studies rely on additional syntax-driven attention components to enhance the transformer, which require more parameters and additional syntactic parsing in downstream tasks. |
| Approach: | They propose a syntax-guided contrastive learning method which does not change the transformer architecture and does not alter the transformer structure. |
| Outcome: | The proposed method achieves consistent improvements in a variety of tasks including grammatical error detection, entity tasks, structural probing and GLUE. |
FaiMA: Feature-aware In-context Learning for Multi-domain Aspect-based Sentiment Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for aspect-based sentiment analysis are limited and integrating with existing techniques is difficult. |
| Approach: | They propose a framework that utilizes in-context learning as a feature-aware mechanism that facilitates adaptive learning in multi-domain ABSA tasks. |
| Outcome: | The proposed framework achieves significant performance improvements in multiple domains compared to baselines, increasing F1 by 2.07% on average. |
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English (2024.lrec-main)
Copied to clipboard
| Challenge: | Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research. |
| Approach: | They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse. |
| Outcome: | The proposed dataset is the largest of its kind that is publicly accessible. |