Challenge: Existing methods for style transfer between Singlish and Standard English lack explainability and fine-grained control.
Approach: They propose a multi-agent framework where large language models act as expert agents for each linguistic aspect.
Outcome: The proposed model enables precise, interpretable transformations, advancing explainability in NLP for Singlish.

Similar Papers

Singlish Where Got Rules One? Constructing a Computational Grammar for Singlish (2022.lrec-1)

Copied to clipboard

Challenge: Singlish is a variety of English spoken in Singapore and has many non-standard features.
Approach: They propose to use Singlish as a branch of English grammar to implement new rules and add new lexical types to it.
Outcome: The proposed grammar is based on the existing rules and lexical types from the English resource grammar and compared with the standard English grammar.
Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer (N18-1)

Copied to clipboard

Challenge: a lack of training and evaluation datasets, benchmarks and automatic metrics has blocked progress in this field.
Approach: They propose to use a grammarly's Yahoo Answers Formality corpus to create the largest corpus for a particular style . they also propose to apply machine translation metrics to the task .
Outcome: The proposed model can be used to train and evaluate a text in a particular style . the proposed model is based on the existing model and can be applied to other tasks .
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to transfer text style focus on sentence-level data, limiting performance . current LLMs struggle to generate public speaking texts that align with human preferences .
Approach: They propose a task to transform official texts into public-speaking styles by analyzing real-world data.
Outcome: The proposed task aims to transform public speaking texts into public-speaking styles . the proposed framework analyzes characteristics and identifies problems of stylized texts .
Towards Robust and Semantically Organised Latent Representations for Unsupervised Text Style Transfer (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies show that auto-encoders perform language generation, smooth sentence interpolation, and style transfer over unseen attributes using unlabelled datasets in a zero-shot manner.
Approach: They propose a discrete token-based perturbation approach to map "similar" sentences close by in latent space.
Outcome: The proposed model can generate and perform language generation, style transfer and sentence interpolation tasks on unlabelled datasets in a zero-shot manner.
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Current ConvXAI systems are based on intent recognition to accurately identify the user’s desired intention and map it to an explainability method.
Approach: They propose a multilingual extension of the CoXQL dataset spanning five typologically diverse languages, including one low-resource language.
Outcome: The proposed model enables multilingual generalization in a multilingual dataset spanning five typologically diverse languages, including one low-resource language.
Formality Style Transfer for Noisy, User-generated Conversations: Extracting Labeled, Parallel Data from Unlabeled Corpora (D19-55)

Copied to clipboard

Challenge: Typical datasets used for style transfer in NLP contain aligned pairs of two opposite extremes of a style.
Approach: They propose a technique to derive a dataset of aligned pairs from an unlabeled corpus by using an auxiliary dataset, allowing for in-domain training.
Outcome: The proposed method significantly outperforms OpenNMT’s Seq2Seq model trained on the Yahoo Formality Dataset and 6 novel datasets.
Comparing Styles across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Communication practices vary across cultures. Inherent differences in how people think and behave influence cultural norms.
Approach: They propose a framework to extract stylistic differences from multilingual language models (LMs) they use a multilingual lexica to consolidate feature importances into comparable lexical categories .
Outcome: The proposed framework generates comprehensive style lexica in any language and consolidates feature importances from LMs into comparable lexical categories.
Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across Modalities (2026.acl-long)

Copied to clipboard

Challenge: Amidst the rapid advances of large language models, most LLMs struggle with mixed-language inputs, limited Code-switching datasets, and evaluation biases.
Approach: They propose a roadmap for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual intelligence.
Outcome: The proposed frameworks are based on 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages.
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore’s Low-Resource Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have transformed natural language processing, but their safety mechanisms remain under-explored in low-resource, multilingual settings.
Approach: They propose a red-teaming approach to probe LLM vulnerabilities in Singapore's diverse linguistic context using a dataset and evaluation framework.
Outcome: The proposed framework systematically probes LLM vulnerabilities in three real-world scenarios including Singlish, Chinese, Malay, and Tamil.
TyDiP: A Dataset for Politeness Classification in Nine Typologically Diverse Languages (2022.findings-emnlp)

Copied to clipboard

Challenge: Whether politeness phenomena and strategies are universal across languages or not have been controversial among sociologists and linguists.
Approach: They create a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples.
Outcome: The proposed model shows a robust zero-shot transfer ability, but falls short of estimated human accuracy significantly.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations