Papers by Can Udomcharoenchaikit

16 papers
Towards Better Understanding of Program-of-Thought Reasoning in Cross-Lingual and Multilingual Environments (2025.findings-acl)

Copied to clipboard

Challenge: Multi-step reasoning is essential for large language models, yet multilingual performance remains challenging.
Approach: They propose a framework to evaluate Program-of-Thought (PoT) prompting by separating multilingual reasoning from code execution to examine impact of fine-tuning on question-reasoning alignment and reasoning quality.
Outcome: The proposed framework outperforms CoT fine-tuned models in multilingual settings and shows strong correlation between reasoning quality and answer accuracy.
Mitigating Spurious Correlation in Natural Language Understanding with Counterfactual Inference (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to debias NLU models rely on superficial patterns to produce correct predictions . lexical overlap and annotation artifacts can be used to make shortcuts .
Approach: They propose a causal analysis framework to help debias NLU models by defining causal relationships and utilizing counterfactual inference to mitigate bias.
Outcome: The proposed framework can improve robustness across three NLU tasks while maintaining high in-distribution performance.
McCrolin: Multi-consistency Cross-lingual Training for Retrieval Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches struggle with consistency across multiple languages and multi-size input scenarios.
Approach: They propose a cross-lingual training framework that leverages multi-task learning to enhance cross-linguistic consistency and ranking stability.
Outcome: The proposed training framework outperforms competitors on various input sizes and architectures.
An Empirical Study of Multilingual Reasoning Distillation for Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing efforts to distill reasoning capabilities have focused mainly on English, leaving multilingual distillation underexplored.
Approach: They propose a method that incorporates incorrect rationales as additional guidance to improve multilingual reasoning in large language models.
Outcome: Empirical results show that d-CoT-nR significantly surpasses the baseline, improving accuracy in unseen languages and correctness in step-by-step reasoning.
ConGen: Unsupervised Control and Generalization Distillation For Sentence Representation (2022.findings-emnlp)

Copied to clipboard

Challenge: Sentence representations are essential in many NLP tasks operating at the sentence level.
Approach: They propose an unsupervised sentence representation method to reduce the supervised-unsupervised performance gap for smaller models.
Outcome: The proposed method outperforms supervised training on STS, text classification, and natural language inference tasks on smaller models.
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai (2024.acl-srw)

Copied to clipboard

Challenge: Xue et al., 2024) have demonstrated that large language models can perform at human level across multitudes of tasks and domains.
Approach: They propose a seed-free framework for generating synthetic instruction-tuning data that incorporates fluency, diversity, and cultural context.
Outcome: The proposed framework achieves competitive performance using only 5,000 instructions compared to state-of-the-art Thai LLMs trained on hundreds of thousands of instructions.
Typo-Robust Representation Learning for Dense Retrieval (2023.acl-short)

Copied to clipboard

Challenge: Dense retrieval is a fundamental building block of information retrieval applications.
Approach: They propose a method that aligns misspelled queries with their pristine counterparts to improve contrast between each query and its surrounding queries.
Outcome: The proposed method outperforms the competitors in all cases with misspelled queries.
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia (2025.acl-long)

Copied to clipboard

Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz, Tack Hwa Wong, Mohammad Rifqi Farhansyah, Thant Thiri Maung, Frederikus Hudi, David Anugraha, Muhammad Ravi Shulthan Habibi, Muhammad Reza Qorib, Amit Agarwal, Joseph Marvin Imperial, Hitesh Laxmichand Patel, Vicky Feliren, Bahrul Ilmi Nasution, Manuel Antonio Rufino, Genta Indra Winata, Rian Adam Rajagede, Carlos Rafael Catalan, Mohamed Fazli Mohamed Imam, Priyaranjan Pattnayak, Salsabila Zahirah Pranida, Kevin Pratama, Yeshil Bangera, Adisai Na-Thalang, Patricia Nicole Monderin, Yueqi Song, Christian Simon, Lynnette Hui Xian Ng, Richardy Lobo Sapan, Taki Hasan Rafi, Bin Wang, null Supryadi, Kanyakorn Veerakanjana, Piyalitt Ittichaiwong, Matthew Theodore Roque, Karissa Vincentio, Takdanai Kreangphet, Phakphum Artkaew, Kadek Hendrawan Palgunadi, Yanzhi Yu, Rochana Prih Hastuti, William Nixon, Mithil Bangera, Adrian Xuan Wei Lim, Aye Hninn Khine, Hanif Muhammad Zhafran, Teddy Ferdinan, Audra Aurora Izzani, Ayushman Singh, Evan Evan, Jauza Akbar Krito, Michael Anugraha, Fenal Ashokbhai Ilasariya, Haochen Li, John Amadeo Daniswara, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Can Udomcharoenchaikit, Fadil Risdian Ansori, Mahardika Krisna Ihsani, Giang Nguyen, Anab Maulana Barik, Dan John Velasco, Rifo Ahmad Genadi, Saptarshi Saha, Chengwei Wei, Isaiah Edri W. Flores, Kenneth Chen Ko Han, Anjela Gail D. Santos, Wan Shen Lim, Kaung Si Phyo, Tim Santos, Meisyarah Dwiastuti, Jiayun Luo, Jan Christian Blaise Cruz, Ming Shan Hee, Ikhlasul Akmal Hanif, M.Alif Al Hakim, Muhammad Rizky Sya’ban, Kun Kerdthaisong, Lester James Validad Miranda, Fajri Koto, Tirana Noor Fatyanosa, Alham Fikri Aji, Jostin Jerico Rosal, Jun Kevin, Robert Wijaya, Onno P. Kampman, Ruochen Zhang, Börje F. Karlsson, Peerat Limkonchotiwat
Challenge: Southeast Asia is underrepresented in vision-language research . SEA-VL is an open-source initiative dedicated to developing culturally relevant datasets for SEA languages.
Approach: They propose to use crowdsourced, automated image crawling and synthetic image generation to develop culturally relevant datasets for SEA languages.
Outcome: The proposed datasets capture SEA cultural nuances and contexts better than existing datasets.
Identifying and Mitigating Annotation Bias in Natural Language Understanding using Causal Mediation Analysis (2024.findings-acl)

Copied to clipboard

Challenge: Current NLU models obtain state-of-the-art accuracy on in-distribution benchmarks, but they use annotation bias to make predictions, negatively affecting the models' generalizability.
Approach: They apply causal mediation analysis to gauge how much each component mediates annotation biases and use causal-grounded masking and gradient unlearning to mitigate bias.
Outcome: The proposed methods improve the model's robustness against annotation bias even after employing other training-time debiasing techniques.
Thai Nested Named Entity Recognition Corpus (2022.findings-acl)

Copied to clipboard

Challenge: a new dataset for Named Entity Recognition (NER) is proposed for Thailand.
Approach: They propose to use Thai N-NER to extract named entities from text . they propose to include a nested structure that can be used to improve NER .
Outcome: The proposed dataset is the largest non-English N-NER dataset and the first non- English one with fine-grained classes.
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models rely on translations, missing cultural and domain specificity.
Approach: They present a human-authored dataset for evaluation and instruction tuning in Thai . findings highlight need for culturally and professionally grounded instruction data .
Outcome: a human-authored dataset for evaluation and instruction tuning in Thai outperforms translation-based models . findings highlight need for culturally and professionally grounded instruction data .
Topic-Regularized Authorship Representation Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing techniques for authorship attribution have focused on out-of-distribution in topics or authors.
Approach: They propose a framework that creates authorship representation with reduced reliance on topic-specific information to handle a large number of unseen authors and topics.
Outcome: The proposed framework has improved over baselines in 4 out of 6 cases.
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning (2026.acl-short)

Copied to clipboard

Challenge: Existing methods to assess memorization in federated learning focus on one sample at a time . centralized learning does not eliminate the risk of memorizing large language models .
Approach: They propose a framework that quantifies both intra- and inter-client memorization in FL . they use fine-grained cross-sample memorisation measurement across all clients .
Outcome: The proposed framework quantifies both intra- and inter-client memorization in FL using fine-grained cross-sample memorisation measurement across all clients.
CL-ReLKT: Cross-lingual Language Knowledge Transfer for Multilingual Retrieval Question Answering (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to cross-lingual question answering use sentence embedding to map documents and questions in multiple languages . a novel cross-linguistic approach to cross language-retrieval question answering is proposed . our method outperforms competitors in 19 out of 21 settings of CL-ReQA .
Approach: They propose a cross-lingual language knowledge transfer framework for cross-linguistic question answering . they use a multilingual sentence embedding technique to create a linguistic embeddable space .
Outcome: The proposed method outperforms current state-of-the-art methods in 19 out of 21 settings of CL-ReQA.
MIST: Mutual Information Maximization for Short Text Clustering (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for clustering short texts are inadequate due to the limited amount of information provided by each text sample.
Approach: They propose a Mutual Information Maximization Framework for Short Text Clustering which maximizes mutual information between representations on sequence and token levels.
Outcome: The proposed framework outperforms the state-of-the-art method in terms of Accuracy or Normalized Mutual Information in most cases.
Efficient Overshadowed Entity Disambiguation by Mitigating Shortcut Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Entity disambiguation (ED) is crucial in natural language processing tasks such as question-answering and information extraction.
Approach: They propose a method to reduce computational overhead on overshadowed entities by addressing shortcut learning.
Outcome: The proposed method achieves state-of-the-art performance without compromising inference speed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations