Papers by Hieu Man

9 papers
ULLME: A Unified Framework for Large Language Model Embeddings with Generation-Augmented Learning (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing frameworks for large language model embeddings have limited support for only a limited range of architectures and fine-tuning strategies.
Approach: They propose a framework that enables bidirectional attention across various LLMs and supports a range of fine-tuning strategies.
Outcome: The proposed framework enables bidirectional attention across various LLMs and supports a range of fine-tuning strategies.
Hierarchical Selection of Important Context for Generative Event Causality Identification with Optimal Transports (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for Event Causality Identification (ECI) rely on external toolkits or human annotation to obtain training signals.
Approach: They propose a generative framework that leverages Optimal Transport to automatically select the most important sentences and words from full documents.
Outcome: The proposed framework can predict causal relation between two events in text without external tools.
Multilingual SubEvent Relation Extraction: A Novel Dataset and Structure Induction Method (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for subevent relation extraction (SRE) focus on sequential order of words in texts to enhance representation learning.
Approach: They propose a method that learns to induce effective graph structures for input texts . they use word alignment frameworks with dependency paths and optimal transport .
Outcome: The proposed method is able to induce effective graph structures for input texts to boost representation learning.
Event Causality Identification via Generation of Important Context Words (2022.starsem-1)

Copied to clipboard

Challenge: Prior work focused on identifying causal relation between two event mentions . current models do not output important contexts for causal prediction of two mentions.
Approach: They propose to use dependency path generation as a complementary task for ECI.
Outcome: The proposed model can generate both causal relation and dependency path words from input sentences.
Reasoning with Memory: Adaptive Information Management for Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Multi-hop reasoning remains a fundamental challenge for Retrieval-Augmented Generation systems.
Approach: They propose a framework that provides a dynamic cognitive workspace for multi-hop reasoning . it uses an explicit working memory that persists across retrieval cycles and is continuously updated .
Outcome: The proposed framework achieves state-of-the-art performance over existing systems on eight QA benchmarks.
Contextualized Soft Prompts for Extraction of Event Arguments (2023.findings-acl)

Copied to clipboard

Challenge: Existing prompt-based methods for event argument extraction rely on discrete and manually-designed prompts that cannot exploit specific context for each example.
Approach: They propose a prompt-based method that introduces soft prompts to facilitate encoding of individual example context and multiple relevant documents to boost EAE.
Outcome: The proposed method extensively evaluates on benchmark datasets to demonstrate its benefits with state-of-the-art performance.
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing training datasets for large language models are often not fully disclosed.
Approach: They propose a multilingual dataset with 6.3 trillion tokens in 167 languages . they use a pipeline of multiple stages to achieve the best quality for model training .
Outcome: The proposed dataset is cleaned and deduplicated to achieve the best quality for model training . lack of transparency has hindered research on attributing and addressing hallucination and bias issues . 6.3 trillion tokens in 167 languages are used to train multilingual LLMs .
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have led to significant breakthroughs in the field.
Approach: They evaluate ChatGPT over multiple tasks with diverse languages and large datasets to provide more comprehensive information for multilingual NLP applications.
Outcome: The proposed model can process and generate texts for multiple languages due to its multilingual training data.
Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI (2026.acl-long)

Copied to clipboard

Challenge: Existing methods struggle with content-style entanglement, leading to poor generalization across domains.
Approach: They propose an explanation-by-design framework that explicitly disentangles style from content through architectural separation-by design.
Outcome: The proposed framework disentangles style from content through architectural separation-by-design.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations