Papers by Aaron Chan

5 papers
Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-text Rationales (2023.acl-long)

Copied to clipboard

Challenge: Existing metrics like task performance of the LM generating the rationales or similarity between generated and gold rationale are not good indicators of their human utility.
Approach: They propose to use a large language model to generate rationales with better human utility by estimating its conciseness and novelty.
Outcome: The proposed model can measure human utility to a better extent by estimating its usefulness in answering similar unseen instances.
Learning Contextualized Knowledge Structures for Commonsense Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: Recent knowledge graph (KG) augmented models have achieved notable success on commonsense reasoning tasks.
Approach: They propose a KG-augmented model that contextualizes extracted and generated knowledge by reasoning over both within a single graph structure.
Outcome: The proposed model outperforms existing models on four commonsense reasoning benchmarks and a user study on edge validness and helpfulness.
RESPROMPT: Residual Connection Prompting Advances Multi-Step Reasoning in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Chain-of-thought (CoT) has impressively unlocked the reasoning potential of large language models (LLMs), but it falls short when tackling problems that require multiple reasoning steps.
Approach: They propose a new prompting strategy that advances multi-step reasoning in LLMs by integrating necessary connections into prompts.
Outcome: The proposed strategy improves multi-step reasoning accuracy and improves reasoning accuracy across math, sequential, and commonsense domains.
XMD: An End-to-End Framework for Interactive Explanation-Based Debugging of NLP Models (2023.acl-demo)

Copied to clipboard

Challenge: Existing models are susceptible to learning spurious biases that do not reflect the underlying task.
Approach: They propose an open-source framework for explanation-based model debugging that allows users to provide various forms of feedback on model explanations.
Outcome: The proposed framework improves model’s OOD performance on text classification tasks by up to 18%.
ER-Test: Evaluating Explanation Regularization Methods for Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Explanation regularization (ER) aims to improve NLM generalization by pushing the NLM’s machine rationales to align with human rationale.
Approach: They propose a framework for evaluating ER models’ OOD generalization along three dimensions: unseen datasets, contrast set tests, and functional tests.
Outcome: The proposed framework evaluates ER models’ OOD generalization across unseen datasets, contrast set tests, and functional tests.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations