Papers by Wang-Chiew Tan

8 papers
OpinionDigest: A Simple Framework for Opinion Summarization (2020.acl-main)

Copied to clipboard

Challenge: Abstractive opinion summarization framework outperforms competitors' summarizing frameworks . extractive approaches produce well-formed text, but selecting the most popular opinions is challenging .
Approach: They propose an abstractive opinion summarization framework that trains a Transformer model to reconstruct reviews from extracted opinions.
Outcome: The proposed framework outperforms baselines on Yelp and shows promising customization capabilities.
Convex Aggregation for Opinion Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in text autoencoders have significantly improved the quality of the latent space, allowing models to generate consistent text from aggregated latent vectors.
Approach: They develop a framework which searches input-output word overlap for latent vector aggregation.
Outcome: The proposed framework improves the quality of the latent space and establishes state-of-the-art performance on two opinion summarization benchmarks.
Essentia: Mining Domain-specific Paraphrases with Word-Alignment Graphs (D19-53)

Copied to clipboard

Challenge: Existing methods for mining general-purpose paraphrases are often based on statistical methods, but domain-specific corpora are too small to fit statistical methods.
Approach: They propose a method to mine paraphrases from a small set of sentences that roughly share the same topic or intent.
Outcome: The proposed method obtains high quality paraphrases as evaluated by crowd workers.
TimelineQA: A Benchmark for Question Answering over Timelines (2023.findings-acl)

Copied to clipboard

Challenge: Existing question answering techniques for lifelogs do not provide accurate answers . augmented reality glasses have led to the creation of personal assistants .
Approach: They propose to use a benchmark to query lifelogs to find out what happened in real life . they find that extractive QA systems out-perform retrieval-augmented QA techniques .
Outcome: The proposed method outperforms state-of-the-art retrieval-augmented QA systems in atomic queries and multi-hop queries.
SubjQA: A Dataset for Subjectivity and Review Comprehension (2020.emnlp-main)

Copied to clipboard

Challenge: Subjectivity is the expression of internal opinions or beliefs which cannot be objectively observed or verified.
Approach: They develop a dataset which investigates subjectivity in question answering . they find that subjectivity is an important feature in the case of QA .
Outcome: The proposed dataset shows that subjectivity is an important feature in question answering (QA) it also shows that subjective questions and answers can have more complex interactions than previously thought.
HappyDB: A Corpus of 100,000 Crowdsourced Happy Moments (L18-1)

Copied to clipboard

Challenge: Recent research has focused on developing technologies that help users incorporate the findings of the science of happiness into their daily lives.
Approach: They crowd-sourced HappyDB, a corpus of 100,000 happy moments, and applied several state-of-the-art analysis techniques to analyze HappyDB.
Outcome: The proposed technology can understand how people express their happy moments in text and analyze them using state-of-the-art techniques.
Open Information Extraction from Question-Answer Pairs (N19-1)

Copied to clipboard

Challenge: Existing work on OpenIE extracts structured data from sentences . a system for extracting tuples from question-answer pairs solves this problem .
Approach: They propose a system for extracting tuples from question-answer pairs . they use distributed representations of a question and an answer to generate knowledge facts .
Outcome: The proposed system extracts meaningful structured tuples from question-answer pairs . it can find new and interesting facts to extend knowledge bases, the authors show .
Reimagining Retrieval Augmented Language Models for Answering Queries (2023.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are expensive to train, deploy, and maintain, both financially and in terms of environmental impact.
Approach: They present a reality check on large language models and compare their predictions to retrieval-augmented language models.
Outcome: The proposed models fare better on question answering tasks and have become the foundation of impressive demos like Chat-GPT.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations