Papers by Shiyu Hong

5 papers
Self-Supervised Learning for Contextualized Extractive Summarization (P19-1)

Copied to clipboard

Challenge: Existing models for extractive summarization are usually trained from scratch with a cross-entropy loss . previous work builds an end-to-end system to learn to choose sentences without explicitly modeling document context .
Approach: They propose three auxiliary pre-training tasks that learn to capture the document context in a self-supervised fashion.
Outcome: The proposed models outperform existing models on a CNN/DM dataset.
Simple yet Effective Bridge Reasoning for Open-Domain Multi-Hop Question Answering (D19-58)

Copied to clipboard

Challenge: Existing work on open-domain multi-hop question answering relies on off-the-shelf information retrieval techniques to retrieve answer passages.
Approach: They propose a new subproblem for open-domain multi-hop question answering . they aim to recognize the anchor from a set of start passages with a reading comprehension model .
Outcome: The proposed method significantly improves the baseline method on the open-domain hotpotQA benchmark.
Sentence Embedding Alignment for Lifelong Relation Extraction (N19-1)

Copied to clipboard

Challenge: Existing approaches to relation extraction require a fixed set of relations . Existing methods assume a closed set of relationships and perform once-and-for-all training on a set of datasets.
Approach: They propose to improve the stochastic gradient methods with a replay memory to alleviate the forgetting problem by anchoring the sentence embedding space.
Outcome: The proposed method outperforms state-of-the-art methods on multiple benchmarks.
HiCuLR: Hierarchical Curriculum Learning for Rhetorical Role Labeling of Legal Documents (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches overlook the varying difficulty levels inherent in legal document discourse styles and rhetorical roles.
Approach: They propose a hierarchical curriculum learning framework for RRL that nests two curricula: Rhetorical Role-level Curriculum (RC) on the outer layer and Document-level curriculum (DC) on inner layer.
Outcome: The proposed framework is based on four legal document datasets and shows that it is complementary to existing models.
TWEETQA: A Social Media Focused Question Answering Dataset (P19-1)

Copied to clipboard

Challenge: Social media is becoming an important realtime information source, especially during natural disasters and emergencies.
Approach: They present a large-scale dataset for question answering over social media data . they gather tweets used by journalists and ask human annotators to write questions upon them .
Outcome: The proposed dataset shows that neural models that perform well on formal texts are limited in their performance . the proposed model is still lagging behind human performance with a large margin .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations