Papers by Shuhan Sun

4 papers
A Self-verified Method for Exploring Simile Knowledge from Pre-trained Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) have succeeded in natural language processing because they learn generic knowledge from a large corpus.
Approach: They propose a method that allows pre-trained language models to explore simile knowledge from PLMs . they enhance PLM models with a multi-level simile recognition task that evaluates similes aplenty .
Outcome: The proposed method can explore more accurate simile knowledge for PLMs.
Investigating Value-Reasoning Reliability in Small Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: sLLMs have been widely deployed in practical applications, but little attention has been paid to their value-reasoning abilities, particularly in terms of reasoning reliability.
Approach: They propose a systematic evaluation framework for assessing the Value-Reasoning Reliability of small Large Language models (sLLMs) . framework includes three core tasks: Repetition Consistency task, Interaction Stability task, and Open-ended Expression Consistencies task.
Outcome: The proposed framework incorporates self-reported confidence scores to evaluate the model’s value reasoning reliability from two perspectives: the model's self awareness of its values, and its value-based decision-making.
Efficient Knowledge Infusion via KG-LLM Alignment (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for knowledge infusion face knowledge mismatch and poor information compliance of LLMs with knowledge graphs.
Approach: They propose a three-stage alignment strategy to enhance the LLM's capability to utilize information from knowledge graphs.
Outcome: The proposed method outperforms baselines on biomedical question-answering datasets and outperformed existing methods.
I run as fast as a rabbit, can you? A Multilingual Simile Dialogues Datasets (2023.findings-acl)

Copied to clipboard

Challenge: A simile is a figure of speech that compares two different things via shared properties.
Approach: They propose a multilingual simile dialogue dataset that can be used to study similes in real-life scenarios.
Outcome: The proposed dataset is the largest manually annotated simile dataset and contains both English and Chinese data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations