Papers by Shengjie Li

14 papers
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)

Copied to clipboard

Challenge: Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality .
Approach: aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers .
Outcome: a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward .
Conundrums in Cross-Prompt Automated Essay Scoring: Making Sense of the State of the Art (2024.acl-long)

Copied to clipboard

Challenge: Automated essay scoring (AES) is a task of assigning a single score to an essay . authors abandon sophisticated neural architectures and develop a simple feature-based approach .
Approach: a team of researchers develop a feature-based approach to cross-prompt automated essay scoring that adopts a simple neural architecture.
Outcome: a new approach to cross-prompt automated essay scoring can achieve state-of-the-art results.
Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing Process Reward Models (PRMs) are vulnerable to reward hacking and require expensive, large-scale annotation of reasoning steps.
Approach: They propose a reward model approach which evaluates both individual and consecutive reasoning steps from fine-grained and coarse-grounded level.
Outcome: Empirical results show that the proposed model performs better than existing PRMs and is more robust than existing models.
Summarizing Dialogues with Negative Cues (2022.coling-1)

Copied to clipboard

Challenge: Abstractive dialogue summarization aims to convert long dialogue content into its short form where the salient information is preserved while the redundant pieces are ignored.
Approach: They propose to have the model perceive the redundant parts of an input dialogue history during the training phase.
Outcome: The proposed method significantly outperforms baselines on the semantic matching and factual consistent based metrics.
Cross-modal Coherence Modeling for Caption Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for image captioning do not guarantee consistent image-text relations . current models do not provide enough data for training robust captioning models .
Approach: They use an annotation protocol specifically devised for capturing image–caption coherence relations to study image captioning.
Outcome: The proposed protocol improves image captioning models with coherence relations . the dataset is large enough to alleviate content hallucinations, the authors show .
ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in automated essay scoring have limited the generalizability of models trained on ASAP.
Approach: They propose to annotate persuasive student essays with holistic and trait-specific scores in a corpus of persuasive student essay annotated with ICLE++.
Outcome: The proposed model can be used to evaluate models for newer AES problems such as multi-trait scoring and cross-prompt scoring.
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Current evaluation methods for large language models rely on static benchmarks . limited knowledge coverage and fixed difficulties hinder the targeted optimizations resulting in superficial evaluations of LLMs - a problem that has been addressed by JudgeAgent .
Approach: They propose a knowledge-driven and dynamic evaluation framework for large language models . judgeAgent leverages LLM agents equipped with context graphs to traverse knowledge structures .
Outcome: The proposed framework can achieve comprehensive evaluations and facilitate effective model iterations.
Cross-Prompt Automated Essay Scoring of Multiple Traits: Making Sense of the State of the Art (2026.acl-long)

Copied to clipboard

Challenge: despite recent progress in cross-prompt essay scoring, there is little analysis of what makes a state-of-the-art cross-propert scorer work well.
Approach: They propose to apply transductive learning to cross-prompt scoring for the first time . they propose to train a model that can offer good performance when applied to unseen prompts .
Outcome: The proposed model could be used in the rarely-studied classroom setting without additional training data.
BayesKD: Bayesian Knowledge Distillation for Compact LLMs in Constrained Fine-tuning Scenarios (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized various domains with their remarkable capabilities, but their massive parameter sizes pose significant challenges for fine-tuning and inference.
Approach: They propose a Bayesian Knowledge Distillation framework for compact Large Language Models in resource-constrained fine-tuning scenarios that employs Logits Dual-Scaling, Knowledge Alignment Module, and Bayes Distillations Optimization.
Outcome: The proposed framework outperforms baseline methods on various state-of-the-art LLMs, including LLaMA, Qwen2, Bloom, and Vicuna.
Graph-Based Multi-Trait Essay Scoring (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work on Automated Essay Scoring (AES) models essay as word sequence, but new approach uses graph-attention network approach to model essay traits.
Approach: They propose a graph-attention network approach to automate essay scoring that models interactions among essay traits as a graphical graph.
Outcome: The proposed approach outperforms competing approaches on the ASAP++ dataset . it allows for multiple-task scoring, allowing for more detailed feedback on essays .
VLP: Vision-Language Preference Learning for Embodied Manipulation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to reward engineering are time-consuming and expensive to collect human preference labels.
Approach: They propose a vision-language preference learning framework which learns from human feedback . they define three types of language-conditioned preferences and construct a visual preference dataset .
Outcome: The proposed framework outperforms baselines on embodied manipulation tasks and can be applied to other tasks.
AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown impressive performance on a range of tasks, yet advanced instruction following (IF) remains a significant challenge.
Approach: They propose a benchmark that features over 1,600 prompts and expert-curated rubrics that assess LLMs’ ability to follow complex, multi-turn, and system-level instructions.
Outcome: The proposed framework improves instruction-following abilities of large language models, achieving a 6.7% gain on AdvancedIF and strong results on public benchmarks.
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Despite the success of transformer-based large language models, understanding and enhancing their mathematical capabilities remains a significant challenge.
Approach: They propose to use numerical precision as a key factor that influences LLMs' effectiveness in arithmetical tasks to determine their effectiveness.
Outcome: The proposed models perform better in arithmetic tasks than transformer-based models with standard numerical precision.
End-to-End Neural Discourse Deixis Resolution in Dialogue (2022.emnlp-main)

Copied to clipboard

Challenge: Lexical overlap is a strong indicator of entity coreference, both among names and in the resolution of nominals.
Approach: They propose to extend their span-based entity coreference model to exploit task-specific characteristics of discourse deixis resolution in dialogue.
Outcome: The proposed model achieves state-of-the-art results on the four datasets in the CODI-CRAC 2021 shared task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations