Papers by Chen Lyu

55 papers
String Editing Based Chinese Grammatical Error Diagnosis (2022.coling-1)

Copied to clipboard

Challenge: Chinese Grammatical Error Diagnosis (CGED) suffers from numerous types of grammatical errors and insufficiency of training data.
Approach: They propose a string editing based CGED model that uses a unified workflow to handle various types of grammatical errors.
Outcome: The proposed model outperforms existing models on Chinese datasets in many aspects.
ANAH: Analytical Annotation of Hallucinations in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: a comprehensive and fine-grained measurement of the hallucination is crucial for LLMs' wide applications.
Approach: They propose a dataset that offers ANalytical Annotation of Hallucinations in Large Language Models.
Outcome: The proposed dataset can be used to train and evaluate hallucination annotators.
Collaborative Performance Prediction for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are one of the most important AI research powered by largescale parameters, high computational resources, and massive training data.
Approach: They propose a framework that leverages historical performance of large language models and other design factors to improve prediction accuracy.
Outcome: The proposed framework surpasses scaling laws in predicting performance of large language models . it also facilitates a detailed analysis of factor importance, an area previously overlooked .
LLM-Rec: Personalized Recommendation via Prompting Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have showcased their remarkable ability to harness commonsense knowledge and reasoning.
Approach: They propose a novel approach which incorporates four distinct prompting strategies of text enrichment for improving personalized text-based recommendations.
Outcome: The proposed approach improves recommendation quality and even basic MLP models achieve comparable or even better results than complex content-based methods.
Generation-Augmented and Embedding Fusion in Document-Level Event Argument Extraction (2025.coling-main)

Copied to clipboard

Challenge: Document-level event argument extraction is a crucial task that aims to extract arguments from the entire document, beyond sentence-level analysis.
Approach: They propose a novel approach to document-level event argument extraction that integrates predefined templates and generative language models into a foundational embedding derived from a classification model.
Outcome: The proposed approach is more effective than baseline models and data-efficient in low-resource scenarios.
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference (2026.acl-long)

Copied to clipboard

Challenge: Existing pruning methods ignore prefill-decode (PD) disaggregation in practice.
Approach: They propose a pruning method that is highly integrated with prefill-decode (PD) disaggregation, enabling more precise pruning of blocks.
Outcome: The proposed method achieves strong performance in both PD disaggregation and PD unified settings, and can be extended to other non-block pruning methods.
Think Wider, Detect Sharper: Reinforced Reference Coverage for Document-Level Self-Contradiction Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to document-level contradiction detection (DSCD) only gain marginal improvement and often introduce inconsistencies across repeated responses.
Approach: They propose a method that combines supervised fine-tuning and reinforcement learning to enhance document-level contradiction detection (DSCD) they propose to use a task-specific reward function to expand the model’s reasoning scope, boosting both accuracy and consistency.
Outcome: The proposed method significantly boosts Llama 3.1-8B-Instruct’s accuracy from 38.5% to 51.1%, and consistency from 59.6% to76.2%.
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation.
Approach: They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation.
Outcome: The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics.
CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on functional relevance while neglecting code quality.
Approach: They propose a multilingual benchmark to evaluate quality-aware code retrieval . they include fine-grained quality annotations over 42,725 queries and 134,907 code snippets .
Outcome: The proposed benchmarks show that state-of-the-art models fail to separate buggy or insecure code from robust counterparts.
Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches for personalizing large language models require modifying parameters.
Approach: They propose a lightweight approach to personalizing large language models via retrieval augmentation . relevance serves as an unreliable proxy for utility, they argue .
Outcome: The proposed framework outperforms strong heuristic and retrieval-augmented baselines on nine personalization tasks.
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective (2026.findings-acl)

Copied to clipboard

Challenge: Code Large Language Models (CLLMs) are reshaping how software is built, maintained, and evolved.
Approach: They propose to use BPE tokenization to inadvertently leak code secrets . they propose to mitigate the gibberish bias by using a newer tokenizer .
Outcome: The proposed model is based on a novel method that can be used to detect and mitigate gibberish bias in CLLMs.
Attention-Enhancing Backdoor Attacks Against BERT-based Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing textual backdoor attacks focus on generating stealthy triggers or modifying model weights.
Approach: They propose a Trojan Attention Loss (TAL) which enhances the Trojan behavior by directly manipulating attention patterns.
Outcome: The proposed method improves the effectiveness of the backdoor attacks on different backbone models and tasks.
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records (2026.acl-long)

Copied to clipboard

Challenge: GUI agents have shown strong performance under explicit and completion instructions, but real-world deployment requires aligning with users’ more complex implicit intents.
Approach: They propose a task that requires agents to leverage long-term user records as persistent context to resolve omitted preferences in vague instructions and anticipate latent routines by user state for proactive assistance.
Outcome: The proposed task improves execution and proactive performance by 15.7% and 7.3% under explicit and completion instructions.
Extracted BERT Model Leaks More Information than You Think! (2022.emnlp-main)

Copied to clipboard

Challenge: Existing pre-trained language models are vulnerable to model extraction attacks . model extraction can cause severe privacy leakage even when victim models are facilitated with state-of-the-art defensive strategies.
Approach: They propose to launch an attribute-inference attack against an extracted BERT model to prevent privacy leakage.
Outcome: The proposed attack can cause severe privacy leakage even when victim models are facilitated with state-of-the-art defensive strategies.
From Selection to Generation: A Survey of LLM-based Active Learning (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used for selection and training of data for active learning.
Approach: They propose an intuitive taxonomy that categorizes LLM-based active learning techniques and discuss the transformative roles they can play in the active learning loop.
Outcome: The proposed model can generate entirely new data instances and provide more cost-effective annotations with fewer labeled data instances.
Streamlining the Collaborative Chain of Models into A Single Forward Pass in Generation-Based Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in "Chain of Models" approach increase resource demands as each model must be deployed separately.
Approach: They propose a prompt-tuning method that enables models to share hidden states . they modify input and attention masks during training to eliminate redundant forward passes .
Outcome: Empirical results show that FTHSS matches the performance of traditional model chains while improving inference efficiency.
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches lack robustness to handle complex edge cases and generalizability across different domains.
Approach: They develop an accurate and lightweight verifier model for evaluation and outcome reward that matches unstructured outputs against standard answers.
Outcome: The proposed model can process multiple answer types including multi-subproblems, formulas, and sequence answers while identifying abnormal/invalid responses.
Event Extraction as Multi-turn Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to event extraction fail to model rich interactions among event types and arguments of different roles.
Approach: They propose a new paradigm that formulates event extraction as multi-turn question answering . they propose to use reading comprehension problems to extract triggers and arguments .
Outcome: The proposed approach outperforms current state-of-the-art on argument extraction tasks . it makes full use of dependency among arguments and event types, and generalizes well .
Repulsive Attention: Rethinking Multi-head Attention as Bayesian Inference (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that multi-head attention is an effective module in deep neural networks, but there are no explicit mechanisms guaranteeing this property.
Approach: They propose a non-parametric approach that explicitly improves the repulsiveness in multi-head attention and consequently strengthens model’s expressiveness.
Outcome: The proposed approach improves the repulsiveness in multi-head attention and strengthens model’s expressiveness.
RESIN: A Dockerized Schema-Guided Cross-document Cross-lingual Cross-media Information Extraction and Event Tracking System (2021.naacl-demos)

Copied to clipboard

Challenge: We present a new information extraction system that can construct temporal event graphs from news documents.
Approach: They propose a temporal event graph extraction system that can extract news documents . they extend the system from sentence-level event extraction to cross-document cross-media event extraction .
Outcome: The proposed system can extract temporal event graphs from news documents in multiple languages and multiple data modalities.
Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided Decoding (2025.naacl-long)

Copied to clipboard

Challenge: Multi-modal large language models (MLLMs) generate plausible but incorrect content, resulting in hallucinations . recent advances in MLLM technology have demonstrated their outstanding performance in a variety of visual tasks, such as object detection.
Approach: They propose a plug-and-play method which leverages MLLMs’ internal representations to mitigate hallucinations by analyzing input and output tokens.
Outcome: The proposed method exploits MLLMs’ internal representations to mitigate hallucinations.
Relation Classification with Entity Type Restriction (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods regard all relations as candidate relations for the two entities, which leads to inappropriate relations being candidate relations.
Approach: They propose a paradigm which exploits entity types to restrict candidate relations by mutual restrictions.
Outcome: The proposed paradigm improves GCN and SpanBERT on a standard dataset by 6.9 and 4.4 F1 points.
VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Existing video metrics are lagging behind in providing reliable scores over generated videos due to lack of large-scale human-annotated dataset.
Approach: They propose to use VideoFeedback to train a human-annotated multi-aspect score over 37.6K synthesized videos from 11 existing video generative models.
Outcome: The proposed model outperforms the prior best metrics by 50 points in the test.
Explanation in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events.
Approach: They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations.
Outcome: The proposed methods are based on the models of large language models (LLMs) and their opaque nature.
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for agentic repository-level code understanding overlook long tail topics and rely on memorized knowledge.
Approach: They propose a repository-level agentic code understanding benchmark that uses long-tail repositories with executable environments to enforce topical balance.
Outcome: Empirically, a Qwen3-8B model trained with the proposed benchmark outperforms GPT-4o by 2.3 points.
Capturing Argument Interaction in Semantic Role Labeling with Capsule Networks (D19-1)

Copied to clipboard

Challenge: State-of-the-art SRL models do not model non-local interaction between arguments . e.g., LSTMs do not allow for efficient inference .
Approach: They propose a new approach to model interactions between arguments using capsule networks . they analyze errors in the refinement procedure by capturing intuition in a flexible way .
Outcome: The proposed model outperforms the baseline model on all 7 languages and achieves state-of-the-art results on 5 languages including English.
Neural Graph Matching Networks for Chinese Short Text Matching (2020.acl-main)

Copied to clipboard

Challenge: Chinese word segmentation can be erroneous, ambiguous or inconsistent, causing performance problems.
Approach: They propose a sentence matching framework that uses paired word lattices as input instead of a character sequence.
Outcome: The proposed framework outperforms the state-of-the-art short text matching models on two Chinese datasets.
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on majority voting or criteria expansion to capture detailed and detailed details, often leading to incomplete outcomes.
Approach: They propose a method which introduces additional crowd responses to compare with the candidate responses, thereby exposing deeper and more comprehensive details within the candidate answers.
Outcome: Experiments show that the proposed method improves evaluation reliability and achieves an average gain of 6.7% across five benchmarks.
Training Language Models to Critique With Multi-agent Feedback (2025.findings-emnlp)

Copied to clipboard

Challenge: utilizing human annotations can enhance critique ability, but model-generated critiques suffer from inherent flaws due to complexity of critique . a new framework that leverages multi-agent feedback improves critique ability .
Approach: They propose a framework that leverages multi-agent feedback to improve critique ability . they propose to use supervised fine-tuning and reinforcement learning to improve this capability .
Outcome: The proposed framework improves critique ability in both supervised fine-tuning and reinforcement learning stages.
Stable Language Guidance for Vision–Language–Action Models (2026.acl-long)

Copied to clipboard

Challenge: Existing vision-Language-Action models are notoriously brittle to linguistic perturbations.
Approach: They propose a probabilistic framework that disentangles physical affordance from semantic execution.
Outcome: The proposed framework disentangles physical affordance from semantic execution.
SoMeLVLM: A Large Vision Language Model for Social Media Processing (2024.findings-acl)

Copied to clipboard

Challenge: Genereal domain large models lack nuanced multimodal understanding of social media . general domain models focus more on text than other modalities, which is not consistent with real-world user habits.
Approach: They propose a Large Vision Language Model for Social Media Processing that combines five key capabilities to understand and generate real social media behavior.
Outcome: The proposed model achieves state-of-the-art performance in multiple social media tasks.
Task-Agnostic Detector for Insertion-Based Backdoor Attacks (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for textual backdoor detection are task-specific and less effective beyond sentence classification.
Approach: They propose a task-agnostic method for backdoor detection that leverages final layer logits and an efficient pooling technique.
Outcome: TABDet can jointly learn from diverse task-specific models, demonstrating superior detection efficacy over traditional methods.
Retrieve-Plan-Generation: An Iterative Planning and Answering Framework for Knowledge-Intensive LLM Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) often produce factual errors due to limited internal knowledge.
Approach: They propose a retrieval-augmented generation framework that generates plan tokens to guide subsequent generation.
Outcome: The proposed framework improves the accuracy of large language models with external knowledge sources.
AwarenessBench: Assessing Cognitive Capabilities of Language Models (2026.acl-long)

Copied to clipboard

Challenge: Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities.
Approach: They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness .
Outcome: Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds .
Multi-Defendant Legal Judgment Prediction via Hierarchical Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for predicting judgment results for multiple defendants are ineffective.
Approach: They propose a method to predict the judgment results for each defendant in multi-defendant cases . they formalize the multi-diffendant judgment process as hierarchical reasoning chains .
Outcome: The proposed method can predict the judgment results for multiple defendants in multi-defendant cases.
Enhancing Code Generation Performance of Smaller Models by Distilling the Reasoning Ability of LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have made significant advances in code generation through the ‘Chain-of-Thought’ prompting technique.
Approach: They propose a framework which aims to transfer LLMs’ reasoning capabilities to smaller models through distillation.
Outcome: The proposed framework improves the smaller model's code generation performance by over 130% on the APPS benchmark.
Few-Shot Text Classification with Edge-Labeling Graph Neural Network-Based Prototypical Network (2020.coling-main)

Copied to clipboard

Challenge: a few-shot text classification method is proposed to solve the few-sshot text problem . supervised learning methods require large corpus of labeled data, making them hindered in practical application.
Approach: They propose a few-shot text classification method that takes advantage of advanced pre-trained language models to extract the semantic features of each document.
Outcome: The proposed method achieves state-of-the-art on sentiment analysis and relation datasets.
VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration (2025.findings-acl)

Copied to clipboard

Challenge: Existing safety calibration methods focus on model undersafety, where the model responds to hazardous queries, while neglecting oversafetiness, where models refuse to answer safe queries.
Approach: They propose safety calibration which addresses both undersafety and oversafetiness by comparing model responses to a novel dataset of 3,600 image-text pairs.
Outcome: The proposed methods have been used to evaluate safety calibration across image-centric and text-centric scenarios.
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Hallucination is a significant challenge for large language models, but current methods struggle when non-factual information arises in the early or mid-sequence of outputs, reducing their reliability.
Approach: They propose a method that captures the full dynamics of large language models by using neural differential equations to assess the truthfulness of statements.
Outcome: The proposed method achieves 14% improvement in AUC-ROC on the True-False dataset compared to state-of-the-art methods.
A Study of the Attention Abnormality in Trojaned BERTs (2022.naacl-main)

Copied to clipboard

Challenge: In computer vision, the trigger can be a fixed pattern overlaid on the images or videos.
Approach: They propose an attention-based Trojan detector to distinguish Trojaned models from clean ones by observing the attention focus drifting behavior of Trojanes.
Outcome: The proposed detector is based on transformer’s attention and can distinguish Trojan models from clean ones.
GUI Agents: A Survey (2025.findings-acl)

Copied to clipboard

Challenge: Large Foundation Models (LFMs) have transformed the landscape of AI research and day-to-day life.
Approach: They propose a framework that delineates GUI agents' perception, reasoning, planning, and acting capabilities.
Outcome: The proposed framework delineates their perception, reasoning, planning, and acting capabilities.
Glyph Enhanced Chinese Character Pre-Training for Lexical Sememe Prediction (2021.findings-emnlp)

Copied to clipboard

Challenge: Sememes are defined as the atomic units to describe the semantic meaning of concepts.
Approach: They propose a method which incorporates internal Chinese character information to help sememe prediction.
Outcome: The proposed method outperforms existing non-external information models on howNet, a famous sememe knowledge base.
Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics (2025.acl-long)

Copied to clipboard

Challenge: Existing studies assessing the spatial abilities of VLMs lack a solid theoretical foundation and lack measurable data.
Approach: They propose a psychometric framework defining five basic spatial abilities in Visual Language Models.
Outcome: The proposed framework defines five basic spatial abilities in Visual Language Models (VLMs) it provides a comprehensive evaluation benchmark and methodological perspective for embodied AI development .
KnowTuning: Knowledge-aware Fine-tuning for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are a default solution for many natural language processing tasks.
Approach: They propose a knowledge-aware fine-tuning method to improve LLMs' knowledge awareness . they propose augmentation and comparison stages to improve accuracy and reliability .
Outcome: The proposed method generates more facts with less factual error rate under fine-grained facts evaluation.
Multi-view Classification Model for Knowledge Graph Completion (2020.aacl-main)

Copied to clipboard

Challenge: Existing knowledge graph completion models only evaluated candidate triples from content information.
Approach: They propose a multi-view classification model where multiple views are performed based on both content and context information for candidate triple evaluation.
Outcome: The proposed model improves on two representative datasets and improves performance.
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: evolving generic Large Language Models into specialized Large Reasoning Models requires effective post-training.
Approach: They propose a plasticity-ceiling framework to harness expert trajectories . they establish the Sequential SFT-then-RL pipeline as the superior standard .
Outcome: The proposed framework overcomes stability and premature convergence deficits in synchronized approaches.
MedCoach: Enhancing Medical Reasoning in LLMs via Knowledge Graph-Augmented Chain-of-Thought Distillation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for training specialized reasoning models for the medical domain are limited due to the scarcity of high-quality, large-scale Chain-of-Thought (CoT) data.
Approach: They propose a framework that introduces a dedicated coach role to guide the student model through question decomposition.
Outcome: The proposed framework smooths the learning curve in medical reasoning by facilitating domain adaptation before advancing to complex long-chain reasoning.
Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Medical Multi-Modal Large Language Models (Med-MLLMs) are a promising new form of artificial general intelligence due to their ability to tackle complex tasks.
Approach: They propose a new benchmark that comprehensively assesses medical multi-modal large language models in terms of distinct medical specialties and different diagnostic capacities.
Outcome: The proposed model covers 15 medical specialties and different diagnostic capacities, and excludes overlap with existing VQA dataset.
EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Vision Language Models (LVLMs) show promise for pathological diagnosis, yet their application in clinical settings faces critical challenges of multimodal hallucination and biased responses.
Approach: They propose a framework that integrates medical expertise into preference alignment.
Outcome: The proposed framework outperforms existing pathological LVLMs while maintaining pathological accuracy.
All Languages Matter: On the Multilingual Safety of LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Existing safety benchmarks only concern the safety in one language, e.g. the majority language in the pretraining data such as English.
Approach: They propose a prompting method to improve multilingual safety of ChatGPT by enhancing cross-lingual generalization of safety alignment.
Outcome: The proposed method can significantly reduce the ratio of unsafe responses by 42% for non-English queries.
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) fail to effectively guide the planning trajectories during task solving and result in planning hallucinations.
Approach: They propose a novel approach to enhance the planning capabilities of large language models by incorporating explicit action knowledge.
Outcome: The proposed approach can achieve comparable or superior performance to existing baselines on HotpotQA and ALFWorld.
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents (2025.findings-acl)

Copied to clipboard

Challenge: Role-Playing Agents (RPAs) are increasingly popular due to diverse task requirements and agent designs.
Approach: They propose an evidence-based evaluation design guideline for LLM-based RPAs based on agent attributes, task attributes, and evaluation metrics.
Outcome: The proposed evaluation design guideline is based on a systematic review of 1,676 papers published between Jan. 2021 and Dec. 2024.
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches lack flexibility to address diverse and ever-evolving user queries in open domains.
Approach: They propose to evaluate LLMs on open-domain knowledge that requires tools to solve diverse and ever-evolving user queries.
Outcome: The proposed system outperforms baselines in the open domain task-solving benchmark.
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks fail to capture scenarios in which vulnerabilities are introduced by humans . we evaluate 5 popular code agents supported by 5 LLMs on SecureVibeBench .
Approach: They propose a benchmarking tool that compares 105 C/C++ secure coding tasks . they use real-world open-source vulnerabilities and a comprehensive evaluation tool .
Outcome: The proposed benchmarks show that code agents struggle to produce correct and secure code . the best performing agent produces merely 23.8% correct and secured solutions .
CLEVA: Chinese Language Models EVAluation Platform (2023.emnlp-demo)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized natural language processing.
Approach: They propose a Chinese-based platform that assesses Chinese LLMs using a standardized workflow and a unique sampling strategy.
Outcome: CLEVA evaluates Chinese LLMs on a standardized workflow and a competitive leaderboard with minimal coding.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations