Papers by Che Lin

5 papers
RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation (2024.emnlp-demo)

Copied to clipboard

Challenge: Xia et al., 2018) demonstrate that a large language model can generate and maintain high-quality code documentation.
Approach: They propose a large language model powered open-source framework for generating, maintaining, and updating code documentation.
Outcome: The proposed framework generates high-quality documentation for the entire project.
A Compare-and-contrast Multistage Pipeline for Uncovering Financial Signals in Financial Reports (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have included attempts to efficiently and effectively comprehend lengthy financial documents.
Approach: They propose a signal-highlighting task that analyzes relationships between financial reports . they also create and publicly release a human-annotated dataset for the task .
Outcome: The proposed pipeline is based on a human-annotated dataset and validates its effectiveness.
Editing the Mind of Giants: An In-Depth Exploration of Pitfalls of Knowledge Editing in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Knowledge editing is a promising technique for updating factual knowledge in large language models (LLMs) but studies have identified side effects such as knowledge distortion and the deterioration of general abilities that have emerged after editing.
Approach: They propose to evaluate the side effects of knowledge editing in large language models using metrics and benchmarks.
Outcome: The results of the study highlight the limitations of current knowledge editing methods and outline potential research directions.
NASH: Numerically Aware Scoring Heuristic for Robust Semantic Similarity (2026.findings-acl)

Copied to clipboard

Challenge: Numerical precision is critical in financial NLP, yet embedding-based semantic similarity metrics exhibit numerical blindness.
Approach: They propose a model-agnostic metric that decouples numerical verification from textual semantic evaluation.
Outcome: The proposed metric improves numerical sensitivity while maintaining general semantic performance.
Breaking Boundaries in Retrieval Systems: Unsupervised Domain Adaptation with Denoise-Finetuning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing domain adaptation methods for dense retrieval models use unadapted rerank models, leading to imprecise labels.
Approach: They propose to adapt a rerank model to the target domain before using it for label generation.
Outcome: The proposed model achieves better results across three retrieval datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations