Papers by Jiaming Yuan

6 papers
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks.
Approach: They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement.
Outcome: The findings highlight the future directions in medical reasoning, physical system integration, and training simulations.
Towards IP Intelligence: Benchmarking Large Language Models on Intellectual Property Knowledge and Practice (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets and benchmarks focus only on patents or cover limited aspects of the IP field, lacking alignment with real-world scenarios.
Approach: They propose a bilingual IP task taxonomy and a large-scale bilingual benchmark to evaluate LLMs in real-world IP practice.
Outcome: The proposed model achieves only 75.8% accuracy, indicating room for improvement . open-source IP and law-oriented models lag behind closed-source general-purpose models .
Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for dataset poisoning require full-dataset poison, which breaks code compilability.
Approach: They propose a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths.
Outcome: The proposed method contaminates 10% of the dataset while maintaining 100% compilability and functional correctness.
Deciphering Undersegmented Ancient Scripts Using Phonetic Prior (2021.tacl-1)

Copied to clipboard

Challenge: a number of undeciphered languages are still undecipherated, igniting fierce scientific debate . a recent study shows that NLP methods can successfully decipher lost languages .
Approach: They propose a decipherment model that incorporates phonetic geometry into word segmentation and cognate alignment . they use the International Phonetic Alphabet to learn character embeddings based on historical sound change .
Outcome: The proposed model shows that it can decipher both deciphered and undeciphered languages.
Changing the Mind of Transformers for Topically-Controllable Language Generation (2021.eacl-main)

Copied to clipboard

Challenge: Existing interactive writing assistants do not allow authors to guide text generation in desired topical directions.
Approach: They propose a framework that displays multiple candidate upcoming topics and generates a text generation model that adheres to the chosen topics.
Outcome: The proposed model generates fluent sentences related to the selected topics, as judged by automated metrics and crowdsourced workers.
Neural Decipherment via Minimum-Cost Flow: From Ugaritic to Linear B (P19-1)

Copied to clipboard

Challenge: Existing methods for decipherment of lost languages are limited by limited data and scarce quantities of ancient text.
Approach: They propose a neural approach for automatic decipherment of lost languages . they use an expressive sequence-to-sequence model to capture character-level correspondences between cognates .
Outcome: The proposed approach improves on the decipherment of Ugaritic and Linear B in ancient Greek . the proposed approach is highly customized for a given language pair and does not generalize to other lost languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations