Papers by James Ren

4 papers
SLENDER: Structured Outputs for SLM-based NER in Low-Resource Englishes (2025.acl-industry)

Copied to clipboard

Challenge: Named Entity Recognition (NER) for low-resource variants of English remains challenging, as most models are trained on datasets predominantly focused on American or British English.
Approach: They propose a new output format for Named Entity Recognition (NER) that achieves a three-fold reduction in inference time compared to JSON format.
Outcome: The proposed output format achieves a three-fold reduction in inference time on average compared to JSON format, which is widely used for structured outputs.
Bi-Phone: Modeling Inter Language Phonetic Influences in Text (2023.acl-long)

Copied to clipboard

Challenge: Increasingly, people are forced to use the Web in languages they have low literacy in due to technology asymmetries.
Approach: They propose a method to mine phoneme confusions for pairs of L1 and L2 and plug them into a generative model for synthetically producing corrupted L2 text.
Outcome: The proposed method corrupts the popular language understanding benchmark SuperGLUE and improves performance.
AutoTriggER: Label-Efficient and Robust Named Entity Recognition with Auxiliary Trigger Extraction (2023.eacl-main)

Copied to clipboard

Challenge: Named entity recognition models have shown impressive results in overcoming label scarcity and generalizing to unseen entities by leveraging distant supervision and auxiliary information such as explanations.
Approach: They propose a framework that automatically generates and leverages “entity triggers” which are human-readable cues in the text that help guide the model to make better decisions.
Outcome: The proposed framework outperforms the RoBERTa-CRF baseline by nearly 0.5 F1 points on three well-studied datasets.
Protein Large Language Models: A Comprehensive Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on specific aspects or applications, but this study provides a comprehensive overview of Protein-specific large language models.
Approach: This paper proposes a structured taxonomy of state-of-the-art ProteinLLMs . they analyze how they leverage large-scale protein sequence data for improved accuracy .
Outcome: The proposed model covers their architectures, training datasets, evaluation metrics, and diverse applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations