Papers by Peter Chen

8 papers
A Cross-lingual Messenger with Keyword Searchable Phrases for the Travel Domain (C18-2)

Copied to clipboard

Challenge: Query Translator is a cross-lingual messaging app for the travel domain that automatically translates conversations . the application addresses common cross-linguistic communication issues such as translation accuracy, speed, privacy and personalization.
Approach: They present a cross-lingual messaging app that automatically translates conversations while supporting keyword-to-sentence matching.
Outcome: The proposed app translates conversations while supporting keyword-to-sentence matching.
BIG-Bench Extra Hard (2025.acl-long)

Copied to clipboard

Challenge: Current benchmarks for large language model reasoning focus on math and coding abilities, leaving a gap in evaluating broader reasoning proficiencies.
Approach: They propose a benchmark to evaluate general reasoning in large language models . they use BIG-Bench and its harder version BIG-Benefit Hard to assess general reasoning .
Outcome: The new benchmark pushes the boundaries of LLM reasoning evaluation.
Spoken Document Retrieval for an Unwritten Language: A Case Study on Gormati (2025.findings-emnlp)

Copied to clipboard

Challenge: Speakers of unwritten languages have the potential to benefit from speech-based automatic information retrieval systems.
Approach: They propose a speech embedding technique that facilitates a zero-shot speech-based automatic information retrieval system for unwritten languages.
Outcome: The proposed method achieves a Top 5 retrieval rate of 87.9% on a corpus of Gormati, an unwritten language, that was collected in partnership with an agrarian Banjara community in Maharashtra State, India.
A Fully Generative Motivational Interviewing Counsellor Chatbot for Moving Smokers Towards the Decision to Quit (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are being used to provide automated talk therapy . however, it is crucial to know if they would be effective and adhere to known standards.
Approach: They propose to use large language models to automate talk therapy with a focus on tobacco addiction.
Outcome: The proposed chatbot showed adherence to MI standards in 98% of utterances, higher than human counsellors.
Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings (2023.acl-short)

Copied to clipboard

Challenge: Recent research has called into question the necessity of positional embeddings in transformer language models.
Approach: They propose to discard positional embeddings in transformer language models to facilitate more efficient pretraining.
Outcome: The proposed model encodes strong positional information through shrinkage of self-attention variance.
LLMs cannot find reasoning errors, but can correct them given the error location (2024.findings-acl)

Copied to clipboard

Challenge: Recent attempts to self-correct logical or reasoning errors often cause correct answers to become incorrect, resulting in poor performance overall.
Approach: They propose to use a backtracking setup to test the correction abilities of LLMs on their mistake-finding ability to find logical mistakes.
Outcome: The proposed model improves on 5 reasoning tasks, showing that it can correct logical mistakes without ground truth labels or training data.
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)

Copied to clipboard

Challenge: With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages.
Approach: They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA.
Outcome: The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices.
Generating Logical Forms from Graph Representations of Text and Entities (P19-1)

Copied to clipboard

Challenge: Recent approaches to semantic parsing have cast it as a sequence-to-sequence task, with strong results.
Approach: They propose a Graph Neural Network architecture to incorporate information about relevant entities and their relations during parsing.
Outcome: The proposed approach outperforms the state-of-the-art in several tasks without pre-training and outperformed existing approaches when combined with BERT pre-trainment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations