Papers by Raymond Ng

7 papers
Stigma Annotation Scheme and Stigmatized Language Detection in Health-Care Discussions on Social Media (2020.lrec-1)

Copied to clipboard

Challenge: a large amount of research has been done on the interpretation and influence of stigma on human behaviour and health.
Approach: They develop an annotation scheme and improve the annotation process for stigma identification . they aim to distinguish stigmatised language from non-stigmatised using machine learning and NLP .
Outcome: The proposed method improves the annotation process for stigma identification . the results show that the method performs better than other models .
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages? (2026.acl-long)

Copied to clipboard

Challenge: SEA-BED examines how multilingual text embeddings perform across tasks and languages . performance gaps arise from data coverage, training objectives, and architectural design, authors say .
Approach: They propose a large-scale benchmark covering 10 SEA languages and diverse embedding tasks.
Outcome: The proposed model performs poorly across languages and tasks, but language-task analyses reveal inconsistencies . the results suggest that performance gaps arise from limitations in data coverage, training objectives, and architectural design.
Discourse Analysis and Its Applications (P19-4)

Copied to clipboard

Challenge: Discourse processing is a suite of NLP tasks to uncover linguistic structures from texts at several levels, which can support many downstream applications.
Approach: They present a set of tasks to uncover linguistic structures from texts at several levels, which can support many downstream applications.
Outcome: The tutorial covers the basic concepts of discourse analysis and linguistic structures in monologue vs. conversation, synchronous v. asynchronous conversation, and key linguistic structure in discourse analysis.
COMET-M: Reasoning about Multiple Events in Complex Sentences (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing commonsense models that generate event-centric inferences for simple sentences struggle with the complexity of multi-event sentences prevalent in natural text.
Approach: They propose a commonsense model that generates inferences for a target event within a complex sentence using a multi-event inference dataset.
Outcome: The proposed model produces inferences for a target event within a complex sentence taking the complete context into account.
What happens before and after: Multi-Event Commonsense in Event Coreference Resolution (2023.eacl-main)

Copied to clipboard

Challenge: Existing event coreference models cluster event mentions pertaining to the same event, but they fail to leverage commonsense inferences for lexically-divergent mentions.
Approach: They propose a model that extends event mentions with temporal commonsense inferences to generate plausible events that happen before and after the target events.
Outcome: The proposed model generates plausible events that happen before and after the target event, and then after it, such as "he was sentenced".
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Reliable multilingual evaluation is difficult and culturally appropriate evaluation is even harder to achieve.
Approach: They propose a multilingual evaluation framework that aims to mitigate these biases by improving translations and annotation practices.
Outcome: The proposed framework improves translation quality and cultural coverage and is culturally sensitive and culturally agnostic.
A High Precision Pipeline for Financial Knowledge Graph Construction (2020.coling-main)

Copied to clipboard

Challenge: Knowledge graphs are a standard for structured knowledge representation in the Semantic Web.
Approach: They propose to extract financial news articles into a knowledge graph by using a financial dictionary.
Outcome: The proposed pipeline extracts 342,000 financial news articles with a precision of 78% at the top-100 extractions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations