Papers by Skatje Myers

6 papers
Building a Broad Infrastructure for Uniform Meaning Representations (2024.lrec-main)

Copied to clipboard

Challenge: This paper reports the first release of the UMR data set for six languages . it includes annotations for six different languages that vary greatly in terms of their linguistic properties and resource availability.
Approach: They report the first release of the UMR data set for six languages . they describe on-going efforts to enlarge the data set and extend it to other languages - including Navajo, Navájo, and Sanapaná .
Outcome: The first release of the UMR data set includes annotations for six languages . the language dataset is available for free and can be extended to other languages if needed .
Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning Representation (2021.acl-long)

Copied to clipboard

Challenge: Compared with general natural language texts, sentences from scientific papers usually possess wider contexts between knowledge elements.
Approach: They propose a novel biomedical Information Extraction model to extract scientific entities and events from English research papers using Abstract Meaning Representation (AMR) they construct a sentence-level knowledge graph from an external knowledge base and encode it to improve the model's understanding of complex scientific concepts.
Outcome: The proposed model can extract scientific entities and events from scientific literature and improve its understanding of complex scientific concepts.
Leveraging Active Learning to Minimise SRL Annotation Across Corpora (2023.starsem-1)

Copied to clipboard

Challenge: In this paper, we investigate the application of active learning to semantic role labeling (SRL) using Bayesian Active Learning by Disagreement (BALD).
Approach: They propose a sentence-focused selection method that is based off of previous methods of using model dropout to approximate a Gaussian process for SRL.
Outcome: The proposed selection method improves on three different domain corpora on three domains with a large and diverse corpus.
PropBank Comes of Age—Larger, Smarter, and more Diverse (2022.starsem-1)

Copied to clipboard

Challenge: The PropBank has been used for semantic role labeling for over 20 years . it includes non-verbal predicates, adjectives, prepositions and multi-word expressions .
Approach: They describe the evolution of the PropBank approach to semantic role labeling over the last 20 years . they describe the substantial effort that has gone into ensuring consistency and reliability of the various annotated datasets and resources .
Outcome: The PropBank has been used for more than 20 years to test semantic role labeling systems.
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications? (2024.findings-emnlp)

Copied to clipboard

Challenge: Numerical data is pivotal for medical questions and answers, but tabular data is not fully integrated into LLMs.
Approach: They examine the effectiveness of vector representations from last hidden states of LLMs for medical diagnostics and prognostics using electronic health record data.
Outcome: The proposed representations outperform those using raw numerical EHR data in medical diagnostics and prognostics.
The Russian PropBank (2020.lrec-1)

Copied to clipboard

Challenge: Using proposition bank for Russian, we can automatically project semantic role labels from English to Russian.
Approach: They propose a proposition bank for Russian that automatically projects semantic role labels from English to Russian.
Outcome: The proposed resource automatically projectes semantic role labels from English to Russian.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations