Papers by Yvette Graham
Improving Document-Level Sentiment Analysis with User and Product Context (2020.coling-main)
Copied to clipboard
| Challenge: | Existing work that improves document-level sentiment analysis by encoding user and product information has been limited to considering only the text of the current review. |
| Approach: | They propose to incorporate all available historical review text belonging to the author of the review in question and investigate the inclusion of his- torical reviews associated with the current product. |
| Outcome: | The proposed model improves on IMDB, Yelp 2013 and Yelpan 2014 datasets by more than 2 percentage points in the best case. |
Assessing Human-Parity in Machine Translation on the Segment Level (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent machine translation shared tasks have shown top-performing systems to tie or outperform human translation. |
| Approach: | They examine the outputs of top-performing systems in a recent machine translation shared task . they find that some systems outperform human translation on average . |
| Outcome: | a new method identifies segments for which human and machine perform poorly . the results show that top-performing systems outperform human translation on average . |
The Second Multilingual Surface Realisation Shared Task (SR’19): Overview and Evaluation Results (D19-63)
Copied to clipboard
| Challenge: | EMNLP’19 Workshop on Multilingual Surface Realisation aims to stimulate the exploration of advanced neural networks for multilingual sentence generation from Universal Dependency (UD) structures. |
| Approach: | They present results from the SR'19 Shared Task, a multilingual surface realisation task organised as part of the EMNLP'19 Workshop on Multilingual Surface Realisation. |
| Outcome: | The SR'19 shared task was organised as part of the EMNLP'19 Workshop on Multilingual Surface Realisation . it consisted of two tracks with different levels of complexity . the shallow track was offered in eleven, and the deep track in three languages . |
Semantic-Aware Dynamic Retrospective-Prospective Reasoning for Event-Level Video Question Answering (2023.acl-srw)
Copied to clipboard
| Challenge: | Event-Level Video Question Answering (EVQA) requires complex reasoning across video events to obtain the visual information needed to provide optimal answers. |
| Approach: | They propose a semantic-aware dynamic retrospective-prospective reasoning approach for video-based question answering that explicitly uses the Semantic Role Labeling (SRL) structure of the question in the dynamic reasoning process. |
| Outcome: | The proposed model outperforms existing models on a trafficQA benchmark dataset. |
Exploiting Rich Textual User-Product Context for Improving Personalized Sentiment Analysis (2023.findings-acl)
Copied to clipboard
| Challenge: | Typical approaches do not exploit the potential of historical reviews or do not make full use of user/product associations. |
| Approach: | They propose to use historical reviews to initialize user and product representations and incorporate textual associations via a user-product cross-context module. |
| Outcome: | The proposed method outperforms existing state-of-the-art models on IMDb, Yelp and Longformer benchmarks. |
Statistical Power and Translationese in Machine Translation Evaluation (2020.emnlp-main)
Copied to clipboard
| Challenge: | a recent paper argues that translationese has been used to describe features of translated text . a translationed text can be more explicit than the original source, authors say . authors recommend reverse-created test data be omitted from future evaluations . |
| Approach: | They propose to omit translationese from future machine translation evaluations . they also re-evaluate a past evaluation claiming human-parity of MT . |
| Outcome: | The proposed analysis shows that translationese does not affect machine translation evaluations. |
Do Stochastic Parrots have Feelings Too? Improving Neural Detection of Synthetic Text via Emotion Recognition (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in generative AI have shone a spotlight on high-performance synthetic text generation technologies. |
| Approach: | They propose to use emotion-driven pretrained language models to generate synthetic text that lacks emotional coherence. |
| Outcome: | The proposed detector achieves significant improvements across a range of synthetic text generators, various sized models, datasets, and domains. |
BERTHA: Video Captioning Evaluation Via Transfer-Learned Human Assessment (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing metrics to evaluate video captioning systems are based on overlap between the caption and the reference sentence, but fail to include the context of the scene. |
| Approach: | They propose a method to evaluate video captioning systems using a deep learning model . the model is based on a language model that has been shown to work well in NLP tasks . |
| Outcome: | The proposed model outperforms the most commonly used metrics in video to text tasks. |
Improving Unsupervised Question Answering via Summarization-Informed Question Generation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Question Generation (QG) is the production of meaningful questions given a set of input passages and corresponding answers. |
| Approach: | They propose a method which uses questions generated heuristically from news summaries as a source of training data for a QG system. |
| Outcome: | The proposed method outperforms previous unsupervised models on three in-domain datasets and three out-of-domain ones. |
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)
Copied to clipboard
| Challenge: | Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed . |
| Approach: | They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues . |
| Outcome: | The proposed method is highly reliable while remaining feasible and low cost. |