Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop

39 papers
Evaluating zero-shot transfers and multilingual models for dependency parsing and POS tagging within the low-resource language family Tupían (2022.acl-srw)

Copied to clipboard

Challenge: Existing studies on NLP applications for low-resource languages have not been done in this area.
Approach: They propose to replicate the transferability of dependency parsers and POS taggers trained on closely related languages within the low-resource language family Tupan.
Outcome: The proposed models replicate the transferability of dependency parsers and POS taggers trained on closely related languages within the low-resource language family Tupan.
RFBFN: A Relation-First Blank Filling Network for Joint Relational Triple Extraction (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods for relational triple extraction ignore semantic information of relations or predict subjects and objects sequentially.
Approach: They propose a relation-first blank filling network to capture semantic information of relations . they transform relations into relation templates with blanks which contain the fine-grained semantic representation of relations.
Outcome: The proposed model outperforms current state-of-the-art methods on public benchmark datasets.
Building a Dialogue Corpus Annotated with Expressed and Experienced Emotions (2022.acl-srw)

Copied to clipboard

Challenge: a human would recognize the emotion of an interlocutor and respond with an appropriate emotion, such as empathy and comfort.
Approach: They propose to build a dialogue corpus annotated with two kinds of emotions . they collect tweets and annotate them with the emotion they put into the utterance .
Outcome: The proposed method shows that it is difficult to recognize experienced emotions and multitask learning is effective.
Darkness can not drive out darkness: Investigating Bias in Hate SpeechDetection Models (2022.acl-srw)

Copied to clipboard

Challenge: a recent study shows that machine learning models are biased and they might make the wrong decisions for the wrong reasons.
Approach: They investigate the impact of social bias on the performance of hate speech detection models . they also investigate the causal effect of intersectional bias on models' unfairness .
Outcome: The proposed model is biased and makes the wrong decisions for the wrong reasons.
Ethical Considerations for Low-resourced Machine Translation (2022.acl-srw)

Copied to clipboard

Challenge: a paper examines the ethical implications of machine translation for low-resourced languages . a value scenario illustrates potential harms that low-rsourced language communities may face .
Approach: They propose to use Armenian as a case study to investigate ethical implications of machine translation for low-resourced languages.
Outcome: The proposed model is based on a value-scenario model of machine translation for low-resourced languages . the model is used to identify potential harms that low-income speakers may face .
Integrating Question Rewrites in Conversational Question Answering: A Reinforcement Learning Approach (2022.acl-srw)

Copied to clipboard

Challenge: Existing approaches to improve QR performance dependencies among dialogue history dependencies are limited.
Approach: They propose a reinforcement learning approach that integrates QR and CQA tasks without corresponding labeled QR datasets.
Outcome: The proposed approach improves existing pipeline approaches in conversational question answering (QA) existing methods depend on assumption of corresponding QR datasets for every CQA dataset, resulting in poor performance.
What Do You Mean by Relation Extraction? A Survey on Datasets and Study on Scientific Relation Classification (2022.acl-srw)

Copied to clipboard

Challenge: Existing RE surveys focus on modeling techniques, but there are few that are based on real-world scenarios.
Approach: They propose to survey RE datasets and revisit the task definition and its adoption by the community.
Outcome: The proposed approach improves the reliability of RE evaluations across multiple datasets and reveals significant discrepancies in annotations.
Logical Inference for Counting on Semi-structured Tables (2022.acl-srw)

Copied to clipboard

Challenge: Natural Language Inference (NLI) tasks require numerical understanding to perform a numerical type of inference, such as counting.
Approach: They propose a logical inference system for reasoning between semi-structured tables and texts that uses logical representations as meaning representations and model checking to handle a numerical type of inference.
Outcome: The proposed system can perform inference with numerical comparatives with tables and texts in English.
GNNer: Reducing Overlapping in Span-based NER Using Graph Neural Networks (2022.acl-srw)

Copied to clipboard

Challenge: Named Entity Recognition (NER) uses sequence labelling and span classification to identify entities.
Approach: They propose a framework that uses Graph Neural Networks to enrich the span representation to reduce the number of overlapping spans during prediction.
Outcome: The proposed framework reduces the number of overlapping spans while maintaining competitive metric performance.
Compositional Semantics and Inference System for Temporal Order based on Japanese CCG (2022.acl-srw)

Copied to clipboard

Challenge: a system for temporal order in Japanese has not been developed for linguistic inference involving temporal expressions.
Approach: They propose a Japanese NLI system that considers temporal order in Japanese . they use axioms for temporal relations and automated theorem provers to perform inference involving temporal orders.
Outcome: The proposed system outperforms logic-based systems and current deep learning models on Japanese datasets.
Combine to Describe: Evaluating Compositional Generalization in Image Captioning (2022.acl-srw)

Copied to clipboard

Challenge: Recent work on compositionality has focused on the ability to combine simpler concepts to understand & generate arbitrarily more complex conceptual structures.
Approach: They propose to use a set of image captioning models to benchmark their compositional generalization properties.
Outcome: The proposed models do not generalize in terms of systematicity and productivity, but are robust to synonym substitutions.
Towards Unification of Discourse Annotation Frameworks (2022.acl-srw)

Copied to clipboard

Challenge: Discourse information is difficult to represent and annotate, and corpora annotated under different frameworks vary considerably.
Approach: They propose to use automatic means to unify discourse structures and relations . they will also explore the application of the unified framework in multi-task learning and graphical models .
Outcome: The proposed method can be used in multi-task learning and graphical models.
AMR Alignment for Morphologically-rich and Pro-drop Languages (2022.acl-srw)

Copied to clipboard

Challenge: Existing AMR aligners for English are not well suited for many languages where many concepts appear from morphologically-semantic elements.
Approach: They propose to use a tree traversal approach to align AMR concepts from morphemes in a Turkish language.
Outcome: The proposed aligner outperforms the existing aligners for English and Portuguese in terms of precision, recall and F1 score.
Sketching a Linguistically-Driven Reasoning Dialog Model for Social Talk (2022.acl-srw)

Copied to clipboard

Challenge: a new study shows that dialog systems that can hold social talk and make sense of conversational content are not efficient for context-sensitive natural language understanding and reasoning.
Approach: They propose a linguistically-informed architecture to handle social talk in English . they propose linguistic models that fit the context-sensitive components into a Bayesian game-theoretic model .
Outcome: The proposed architecture is based on corpus-based methods but does not track what is happening in a conversation.
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)

Copied to clipboard

Challenge: Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate.
Approach: They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs.
Outcome: The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis.
English-Malay Cross-Lingual Embedding Alignment using Bilingual Lexicon Augmentation (2022.acl-srw)

Copied to clipboard

Challenge: Embedings that are pre-trained monolingually are limited to tasks only in its own language.
Approach: They propose to create English-Malay cross-lingual word embeddings using embedd alignment by exploiting existing language resources.
Outcome: The proposed approach improves the quality of the existing English-Malay bilingual lexicon and the effect of Malay word coverage on the quality.
Towards Detecting Political Bias in Hindi News Articles (2022.acl-srw)

Copied to clipboard

Challenge: Political propaganda in recent times has been amplified by media news portals through biased reporting, creating untruthful narratives on serious issues . a dataset for this task was not available, therefore we developed a transformer-based transfer learning method to fine-tune the pre-trained network on our data.
Approach: They propose a transformer-based transfer learning method to fine-tune the pre-trained network on the data for this bias detection.
Outcome: The proposed method fine-tunes the pre-trained network on the data to detect political bias in Hindi news articles.
Restricted or Not: A General Training Framework for Neural Machine Translation (2022.acl-srw)

Copied to clipboard

Challenge: Existing work imposes constraints on beam search decoding, which limits the concurrent processing ability of the model in deployment.
Approach: They propose a general training framework that allows a model to support both restricted and unrestricted translations by adopting an additional auxiliary training process without constraining the decoding process.
Outcome: The proposed training framework is tested on simulated and original benchmarks.
What do Models Learn From Training on More Than Text? Measuring Visual Commonsense Knowledge (2022.acl-srw)

Copied to clipboard

Challenge: Existing evaluation methods to measure what language models learn from multimodal training are lacking.
Approach: They propose two evaluation tasks to measure commonsense knowledge in language models by using visual data to evaluate multimodal models and unimodal baselines.
Outcome: The proposed evaluation tasks show that training on a visual modality improves on the visual commonsense knowledge in language models.
TeluguNER: Leveraging Multi-Domain Named Entity Recognition with Deep Transformers (2022.acl-srw)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a successful and well-researched problem in English due to the availability of resources.
Approach: They propose to use two annotated NER datasets for the Telugu language . they compare the finetuned Telugus model with the existing model in NER .
Outcome: The proposed models outperform existing models on a large dataset of 38,363 sentences on telugu and other languages.
Using Neural Machine Translation Methods for Sign Language Translation (2022.acl-srw)

Copied to clipboard

Challenge: Sign languages are the main medium of exchanging information for the deaf and hard of hearing.
Approach: They propose to use two NMT architectures to train models on parallel German Sign Language corpora . they achieve substantial improvement in BLEU scores for the models trained on the two corporales .
Outcome: The proposed models achieve significant improvements on the two corpora trained on the german sign language . the proposed models outperform the models trained on both corporales .
Flexible Visual Grounding (2022.acl-srw)

Copied to clipboard

Challenge: Existing visual grounding datasets require queries to be answerable, but in multimedia data, many entities cannot be grounded to the image, resulting in unanswerable visual ground.
Approach: They propose a method to ground to a pseudo image region for unanswerable queries . they add a query that cannot be grounded to the image and train it to ground .
Outcome: The proposed model can handle answerable and unanswerable visual grounding with high accuracy on the proposed datasets.
A large-scale computational study of content preservation measures for text style transfer and paraphrase generation (2022.acl-srw)

Copied to clipboard

Challenge: Text style transfer and paraphrases generation are growing areas of NLP . many researchers still use BLEU-like measures to evaluate content preservation .
Approach: They compare 57 different measures based on different principles on 19 annotated datasets . they find that measures relying on cross-encoder models outperform alternative approaches .
Outcome: The proposed methods outperform traditional methods on 19 datasets.
Explicit Object Relation Alignment for Vision and Language Navigation (2022.acl-srw)

Copied to clipboard

Challenge: Existing work on vision and language navigation grounding the landmarks and spatial relations in textual instructions into visual modality is important.
Approach: They propose a neural agent to explicitly align the spatial information in both instruction and visual environment, including landmarks and spatial relationships between the agent and landmarks.
Outcome: The proposed method surpasses the baseline on the R2R dataset and shows that it can explain spatial reasoning and spatial relationships.
Mining Logical Event Schemas From Pre-Trained Language Models (2022.acl-srw)

Copied to clipboard

Challenge: a pre-trained language model is induced into acting as a distribution over stories, a new system is proposed . NESL is a neural event schema learning system that combines large language models, FrameNet parsing, and simple behavioral schemas to bootstrap the learning process.
Approach: They propose a neural event schema learning system that bootstraps the learning process by parsing pre-trained language models into simple behavioral schemas.
Outcome: The proposed system combines large language models, a powerful logical representation of language, and simple behavioral schemas to bootstrap the learning process.
Exploring Cross-lingual Text Detoxification with Large Multilingual Language Models. (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods of textual style transfer are monolingual i.e. designed to work in one exact language.
Approach: They propose to make large multilingual models capable of performing multilingual style transfer without direct fine-tuning in a given language.
Outcome: The proposed model can generate text in a given language without fine-tuning and is able to perform cross-lingual detoxification without direct fine- tuning.
MEKER: Memory Efficient Knowledge Embedding Representation for Link Prediction and Question Answering (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods to embed learning use a standard Neural Networks (NN) backward mechanism, duplicating its memory consumption.
Approach: They propose a memory-efficient KG embedding model that embeds knowledge graphs as 3rd-order binary tensors.
Outcome: The proposed model yields comparable performance on link prediction and KG-based question answering tasks.
Discourse on ASR Measurement: Introducing the ARPOCA Assessment Tool (2022.acl-srw)

Copied to clipboard

Challenge: Automated speech recognition (ASR) models are based on a corpus of audio recordings, but are often small or nonexistent for less common languages and dialects.
Approach: This research proposal will develop a semi-automatic acoustic features extraction system that integrates phonetic transcripts with pronunciation dictionaries.
Outcome: The proposed system will be used to improve language recognition and model feedback in less common languages and dialects.
Pretrained Knowledge Base Embeddings for improved Sentential Relation Extraction (2022.acl-srw)

Copied to clipboard

Challenge: Existing models that perform explicit on-task training of graph embeddings are inadequate.
Approach: They propose to combine pretrained knowledge base graph embeddings with transformer based language models to improve performance on sentential Relation Extraction task.
Outcome: The proposed model outperforms state-of-the-art models on the sentential Relation Extraction task.
Improving Cross-domain, Cross-lingual and Multi-modal Deception Detection (2022.acl-srw)

Copied to clipboard

Challenge: Deception detection is a deliberate choice to mislead to gain some advantage or avoid some penalty.
Approach: They propose to use inter-domain distance to identify suitable source domain for a given target domain to improve cross-domain deception classification and to better understand multi-modal deception detection.
Outcome: The proposed methods will be able to detect deception in cross-domain, cross-lingual and multi-modal settings and will improve multi-modular deception classification.
Automatic Generation of Distractors for Fill-in-the-Blank Exercises with Round-Trip Neural Machine Translation (2022.acl-srw)

Copied to clipboard

Challenge: a fill-in-the-blank exercise involves removing one word from a sentence and generating distractors . a valid distractor is a word that does not fit the context, and distractors are invalid .
Approach: They propose to automatically generate distractors using round-trip neural machine translation . they show that using hundreds of translations for a given sentence generates a rich set of distractors .
Outcome: The proposed method outperforms two strong baselines against a real corpus of cloze exercises and manually checks for validity.
On the Locality of Attention in Direct Speech Translation (2022.acl-srw)

Copied to clipboard

Challenge: Recent advances in NLP have created problems with the complexity of the self-attention layer.
Approach: They propose to substitute standard self-attention with a local efficient one to avoid the computation of attention weights.
Outcome: The proposed model matches the baseline performance and improves efficiency by skipping the computation of weights that standard attention discards.
Extraction of Diagnostic Reasoning Relations for Clinical Knowledge Graphs (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods for analyzing knowledge graphs focus on concept relations and clinical processes.
Approach: They propose to extract clinical knowledge graphs from a wiki and consumer health resource texts by using a clinical reasoning ontology.
Outcome: The proposed methods evaluate the correctness of extracted triples in the zero-shot setting.
Scene-Text Aware Image and Text Retrieval with Dual-Encoder (2022.acl-srw)

Copied to clipboard

Challenge: Existing studies on image and text retrieval using a dual-encoder model have not shown their effectiveness for fast inferences.
Approach: They propose a dual-encoder model that connects vision and language in the same semantic space and integrates scene-text and visual information into a model.
Outcome: The proposed model can interpret scene-text and surrounding visual information better than cross-encoder models.
Towards Fine-grained Classification of Climate Change related Social Media Text (2022.acl-srw)

Copied to clipboard

Challenge: a new study examines the fine-grained classification and classification of climate change-related social media text.
Approach: They propose to use two datasets to analyze climate change-related social media text and propose a fine-grained classification based on the proposed dataset.
Outcome: The proposed datasets are compared with existing datasets and benchmarked using the best-performing model.
Deep Neural Representations for Multiword Expressions Detection (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods for multiword expression detection are based on sequence labeling and statistical measures.
Approach: They propose a weakly supervised method for multiword expressions extraction . they use a lexicon of English multiword lexical units as a reference knowledge base .
Outcome: The proposed method can be easily applied to other languages.
A Checkpoint on Multilingual Misogyny Identification (2022.acl-srw)

Copied to clipboard

Challenge: a study on hate speech against minorities in Italian tweets found that 1 women are the most targeted group.
Approach: They propose to train monolingual transformers and multilingual transformer models with monolingual data in English, Italian, and Spanish to detect misogyny in tweets.
Outcome: The proposed model achieves state-of-the-art on English, Italian, and Spanish.
Using dependency parsing for few-shot learning in distributional semantics (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods for few-shot learning use dependency parsing information to learn meaning of rare words based on limited amount of context sentences.
Approach: They propose dependency parsing for few-shot learning to learn meaning of rare words . they use word embedding models as background spaces for few shot learning .
Outcome: The proposed methods enhance the additive baseline model by using dependencies.
A Dataset and BERT-based Models for Targeted Sentiment Analysis on Turkish Texts (2022.acl-srw)

Copied to clipboard

Challenge: Sentiment analysis is a field that is growing due to the availability of the Internet and the growing number of online platforms.
Approach: They propose an annotated Turkish dataset suitable for targeted sentiment analysis.
Outcome: The proposed models outperform the traditional models for the targeted sentiment analysis task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations