Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop

60 papers
Distributed Knowledge Based Clinical Auto-Coding System (P19-2)

Copied to clipboard

Challenge: Codification of free-text clinical narratives has long been recognised to be beneficial for secondary uses such as funding, insurance claim processing and research.
Approach: They propose to use NLP and related machine learning techniques to assign ICD-10-AM and ACHI codes to clinical records using local and international standards.
Outcome: The proposed system utilises NLP and ML techniques to assign ICD-10-AM and ACHI codes to clinical records while adhering to local and international standards.
Robust to Noise Models in Natural Language Processing Tasks (P19-2)

Copied to clipboard

Challenge: Existing spelling correction systems are far from perfect for noise-sensitive texts . a new way to handle noise is to make models robust to noise.
Approach: They propose a robust to noise word embeddings model which outperforms existing models in different tasks.
Outcome: The proposed model outperforms existing models in three downstream tasks and shows improvements in noise robustness over existing models.
A Computational Linguistic Study of Personal Recovery in Bipolar Disorder (P19-2)

Copied to clipboard

Challenge: Mental health research can benefit from computational linguistics methods given the abundant availability of language data in the internet and advances in computational tools.
Approach: They will collect and analyse social media data of individuals diagnosed with bipolar disorder with regard to their recovery experiences.
Outcome: The proposed method will analyse first-person accounts shared online in large quantities representing unstructured settings and a more heterogeneous, multilingual population to draw a better picture of the aspects and mechanisms of recovery in bipolar disorder.
Measuring the Value of Linguistics: A Case Study from St. Lawrence Island Yupik (P19-2)

Copied to clipboard

Challenge: a recent study has called into question the utility of linguistics in the development of computational systems.
Approach: a new research proposes to integrate linguistics into a neural morphological analyzer for a polysynthetic language . the researchers propose to use linguistic elements to improve performance in low-resource settings .
Outcome: The proposed analysis shows that linguistics can improve performance in low-resource and high-resolution settings.
Not All Reviews Are Equal: Towards Addressing Reviewer Biases for Opinion Summarization (P19-2)

Copied to clipboard

Challenge: Existing research focuses on mining for opinions from review texts and ignores reviewers.
Approach: They propose to model reviewer biases from review texts and learn a bias-aware opinion representation.
Outcome: The proposed method includes balanced opinions from reviewers with different biases and preferences.
Towards Turkish Abstract Meaning Representation (P19-2)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) abstracts away from syntactic features such as word order and does not annotate every constituent in a sentence.
Approach: They have built a first Turkish AMR corpus by hand-annotating 100 sentences from the novel "The Little Prince" they will use the results to prepare a Turkish AML annotation specification for future annotators.
Outcome: The results of the study compare Turkish AMRs with English AMR annotations . the proposed framework is expected to be used in training future annotators.
Gender Stereotypes Differ between Male and Female Writings (P19-2)

Copied to clipboard

Challenge: a new study quantitatively evaluates gender stereotypes in written language . female writings contain fewer gender stereotype scores than male writings .
Approach: They quantitatively evaluate and analyze gender stereotypes in written language . they compare writings by female authors with writings from male authors .
Outcome: The results show that writings by female authors have lower gender stereotype scores . the authors plan on using more datasets over the past century to study gender stereotypes .
Question Answering in the Biomedical Domain (P19-2)

Copied to clipboard

Challenge: False positive questions require specific knowledge, common sense or a procedure due to ambiguity or the scope of the question.
Approach: False q is a question answering technique that uses natural language to find an answer . Falsity is based on a lexical gap and quality of answer spans .
Outcome: Using the proposed system, patients can self-diagnose without sacrificing quality of answer spans.
Knowledge Discovery and Hypothesis Generation from Online Patient Forums: A Research Proposal (P19-2)

Copied to clipboard

Challenge: Unprompted patient experiences on patient forums contain a wealth of unexploited knowledge.
Approach: They propose to develop automated methods for mining, aggregating and cross-linking patient knowledge from online forums.
Outcome: The proposed methods could be compared with biomedical literature and provide hypotheses for future clinical research.
Automated Cross-language Intelligibility Analysis of Parkinson’s Disease Patients Using Speech Recognition Technologies (P19-2)

Copied to clipboard

Challenge: PD is the second most common neurodegenerative disorder after Alzheimers disease . speech impairments are one of the earliest manifestations in PD patients .
Approach: They propose to analyze the speech signals of PD patients and healthy control subjects in three different languages: German, Spanish, and Czech.
Outcome: The proposed model can discriminate between PD patients and HC subjects even when the language used for train and test is different.
Natural Language Generation: Recently Learned Lessons, Directions for Semantic Representation-based Approaches, and the Case of Brazilian Portuguese Language (P19-2)

Copied to clipboard

Challenge: Natural Language Generation (NLG) is a promising area in Natural Language Processing (NLP) .
Approach: They present a review of the literature on Natural Language Generation in Brazilian Portuguese.
Outcome: The proposed approaches are based on the Abstract Meaning Representation formalism and have potential future directions.
Long-Distance Dependencies Don’t Have to Be Long: Simplifying through Provably (Approximately) Optimal Permutations (P19-2)

Copied to clipboard

Challenge: Neural models at the sentence level often need to model the interaction between words . however, there is no guarantee that the standard ordering of words is computationally efficient or optimal .
Approach: They propose to use a dependency parse as a proxy for inter-word dependencies in a sentence to simplify the sentence with combinatorial objectives imposed on the sentence-parse pair.
Outcome: The proposed model improves classification accuracy and reduces classification error by 2.0% over the previous state of the art.
Predicting the Outcome of Deliberative Democracy: A Research Proposal (P19-2)

Copied to clipboard

Challenge: Deliberative dialogue is a structured, face-to-face method of public interaction that is fundamental to the concept of deliberative democracy.
Approach: They propose to use a combination of lexical, sentiment, durational and further ‘derivative’ features of adjacency pairs to train traditional classification models.
Outcome: The proposed method improves the accuracy of classification models and prediction tasks and shows that the task of recognising agreement is demanding but possible.
Active Reading Comprehension: A Dataset for Learning the Question-Answer Relationship Strategy (P19-2)

Copied to clipboard

Challenge: Literature in quality learning suggests that task performance should also be evaluated on the undergone process to answer.
Approach: They propose to use the Question-Answer Relationship (QAR) to evaluate a reader's ability to select different sources of information depending on the question type.
Outcome: The proposed model will be used to evaluate reading comprehension with weak supervision.
Paraphrases as Foreign Languages in Multilingual Neural Machine Translation (P19-2)

Copied to clipboard

Challenge: Unlike previous studies that use paraphrases at the word/phrase level, we train on parallel paraphrase training on closely related languages.
Approach: They train on parallel paraphrases in the style of multilingual Neural Machine Translation (NMT) they train on translations of the whole corpus that are consistent in structure as paraphrase versions at the corpus level.
Outcome: The proposed training on paraphrases outperforms the baselines on two languages and improves lexical choice and entropy.
Improving Mongolian-Chinese Neural Machine Translation with Morphological Noise (P19-2)

Copied to clipboard

Challenge: Existing models for Mongolian-Chinese translation are based on recurrent, convolutional neural networks or completely eliminate recurrence connections.
Approach: They propose a adversarial training model to alleviate the UNK problem in Mongolian-Chinese machine translation by adding a screener to the model to emphasize the added Mongolian morphological noise.
Outcome: The proposed model reduces training time and improves accuracy in Mongolian-Chinese translation tasks.
Unsupervised Pretraining for Neural Machine Translation Using Elastic Weight Consolidation (P19-2)

Copied to clipboard

Challenge: Neural machine translation (NMT) uses sequence to sequence architectures, but requires a huge amount of parallel data.
Approach: They use Elastic Weight Consolidation to regularize weights of two language models . they then fine-tune the model on parallel data to avoid forgetting the original task .
Outcome: The proposed method achieves BLEU scores similar to the previous work, but is slower and requires less training data.
Māori Loanwords: A Corpus of New Zealand English Tweets (P19-2)

Copied to clipboard

Challenge: Mori loanwords are widely used in New Zealand English for various social functions by New Zealanders within and outside of the Mi community.
Approach: They present a corpus of New Zealand English tweets containing selected Mori words that are likely to be known by New Zealanders who do not speak Mi.
Outcome: The results show that over 30% of the Mori loanwords in tweets are irrelevant . they were manually annotated and used to train machine learning models to filter out irrelevant tweets.
Ranking of Potential Questions (P19-2)

Copied to clipboard

Challenge: linguistic theories of discourse structure view questions and their answers as the main structuring element in discourse.
Approach: They propose a ranking system for questions by their appropriateness in a dialogue . system implements constraints and principles put forward by linguistic theories .
Outcome: The proposed system implements constraints and principles put forward in the linguistic literature.
Controlling Grammatical Error Correction Using Word Edit Rate (P19-2)

Copied to clipboard

Challenge: Existing models for grammatical error correction only consider the single degree of correction suited for training corpus.
Approach: They propose a neural grammar error correction method that can control the degree of correction by using new training data annotated with word edit rate.
Outcome: The proposed method improves correction accuracy by using training data annotated with word edit rate.
From Brain Space to Distributional Space: The Perilous Journeys of fMRI Decoding (P19-2)

Copied to clipboard

Challenge: Recent work in cognitive neuroscience has introduced models for predicting distributional word meaning representations from brain imaging data.
Approach: They propose to use several alternative measures to evaluate the predicted distributional space against a corpus-derived distributional spatial space.
Outcome: The proposed model performs poorly on the most common metrics, while still delivering promising results.
Towards Incremental Learning of Word Embeddings Using Context Informativeness (P19-2)

Copied to clipboard

Challenge: In this paper, we investigate the task of learning word embeddings from very sparse data in an incremental, cognitively-plausible way.
Approach: They propose a model that incorporates informativeness into a proposed model of nonce learning, using it for context selection and learning rate modulation.
Outcome: The proposed model is based on a proposed model of nonce learning, and it performs well on the task of learning new words from definitions and potentially uninformative contexts.
A Strong and Robust Baseline for Text-Image Matching (P19-2)

Copied to clipboard

Challenge: Text-image matching is one of the most popular methods for training text-image embeddings.
Approach: They propose to use a kNN-margin loss that utilizes hard negatives and is robust to noise . they advocate using Inverted Softmax and Cross-modal Local Scaling during inference .
Outcome: The proposed loss function is robust to noise and pseudo negatives are tolerable . the proposed loss functions improve scores of all metrics by a large margin .
Incorporating Textual Information on User Behavior for Personality Prediction (P19-2)

Copied to clipboard

Challenge: Recent studies have shown that textual information of user posts and user behaviors are useful for predicting the personality of social media users.
Approach: They propose to use textual information of user behaviors to predict personality of Twitter users by taking user behaviors into account.
Outcome: The proposed models can predict personality of users who do not post frequently, while taking user behaviors into account.
Corpus Creation and Analysis for Named Entity Recognition in Telugu-English Code-Mixed Social Media Data (P19-2)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of Information Extraction in NLP.
Approach: They present a Telugu-English code-mixed corpus with the corresponding named entity tags.
Outcome: The proposed model scored 0.96, 0.94 and 0.95 on a Telugu-English code-mixed corpus.
Joint Learning of Named Entity Recognition and Entity Linking (P19-2)

Copied to clipboard

Challenge: Named entity recognition and entity linking are two fundamentally related tasks . most approaches focus on the mention detection part, assuming the correct mentions have been detected .
Approach: They perform joint learning of named entity recognition and entity linking to leverage their relatedness.
Outcome: The proposed model achieves competitive results with the state-of-the-art in both NER and EL tasks.
Dialogue-Act Prediction of Future Responses Based on Conversation History (P19-2)

Copied to clipboard

Challenge: Sequence-to-sequence models are a common approach to develop chatbots, but they are prone to a black-box response generation process.
Approach: They propose a method to predict a DA of the next response based on the history of previous utterances and their DAs.
Outcome: The proposed model achieves 10.8% higher F1-score and 3.0% higher accuracy on DA prediction compared to baseline using only a single utterance .
Computational Ad Hominem Detection (P19-2)

Copied to clipboard

Challenge: ad hominem attacks are introduced in debates as an easy win, but their impact on argumentation is limited . a machine learning approach to detect the personal attack is insufficient, we show .
Approach: They propose a machine learning approach that detects ad hominem attacks using social media data . they propose TF-IDF approaches that are insufficient to detect the personal attack .
Outcome: The proposed method has a recall of 80% for a social media data source.
Multiple Character Embeddings for Chinese Word Segmentation (P19-2)

Copied to clipboard

Challenge: Chinese word segmentation is regarded as character-based sequence labeling task in most current work but it neglects important fact: Chinese characters contain both semantic and phonetic meanings.
Approach: They propose a shared bi-LSTM-CRF model which fuses linguistic features efficiently by sharing the LSTM network during the training procedure.
Outcome: The proposed model achieves state-of-the-art in AS and CityU corpora without external lexical resources.
Attention over Heads: A Multi-Hop Attention for Neural Machine Translation (P19-2)

Copied to clipboard

Challenge: Existing multihop attentions for machine comprehension are recurrent and hierarchical . a proposed multi-hop attention for the Transformer refines the attention for an output symbol many times .
Approach: They propose a multi-hop attention for the Transformer which integrates attentions from each head.
Outcome: The proposed model outperforms the baseline Transformer in terms of translation accuracy and speed.
Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function (P19-2)

Copied to clipboard

Challenge: Existing methods to reduce gender bias in natural language datasets are inadequate.
Approach: They propose a loss function modification approach which equalizes the probabilities of male and female words in the output.
Outcome: The proposed approach outperforms existing methods in several aspects, especially in reducing gender bias in occupation words.
Automatic Generation of Personalized Comment Based on User Profile (P19-2)

Copied to clipboard

Challenge: Experimental results show that our model can generate natural, human-like and personalized comments.
Approach: They propose a model that takes user profile into account when generating comments on social media and integrates it with a gated memory.
Outcome: The proposed model can generate natural, human-like and personalized comments on social media.
From Bilingual to Multilingual Neural Machine Translation by Incremental Training (P19-2)

Copied to clipboard

Challenge: Existing approaches to multilingual neural machine translation are based on task specific models and the addition of one more language is only possible by retraining the whole system.
Approach: They propose a training schedule that scales to more languages without modification of previous components.
Outcome: The proposed training schedule shows close results to state-of-the-art in the WMT task.
STRASS: A Light and Effective Method for Extractive Summarization Based on Sentence Embeddings (P19-2)

Copied to clipboard

Challenge: Summarization is a costly and timedemanding task.
Approach: They propose an extractive text summarization method which leverages the semantic information in existing sentence embedding spaces.
Outcome: The proposed method performs similarly to state-of-the-art extractive methods with effective training and inference time.
Attention and Lexicon Regularized LSTM for Aspect-based Sentiment Analysis (P19-2)

Copied to clipboard

Challenge: End-to-end deep learning systems lack flexibility as one cannot adjust the network to fix an obvious problem.
Approach: They propose a way to leverage lexicon information to make the model more flexible . they also explore the effect of regularizing attention vectors to allow the network to have a broader "focus"
Outcome: The proposed approach leverages lexicon information to make it more flexible and robust.
Controllable Text Simplification with Lexical Constraint Loss (P19-2)

Copied to clipboard

Challenge: Existing models that only consider the sentence level generate words beyond the target level.
Approach: They propose a method to control the level of a sentence in a text simplification task . they add the target grade level as input and weight words in the loss function .
Outcome: The proposed method improves both BLEU and SARI scores and achieves aggressive rewriting.
Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches (P19-2)

Copied to clipboard

Challenge: a study using non-canonical text normalization shows that it can surpass the current best performing system by a large margin.
Approach: They propose a fully automated, context-aware machine translation approach with fewer stages of processing.
Outcome: The proposed approach surpasses the current best-performing system by a large margin . the proposed method is more data-hungry and more data sensitive than other methods .
ARHNet - Leveraging Community Interaction for Detection of Religious Hate Speech in Arabic (P19-2)

Copied to clipboard

Challenge: Existing methods to detect hate speech in Arabic rely on textual cues and social network graphs.
Approach: They propose to use Arabic word embeddings and social network graphs to profile hate speech in Arabic.
Outcome: The proposed model incorporates Arabic Word Embeddings and Social Network Graphs for the detection of religious hate speech in Arabic.
Investigating Political Herd Mentality: A Community Sentiment Based Approach (P19-2)

Copied to clipboard

Challenge: polarities inherent in political speeches and debates pose an important problem today.
Approach: They propose to use community-based graphs to augment hand-crafted features based on topic modeling and emotion detection on debate transcripts.
Outcome: The proposed approach surpasses the benchmark results on the same dataset.
Transfer Learning Based Free-Form Speech Command Classification for Low-Resource Languages (P19-2)

Copied to clipboard

Challenge: Current speech-based user interfaces use data intensive methodologies to recognize free-form speech commands, but this is not viable for low-resource languages, which lack speech data.
Approach: They propose a method to develop a domain-specific speech command classification system using speech data from a high-resource language.
Outcome: The proposed system is robust to low-resource languages with limited speech data . the proposed system achieves significant results for Sinhala and Tamil datasets .
Embedding Strategies for Specialized Domains: Application to Clinical Entity Recognition (P19-2)

Copied to clipboard

Challenge: Off-the-shelf word embeddings tend to perform poorly on texts from specialized domains such as clinical reports.
Approach: They combine off-the-shelf contextual embeddings with static word2vec embedders trained on a small in-domain corpus built from task data to reach and sometimes outperform representations learned from a large corpus in the medical domain.
Outcome: The proposed embedding strategies outperform representations learned from a large corpus in the medical domain.
Enriching Neural Models with Targeted Features for Dementia Detection (P19-2)

Copied to clipboard

Challenge: In the United States, adults over 65 are expected to comprise one-fifth of the population by 2030, and a larger proportion of the . population than those under 18 by 2035.
Approach: They propose a neural model that takes into account both long language samples and hand-crafted linguistic features to distinguish between dementia affected and healthy patients.
Outcome: The proposed model achieves an F1 score of 0.929 on the DementiaBank dataset and the state-of-the-art on the dataset.
English-Indonesian Neural Machine Translation for Spoken Language Domains (P19-2)

Copied to clipboard

Challenge: Neural machine translation (NMT) is a data-driven method that requires a large amount of data to build a robust model.
Approach: They conduct a study on Neural Machine Translation (NMT) for English-Indonesian and Indonesian-English (ID-EN) they build NMT systems using the Transformer model for both translation directions and implement domain adaptation method to train pre-trained NMT on speech language data.
Outcome: The proposed model can learn formal translation outputs for English-Indonesian and Indonesian-English (ID-EN) given a small dataset of speech-styled language and a larger dataset of less formal language, the proposed model will be useful for learning formality level.
Improving Neural Entity Disambiguation with Graph Embeddings (P19-2)

Copied to clipboard

Challenge: Entity Disambiguation (ED) is the task of linking an ambiguous entity mention to a corresponding entry in a knowledge base.
Approach: They propose a method that integrates structured information from the knowledge base with unstructured information from text-based representations.
Outcome: The proposed method improves on a graph of hyperlinks between Wikipedia articles and a state-of-the-art neural ED model.
Hierarchical Multi-label Classification of Text with Capsule Networks (P19-2)

Copied to clipboard

Challenge: In hierarchical multi-label classification, samples are classified into one or multiple class labels organized in a structured label hierarchy.
Approach: They apply and compare shallow capsule networks for hierarchical multi-label text classification and introduce a new real-world scenario dataset.
Outcome: The proposed model outperforms neural networks and non-neural network architectures on a real-world scenario dataset.
Convolutional Neural Networks for Financial Text Regression (P19-2)

Copied to clipboard

Challenge: Recent studies have defined forecasting financial volatility from annual reports as text regression problem.
Approach: They propose to replace word features with word embedding vectors to remove lexicon dependency.
Outcome: The proposed model provides more accurate volatility predictions than lexicon based models.
Sentiment Analysis on Naija-Tweets (P19-2)

Copied to clipboard

Challenge: Existing methods for analysing sentiments in social media do not consider the issue of ambiguity that evolves in their usage.
Approach: They propose to leverage on local knowledge bases and adapted Lesk algorithm to facilitate pre-processing of social media feeds.
Outcome: The proposed framework improves on existing methods in extracting sentiments from Nigeria-origin tweets with an accuracy of 99.17%.
Fact or Factitious? Contextualized Opinion Spam Detection (P19-2)

Copied to clipboard

Challenge: In analytic analysis of fake reviews, we compare a number of techniques to detect them.
Approach: They propose a machine learning approach that fine-tunes contextualised embeddings to detect fake reviews.
Outcome: The proposed approach fine-tunes state of the art contextualised embeddings to show that it is effective in detecting fake reviews and lay the groundwork for future research in this area.
Scheduled Sampling for Transformers (P19-2)

Copied to clipboard

Challenge: Existing studies show that scheduled sampling can be applied to recurrent neural networks to avoid exposure bias.
Approach: They propose to use teacher forced embeddings and model predictions to avoid exposure bias in sequence-to-sequence generation.
Outcome: The proposed technique achieves performance close to a teacher-forcing baseline on two language pairs and is promising for future research.
BREAKING! Presenting Fake News Corpus for Automated Fact Checking (P19-2)

Copied to clipboard

Challenge: a new study shows that fake news spreads faster than mainstream articles on the same topic . however, there is no dataset containing compelling fake and questionable news articles .
Approach: They introduce manually verified corpus of compelling fake and questionable news articles on the USA politics . they plan to extend the corpus in the future and use it for automated fake news detection.
Outcome: The proposed model is based on linguistic features and will be extended in the future . it will be used to improve the existing model and improve the tools in the field of fake news detection .
Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon (P19-2)

Copied to clipboard

Challenge: Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders.
Approach: They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content.
Outcome: The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical.
De-Mixing Sentiment from Code-Mixed Text (P19-2)

Copied to clipboard

Challenge: Code-mixing is the phenomenon of mixing the vocabulary and syntax of multiple languages in the same sentence.
Approach: They propose a hybrid architecture for the task of Sentiment Analysis of English-Hindi code-mixed data using CNNs to generate subword representations for the sentences.
Outcome: The proposed architecture achieves 83.54% accuracy and 0.827 F1 score on a benchmark dataset.
Unsupervised Learning of Discourse-Aware Text Representation for Essay Scoring (P19-2)

Copied to clipboard

Challenge: Existing document embedding approaches focus on capturing sequences of words in documents . however, some document classification and regression tasks need to consider discourse structure of text .
Approach: They propose an unsupervised approach to capture discourse structure in terms of coherence and cohesion for document embedding that does not require expensive parsers or annotation.
Outcome: The proposed method improves essay Organization scoring and Argument Strength scoring.
Multimodal Logical Inference System for Visual-Textual Entailment (P19-2)

Copied to clipboard

Challenge: Recent studies of multimodal inference provide challenging tasks such as visual question answering and visual reasoning.
Approach: They propose an unsupervised multimodal logical inference system that can prove entailment relations between texts and images by combing semantic parsing and theorem proving.
Outcome: The proposed system can handle semantically complex sentences for visual-textual inference.
Deep Neural Models for Medical Concept Normalization in User-Generated Texts (P19-2)

Copied to clipboard

Challenge: a medical concept normalization problem is a challenge since social media texts are ambiguous and noisy . a recent study shows that neural architectures leverage the semantic meaning of the entity mention .
Approach: They propose to map a health-related entity mention to a controlled vocabulary . they use powerful neural networks and contextualized word representation models .
Outcome: The proposed model outperforms existing state-of-the-art models in mapping medical concepts to medical terms . the proposed model is based on recurrent neural networks and contextualized word representation models .
Using Semantic Similarity as Reward for Reinforcement Learning in Sentence Generation (P19-2)

Copied to clipboard

Challenge: Existing models for sentence generation use cross-entropy loss as the loss function . however, cross-etropy is unable to evaluate sentences as a whole and lacks flexibility . et al., 2018: a novel approach to improve sentence generation models .
Approach: They propose a method to train a model using estimated semantic similarity between output and reference sentences to alleviate cross-entropy loss problems.
Outcome: The proposed model improves the BLEU scores from the baseline LSTM NMT model.
Sentiment Classification Using Document Embeddings Trained with Cosine Similarity (P19-2)

Copied to clipboard

Challenge: Existing document embedding models map each document to a dense, low-dimensional vector in continuous vector space.
Approach: They propose to train document embeddings using cosine similarity instead of dot product . they focus on document sentiment classification of long movie reviews .
Outcome: The proposed model improves document embedding accuracy by using cosine similarity instead of dot product on the IMDB dataset while using feature combination with Naive Bayes weighted bag of n-grams achieves 93.68% accuracy.
Detecting Adverse Drug Reactions from Biomedical Texts with Neural Networks (P19-2)

Copied to clipboard

Challenge: Detection of adverse drug reactions in post-marketing period is a crucial challenge for pharmacology.
Approach: They propose to use social media to extract information about adverse drug reactions . they compare four state-of-the-art attention-based neural networks to the F-measure .
Outcome: The proposed methods perform better on four different benchmarks.
Annotating and Analyzing Semantic Role of Elementary Units and Relations in Online Persuasive Arguments (P19-2)

Copied to clipboard

Challenge: Existing studies on the features of persuasiveness focus on lexical features and argumentative features.
Approach: They propose an annotation scheme that captures the semantic role of arguments in a popular online persuasion forum, ChangeMyView.
Outcome: The proposed scheme captures the semantic role of arguments in a popular online persuasion forum, so-called ChangeMyView.
A Japanese Word Segmentation Proposal (P19-2)

Copied to clipboard

Challenge: Current word segmentation methods may produce different segmentations for the same strings . this occurs when strings appear in different sentences .
Approach: They propose to use Japanese word segmentation methods that use a morpheme-based approach to produce different segmentations for the same strings.
Outcome: The proposed method produces much more consistent segmentation than the current morpheme-based one.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations