Proceedings of the Eleventh International Conference on Language Resources and Evaluation (

728 papers
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)

Copied to clipboard

Challenge: Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding.
Approach: They propose to augment an existing (monolingual) corpus: LibriSpeech.
Outcome: The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned.
Evaluating Domain Adaptation for Machine Translation Across Scenarios (L18-1)

Copied to clipboard

Challenge: Statistical machine translation (SMT) has been the dominant approach for the last 20 years, with neural machine translation becoming the new main paradigm in academic research and the industry.
Approach: They propose to compare domain-adapted statistical and neural machine translation systems on three different domains and language pairs with varying degrees of domain specificity and available training data.
Outcome: The proposed system is the best choice for translation, with marked impacts for domains with higher specificity.
Upping the Ante: Towards a Better Benchmark for Chinese-to-English Machine Translation (L18-1)

Copied to clipboard

Challenge: Currently, there is no widely accepted standard for evaluation of machine translation (MT) for Chinese-to-English translation, there are no standard for standardized training sets, development sets, and test sets.
Approach: They propose to use Chinese-to-English machine translation as a benchmark . they build a highly competitive state-of-the-art MT system that outperforms reported results .
Outcome: The proposed system outperforms reported results on NIST OpenMT test sets in almost all papers published in major conferences and journals in computational linguistics and artificial intelligence in the past 11 years.
ESCAPE: a Large-scale Synthetic Corpus for Automatic Post-Editing (L18-1)

Copied to clipboard

Challenge: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Approach: a team of researchers develops a Synthetic Corpus for Automatic Post-Editing . eSCAPE is the largest freely-available Synthetic corpus for automatic post-editing released so far . the results prove that the models always improve MT quality with statistically significant gains .
Outcome: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Evaluating Machine Translation Performance on Chinese Idioms with a Blacklist Method (L18-1)

Copied to clipboard

Challenge: idiom translation is a challenging problem in machine translation because meaning is non-compositional and literal translations are likely to be wrong.
Approach: They propose a method to evaluate the quality of idiom translation of MT systems by a blacklist of literal translations.
Outcome: The proposed method detects that a sizable number of idioms are mistranslated (46.1%) and that literal translation error is a common error type.
Network Features Based Co-hyponymy Detection (L18-1)

Copied to clipboard

Challenge: Existing methods to detect lexical relations have been used to identify them in both supervised and unsupervised ways.
Approach: They propose to use distributional semantic models to detect co-hyponymy relation with high accuracy and various network measures to perform better or at par with the state-of-the-art models.
Outcome: The proposed model performs better or at par with the state-of-the-art models.
Cross-Lingual Generation and Evaluation of a Wide-Coverage Lexical Semantic Resource (L18-1)

Copied to clipboard

Challenge: Neural word embedding models are not interpretable for humans by themselves . we present a method that assigns explicit symbolic semantic features to words .
Approach: They propose a method that assigns explicit symbolic semantic features to words in an embedding model . they use a finite list of terms to make the model interpretable for humans .
Outcome: The proposed method is shown to be very efficient for word embedding models . it can be applied across languages and can be used as a searchable semantic annotation .
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
Integrating Generative Lexicon Event Structures into VerbNet (L18-1)

Copied to clipboard

Challenge: Efforts to use the verb lexicon's semantic representations have revealed a need to revise the form to allow for greater flexibility in representing complex events.
Approach: They propose to restrict the form to first-order representations to simplify use by planners and integrate with the Generative Lexicon's event structure.
Outcome: The proposed representations simplify use by and integration with planners and allow for greater flexibility in representing complex events and for a more nuanced portrayal of the Agent's role.
FontLex: A Typographical Lexicon based on Affective Associations (L18-1)

Copied to clipboard

Challenge: a typographical lexicon provides associations between words and fonts . tens of thousands of fonts are available, and font choice affects perception of text, author and brand.
Approach: They create a typographical lexicon providing associations between words and fonts by using affective evocations and word-emotion relationships.
Outcome: The proposed typographical lexicon provides associations between words and fonts using affective evocations and word-emotion relationships.
Multi-layer Annotation of the Rigveda (L18-1)

Copied to clipboard

Challenge: Using a multi-level annotation, we present a corpus of the R. GVEDA .
Approach: They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links.
Outcome: The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments.
The Natural Stories Corpus (L18-1)

Copied to clipboard

Challenge: Existing corpora of naturalistic text do not contain the low-frequency syntactic constructions needed to distinguish between theories.
Approach: They propose to compare models of language processing by comparing their ability to predict behavioral and neural measures of processing difficulty to corpora of naturalistic text.
Outcome: The proposed corpus contains low-frequency syntactic constructions while sounding fluent to native speakers.
Semi-automatic Korean FrameNet Annotation over KAIST Treebank (L18-1)

Copied to clipboard

Challenge: Annotating FrameNet over raw sentences is an expensive and complex task, because of which we have designed a semi-automatic annotation approach.
Approach: They propose to use Korean FrameNet annotations to build a frame-semantic parser for English using full-text annotation and partially annotated exemplar sentences to train their models.
Outcome: The proposed model is based on a lexical database of the Korean FrameNet, and its current scope, status, and limitations are discussed in the paper.
Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text (L18-1)

Copied to clipboard

Challenge: a new approach to POS tagging noisy user generated text is proposed . word embeddings are trained on a noisy corpus to address both normalization and POS.
Approach: They propose to use word embeddings to normalize text before tagging it, while a gated neural network based tagger handles the remaining errors.
Outcome: The proposed approach normalizes some errors before tagging, while a gated neural network handles the remaining errors.
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)

Copied to clipboard

Challenge: Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability.
Approach: They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect.
Outcome: The proposed model can tag four different dialects with an average accuracy of 89.3%.
A Corpus for Modeling Word Importance in Spoken Dialogue Transcripts (L18-1)

Copied to clipboard

Challenge: a project aims to create a system that uses automatic speech recognition (ASR) to produce real-time text captions of spoken English during in-person meetings with hearing individuals.
Approach: They propose to use automatic speech recognition to produce captions in real-time . they add word-importance annotations to a transcript of a conversational dialogue corpus .
Outcome: The proposed system would produce captions in real-time for people who are deaf or hard-of-hearing . the best performing model has an F-score of 0.60 in an ordinal 6-class word-importance classification task with an agreement (concordance correlation coefficient) of 0.89 with the human annotators.
Dialogue Structure Annotation for Multi-Floor Interaction (L18-1)

Copied to clipboard

Challenge: Existing annotation schemes do not address dialogue structure.
Approach: They propose an annotation scheme for meso-level dialogue structure that clusters utterances from multiple participants and floors into units according to realization of an initiator's intent.
Outcome: The proposed annotation scheme is used to annotate a corpus of human-robot interaction dialogues.
Effects of Gender Stereotypes on Trust and Likability in Spoken Human-Robot Interaction (L18-1)

Copied to clipboard

Challenge: a study investigates the influence of gender stereotypes on trust and likability of humanoid robots . explicit gender and stereotypicality of a task are manipulated to influence robot behavior . future research may look into situational variables that drive stereotypification in robot interaction .
Approach: They investigated the influence of gender stereotypes on trust and likability of robots . they used explicit (name and voice) and implicit (personality) genders to manipulate stereotypical tasks . future research may look into situational variables that drive stereotypization .
Outcome: The findings suggest that gender stereotypes need to be differentiated in robot interaction . the gender and personality characteristics of robots influence trust and likability .
A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)

Copied to clipboard

Challenge: Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener.
Approach: They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen.
Outcome: The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object.
Improving Dialogue Act Classification for Spontaneous Arabic Speech and Instant Messages at Utterance Level (L18-1)

Copied to clipboard

Challenge: Existing methods to detect dialogue act from utterances are limited in Arabic dialects . linguistic knowledge of the speaker is important for understanding spontaneous speech and instant messages .
Approach: They propose a statistical dialogue analysis model to automatically recognize dialogue acts from a textual corpus.
Outcome: The proposed model improves the F-measure by 20% . the proposed model can automatically acquire probabilistic discourse knowledge from a dialogue corpus .
Data Management Plan (DMP) for Language Data under the New General Da-ta Protection Regulation (GDPR) (L18-1)

Copied to clipboard

Challenge: ELRA proposes its own template for the Data Management Plan, which is being updated to take the new law into account.
Approach: They propose a framework for the data management plan to be updated to take the new law into account and propose how it can be integrated into the DMP to increase transparency and spread good practices .
Outcome: The proposed framework will strengthen certain principles related to the processing of personal data, which will also affect many projects in the field of natural language processing.
We Are Depleting Our Research Subject as We Are Investigating It: In Language Technology, more Replication and Diversity Are Needed (L18-1)

Copied to clipboard

Challenge: a recent study suggests that language scientists are depleting natural language diversity as they investigate it . around one third of the world languages are at present vulnerable to becoming extinct according to UNESCO .
Approach: They propose that language scientists should take more into account the diversity of natural language in their research activities . they propose that more replication and reproduction and more language diversity should be taken into account in their work .
Outcome: The proposed research should take into account more replication and reproduction and more language diversity . around one third of the world languages are at present vulnerable to become extinct according to UNESCO .
Lessons Learned: On the Challenges of Migrating a Research Data Repository from a Research Institution to a University Library. (L18-1)

Copied to clipboard

Challenge: a paper describes the feasibility of research data migration from one institution to another . a university library and computing centre takes care of all data and guarantees its access for the foreseeable future.
Approach: They propose a migration concept that allows researchers to migrate research data to another institution . they propose archiving research data in a discipline independent data facility .
Outcome: The proposed migration concept supports the stance of data archives that users can expect high levels of trust and reliability when it comes to data safety and sustainability.
Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data (L18-1)

Copied to clipboard

Challenge: a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say .
Approach: They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives .
Outcome: a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation .
Three Dimensions of Reproducibility in Natural Language Processing (L18-1)

Copied to clipboard

Challenge: a recent editorial on reproducibility in language processing defined three dimensions of reproducibility . authors had already submitted a correction, but there is no consensus on the definitions .
Approach: They propose an ontology of reproducibility in natural language processing to address these problems . they propose to analyze three dimensions of reproducible in natural languages papers . authors propose to use a 'replicability' term to describe the reproducibility of a conclusion, finding, value .
Outcome: The proposed ontology aims to enhance future research and communication about the topic and retrospective meta-analyses.
Content-Based Conflict of Interest Detection on Wikipedia (L18-1)

Copied to clipboard

Challenge: Conflict-of-Interest (CoI) editing is a problem on Wikipedia that is highly subjective . a key feature of Wiki sites is to allow people from all over the world to add or modify articles anonymously and without consequence.
Approach: They frame CoI detection as a binary classification problem and explore features for it . they find that stylometric features outperform other types of features and give an F-measure of 0.63 .
Outcome: The proposed method outperforms other features and gives an F-measure of 0.63 . the proposed method is not certain that the set of non-CoI articles contains any CoI articles .
Word Affect Intensities (L18-1)

Copied to clipboard

Challenge: Existing lexicons of affect only show coarse associations, but are not accurate as human-created ones.
Approach: They propose to use a manually created affect intensity lexicon with real-valued intensity scores for anger, fear, joy, and sadness.
Outcome: The lexicon has real-valued scores for anger, fear, joy, and sadness . anger, fears, and sad words have very similar VAD scores .
Representation Mapping: A Novel Approach to Generate High-Quality Multi-Lingual Emotion Lexicons (L18-1)

Copied to clipboard

Challenge: Existing representational frameworks for emotion encoding are incompatible with semantic polarity, resulting in a large amount of incompatible emotion lexicons.
Approach: They propose to map different emotion representation formats onto each other for mutual compatibility and interoperability of language resources.
Outcome: The proposed method produces (near-)gold quality emotion lexicons even in crosslingual settings.
Unfolding the External Behavior and Inner Affective State of Teammates through Ensemble Learning: Experimental Evidence from a Dyadic Team Corpus (L18-1)

Copied to clipboard

Challenge: a study aimed to understand the relationship between the external behavior and inner affective state of two team members during a demanding operational task.
Approach: They assessed the external behavior and inner affective state of two team members during a demanding operational task.
Outcome: The study compared verbal responses of two team members to their external and internal cues . the results suggest that an association exists between turn taking and inner affective state .
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)

Copied to clipboard

Challenge: a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories .
Approach: They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets .
Outcome: The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets .
When ACE met KBP: End-to-End Evaluation of Knowledge Base Population with Component-level Annotation (L18-1)

Copied to clipboard

Challenge: Automating constructing a Knowledge Base from unstructured text is a goal of natural language processing.
Approach: They propose a method to evaluate a Knowledge Base population from unstructured text . they propose bootstrap resampling to provide statistical significance to the results .
Outcome: The proposed method uses component-level annotations to evaluate Cold Start KBP . it also uses bootstrap resampling to provide statistical significance to the results reported .
Simple Large-scale Relation Extraction from Unstructured Text (L18-1)

Copied to clipboard

Challenge: Knowledge-based question answering relies on the availability of facts, most of which cannot be found in structured sources.
Approach: They propose a method for creating distant (weak) supervision labels for training a large-scale RE system by decoupling the model architecture from the feature design of a state-of-the-art neural network system.
Outcome: The proposed method performs on par with the state-of-the-art model with similar features at 75x reduction in training time.
Joint Learning of Sense and Word Embeddings (L18-1)

Copied to clipboard

Challenge: Existing methods for learning lower-dimensional representations of words using unlabelled data learn a single representation for a word, ignoring the different senses of that word (polysemy).
Approach: They propose a method that jointly learns sense-aware word embeddings using both unlabelled and sense-tagged text corpora.
Outcome: The proposed method outperforms competing methods on word similarity and short-text classification benchmark datasets.
Comparing Pretrained Multilingual Word Embeddings on an Ontology Alignment Task (L18-1)

Copied to clipboard

Challenge: Existing word embeddings capture a string's semantics and can be trained for multiple languages.
Approach: They propose to compare three different multilingual pretrained word embedding repositories with a string-matching baseline and use it to compute semantic similarities of strings in different languages.
Outcome: The proposed method produces correct alignments on a non-standard dataset on all four languages.
A Large Resource of Patterns for Verbal Paraphrases (L18-1)

Copied to clipboard

Challenge: Xu et al., 2015: paraphrases play an important role in natural language understanding . he says it is difficult to propose a paraphrasing relation for natural language processing systems .
Approach: They propose a resource of such paraphrases that can be used to identify hidden paraphrase pairs . they propose to use the resource to identify paraphrase relationships between two words .
Outcome: The proposed resource contains tens of thousands of such pairs and is available for academic purposes.
Building Parallel Monolingual Gan Chinese Dialects Corpus (L18-1)

Copied to clipboard

Challenge: In particular, we manually annotate a Gan Chinese Dialects Corpus (GCDC) including 131.5 hours and 310 documents with 6 different genres, containing news, official document, story, prose, poet, letter and speech, from 19 different Gan regions.
Approach: They propose a scheme to represent Gan Chinese dialects using Chinese character, Chinese Pinyin and Chinese audio forms.
Outcome: The proposed scheme is based on a Gan Chinese Dialects Corpus (GCDC) with 131.5 hours and 310 documents with 6 different genres, containing news, official document, story, prose, poet, letter and speech, from 19 different Gan regions.
A Recorded Debating Dataset (L18-1)

Copied to clipboard

Challenge: Existing research in computational argumentation and debating technologies focuses on argumentation mining, but other tasks are being addressed as well.
Approach: They describe a dataset of debating speeches in English that is used for research . they use an automatic speech recognition system to produce a more "nLP-friendly" text .
Outcome: The proposed dataset contains 60 speeches on various controversial topics, each in five formats corresponding to different stages in production.
Building a Corpus from Handwritten Picture Postcards: Transcription, Annotation and Part-of-Speech Tagging (L18-1)

Copied to clipboard

Challenge: In this paper, we describe the processes and challenges of digitalisation, manual transcription, and manual annotation of over 11,000 postcards.
Approach: They describe the processes and challenges of digitalisation, manual transcription, and manual annotation of over 11,000 postcards written in German and Swiss German.
Outcome: The proposed system outperforms state-of-the-art taggers in the evaluation of the 'picture postcard corpus' containing over 11,000 handwritten postcards .
A Lexical Tool for Academic Writing in Spanish based on Expert and Novice Corpora (L18-1)

Copied to clipboard

Challenge: a corpus of academic texts provides lexical combinations for the production of academic text.
Approach: They describe the extraction of data from a corpus of academic texts and the use of those data to develop a lexical tool oriented to the production of academic text.
Outcome: The proposed tool will provide indications as to how to use vocabulary typical of the academic genre in order to build complete texts.
Framing Named Entity Linking Error Types (L18-1)

Copied to clipboard

Challenge: Named Entity Linking (NEL) and relation extraction forms the backbone of Knowledge Base Population tasks.
Approach: They propose a taxonomy to frame common errors and apply it to four well-known Named Entity Linking systems.
Outcome: The proposed taxonomy was applied to four well-known Named Entity Linking systems on three gold standards.
A FrameNet for Cancer Information in Clinical Narratives: Schema and Annotation (L18-1)

Copied to clipboard

Challenge: Existing natural language processing (NLP) systems for cancer-related information are highly task-specific and often produce incompatible annotations and algorithms.
Approach: They propose a general-purpose natural language processing resource for cancer-related information in clinical notes . the project uses a frame semantic method to emphasize the information presented in the notes themselves .
Outcome: The proposed project emphasizes the information presented in the clinical notes and its linguistic structure.
A New Corpus to Support Text Mining for the Curation of Metabolites in the ChEBI Database (L18-1)

Copied to clipboard

Challenge: a corpus of 200 abstracts and 100 full text papers which have been annotated with named entities and relations in the biomedical domain is part of the OpenMinTeD project.
Approach: They propose to annotate 200 abstracts and 100 full text papers with entities and relations in the biomedical domain as part of the OpenMinTeD project.
Outcome: The proposed corpus can be used within ChEBI to facilitate text and data mining and integrate with the OpenMinTeD text and database platform.
Parallel Corpora for the Biomedical Domain (L18-1)

Copied to clipboard

Challenge: Existing corpora of parallel corporata are being used in the biomedical domain . MT is known to support readers' access to textual documents in a language other than their native language .
Approach: They propose to leverage parallel corpora to implement cross-lingual information retrieval or machine translation tools.
Outcome: The proposed corpus is being used in the biomedical task at the conference on machine translation (WMT'16 and WMT'17) it can be leveraged to provide access to health information in languages other than English.
Medical Entity Corpus with PICO elements and Sentiment Analysis (L18-1)

Copied to clipboard

Challenge: In this paper, we establish a PICO and a sentiment annotated corpus of clinical trial publications.
Approach: They propose to create a phrase-level PICO corpus and a sentence-level sentiment annotated corpus from clinical trial publications.
Outcome: The proposed corpus is annotated on a phrase-level and a sentiment annotation on the same corpus.
Word Embedding Approach for Synonym Extraction of Multi-Word Terms (L18-1)

Copied to clipboard

Challenge: MWTs are motivated combinations that clearly convey the concept they designate.
Approach: They propose a word-embedding-based approach for automatic acquisition of MWT synonyms that manage length variability.
Outcome: The proposed approach improves on two specialized domain corpora and shows that it is more efficient than baseline approaches.
A Large Automatically-Acquired All-Words List of Multiword Expressions Scored for Compositionality (L18-1)

Copied to clipboard

Challenge: Existing literature on semantically idiosyncratic multiword expressions is limited to English . idiomatic expressions are phraseological units consisting of more than one lexeme and exhibit some kind of idiom.
Approach: They propose to make available a large automatically-acquired all-words list of English multiword expressions scored for compositionality.
Outcome: The proposed list improves the BLEU scores of the English multiword expressions.
A Hybrid Approach for Automatic Extraction of Bilingual Multiword Expressions from Parallel Corpora (L18-1)

Copied to clipboard

Challenge: Specific-domain bilingual lexicons are composed of MultiWord Expressions (MWEs) the manual construction of MWEs bilingual dictionaries is costly and time-consuming.
Approach: They propose to use word alignment approaches to automatically construct bilingual lexicons of MWEs from parallel corpora by formalizing the alignment process as an integer linear programming problem.
Outcome: The proposed approach extracts and aligns multiword expressions from parallel corpora and then filters them using linguistic patterns to build bilingual lexicons.
No more beating about the bush : A Step towards Idiom Handling for Indian Language NLP (L18-1)

Copied to clipboard

Challenge: idioms are a part of natural language and are difficult to learn with a parallel corpora database.
Approach: They propose to use a parallel idiom dataset to train two NLP subtasks . they show significant improvement in the two subtask training without the idiomatic dataset .
Outcome: The proposed model improves on baseline models with the idiom dataset for two NLP applications.
Sentence Level Temporality Detection using an Implicit Time-sensed Resource (L18-1)

Copied to clipboard

Challenge: Temporal sense detection of any word is an important aspect for detecting temporality at the sentence level.
Approach: They build a temporal resource based on a semi-supervised learning approach . they use past, present, future, neutral and atemporal senses to tag sentences .
Outcome: The proposed resource is based on a semi-supervised learning approach . it is used to tag sentences with past, present and future temporal senses .
Comprehensive Annotation of Various Types of Temporal Information on the Time Axis (L18-1)

Copied to clipboard

Challenge: Existing studies linking event and time information have been conducted to train and evaluate models.
Approach: They propose an annotation scheme that anchors expressions in text to the time axis comprehensively.
Outcome: The proposed scheme can be utilized for integrated information analysis of events, entities and time.
Systems’ Agreements and Disagreements in Temporal Processing: An Extensive Error Analysis of the TempEval-3 Task (L18-1)

Copied to clipboard

Challenge: Temporal Processing systems are crucial for timelines and storylines . TempEval-3 is the latest evaluation campaign on open-domain TP in english .
Approach: They present a Temporal Processing system that incorporates high level lexical semantic features and uses them to evaluate temporal relation classification.
Outcome: The proposed system achieves the best scores for event detection and temporal relation classification from raw text, but the errors are not as robust as previous systems.
Annotating Temporally-Anchored Spatial Knowledge by Leveraging Syntactic Dependencies (L18-1)

Copied to clipboard

Challenge: Existing approaches to extract spatial knowledge focus on extracting locations of events, someone or something.
Approach: They propose a method to annotate temporally-anchored spatial knowledge on top of OntoNotes by crowdsourcing annotations.
Outcome: The proposed method can be automated and validated using syntactic dependencies and crowdsourced annotations.
Contextualized Usage-Based Material Selection (L18-1)

Copied to clipboard

Challenge: Currently, authentic linguistic examples for a given keyword search are organized alphabetically according to context.
Approach: They propose to use NLP-functionalities to organize usage-based examples from corpora . they group retrieved examples on syntactic grounds, then show semantic similarity within phrasal slots .
Outcome: The proposed system would help language learners and other end users to benefit from a distributional linguistic analysis.
CBFC: a parallel L2 speech corpus for Korean and French learners (L18-1)

Copied to clipboard

Challenge: Using corpora for second language acquisition has become more and more common . corporata are used to study morpho-syntactic phenomena in English as a foreign language .
Approach: They propose to use a bilingual corpus of French learners of Korean and Korean learners of French to provide a translated and annotated corpus to the scientific community.
Outcome: The proposed corpus can be used for a wide array of purposes in the field of theoretical but also applied linguistics.
SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning (L18-1)

Copied to clipboard

Challenge: Learning a second language requires exposition to texts, especially for the acquisition of vocabulary.
Approach: They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia .
Outcome: The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency.
Towards a Diagnosis of Textual Difficulties for Children with Dyslexia (L18-1)

Copied to clipboard

Challenge: a study on diagnosing the textual difficulties of children's books is published . it focuses on the passages of the books that are difficult to understand for underage children .
Approach: They propose to diagnose the difficulties appearing in French children's books . they focus on the subject pronouns "il" and "elle" and detect difficult anaphoras .
Outcome: The proposed method detects half of the difficult anaphorical pronouns in french children's books . authors say it is complementary of previous approaches to support dyslexia .
Coreference Resolution in FreeLing 4.0 (L18-1)

Copied to clipboard

Challenge: FreeLing is an open-source library for NLP with more than fifteen years of existence and a widespread user community.
Approach: They propose to port RelaxCor to FreeLing to solve the integration problems . they propose to use two strategies and a rough evaluation of the integration results .
Outcome: The proposed integration of RelaxCor into FreeLing solves the problems found in a shared task scenario.
BASHI: A Corpus of Wall Street Journal Articles Annotated with Bridging Links (L18-1)

Copied to clipboard

Challenge: Bridging resolution is an under-researched area of NLP where the lack of annotated training data makes the application of statistical models difficult.
Approach: They propose to use a corpus resource for the anaphoric phenomenon of bridging to add briding anamorphs to other gold annotations created as part of the OntoNotes project.
Outcome: The proposed corpus adds bridging anaphors and their antecedents to other gold annotations created as part of the OntoNotes project.
SACR: A Drag-and-Drop Based Tool for Coreference Annotation (L18-1)

Copied to clipboard

Challenge: Several annotation strategies have been proposed to balance scientific needs with annotation speed.
Approach: They introduce SACR, an easy-to-use coreference chain annotation tool . it is used to annotate large corpora for natural language processing applications . paper compares several annotation schemes implemented in existing tools .
Outcome: The proposed tool is used to annotate large corpora for natural language processing applications.
Deep Neural Networks for Coreference Resolution for Polish (L18-1)

Copied to clipboard

Challenge: Existing deep neural networks for coreference resolution for Polish have been used to resolve textual fragments that refer to the same entity in the discourse world.
Approach: They propose a system combining the best deep neural architecture and sieve-based coreference resolvers ordered from most to least precise to achieve the highest results.
Outcome: The proposed system improves the state of the art for Polish by 0.53 F1 points, reaching 81.23 points of the CoNLL metric.
SzegedKoref: A Hungarian Coreference Corpus (L18-1)

Copied to clipboard

Challenge: SzegedKoref is a treebank of Hungarian that contains manual annotation at several linguistic layers.
Approach: They introduce a Hungarian corpus in which coreference relations are manually annotated.
Outcome: The proposed corpus can be used in training and testing machine learning based coreference resolution systems.
A Corpus to Learn Refer-to-as Relations for Nominals (L18-1)

Copied to clipboard

Challenge: Existing work on how to learn refer-to-as relations from large unlabeled corpora lacks coreferential information.
Approach: They propose to use Wikipedia to generate coreferential neural embeddings for nominals . they use coreference resolution as a proxy to evaluate the neural embeds for noun phrases .
Outcome: The proposed dataset can be leveraged to construct representations for coreferential nominals from Wikipedia.
Sanaphor++: Combining Deep Neural Networks with Semantics for Coreference Resolution (L18-1)

Copied to clipboard

Challenge: Coreference resolution is a challenging task in Natural Language Processing . since a few years, the biggest step forward has been made using deep neural networks .
Approach: They propose to improve coreference resolution by adding semantic features to a top-level deep neural network system . they evaluate a shared task dataset and compare it to the state-of-the-art system based on Stanford deep-coref .
Outcome: The proposed system achieves 1.13% gain over the CoNLL 2012 dataset and the state-of-the-art system.
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)

Copied to clipboard

Challenge: ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones.
Approach: They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions.
Outcome: The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format.
ParCorFull: a Parallel Corpus Annotated with Full Coreference (L18-1)

Copied to clipboard

Challenge: Recent research in multilingual coreference and automatic pronoun translation has led to important insights into the problem and some promising results.
Approach: They propose a corpus annotated with full coreference chains that addresses a problem that machine translation and other multilingual natural language processing (NLP) technologies face: translation of coreference across languages.
Outcome: The proposed corpus contains parallel texts for the language pair English-German, two major European languages.
An Application for Building a Polish Telephone Speech Corpus (L18-1)

Copied to clipboard

Challenge: Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance.
Approach: They propose to build a tool for speech corpus collection of a specific domain content.
Outcome: The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks.
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)

Copied to clipboard

Challenge: Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues.
Approach: They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms.
Outcome: The proposed corpus includes parallel text and speech data of 21 Japanese dialects.
Korean L2 Vocabulary Prediction: Can a Large Annotated Corpus be Used to Train Better Models for Predicting Unknown Words? (L18-1)

Copied to clipboard

Challenge: a recent study suggests that a classifier trained on unknown words may yield better results for L2 learners.
Approach: They propose to use a supervised learning classifier to predict word complexity in Korean . they propose to train models on annotated corpus of unknown words with 71 % precision .
Outcome: The proposed model recalls 80 % of unknown words with 71 % precision.
Crowdsourcing-based Annotation of the Accounting Registers of the Italian Comedy (L18-1)

Copied to clipboard

Challenge: CIRESFI project aims to reassess a theatrical heritage that has often been considered inferior to that of the two major, royally-privileged theaters.
Approach: They propose a double annotation system for new handwritten historical documents . crowdsourcing platform is set up to perform labeling and transcription of the documents based on budget data .
Outcome: The proposed system is based on a database of 25,250 pages of registers of the Italian Comedy of the 18th century.
FEIDEGGER: A Multi-modal Corpus of Fashion Images and Descriptions in German (L18-1)

Copied to clipboard

Challenge: Recent years have seen a renewed interest in text-image multi-modality . paired text-picture datasets are often limited to English language text .
Approach: They propose a multi-modal corpus that pairs images and textual descriptions of their content in German to enable study of these challenges.
Outcome: The proposed dataset focuses on the domain of fashion items and their visual descriptions in German.
Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Using a crowdsourcing platform, we collected 18,917 annotations for a less-resourced French regional language, Alsatian.
Approach: They developed a platform that allows people to gather part-of-speech annotations on a variety of corpora and train a first tagger specific to Alsatian.
Outcome: The proposed method is valid for Alsatian and can be adapted to other languages.
Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones.
Approach: They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences.
Outcome: The proposed set of simplified sentences is a good quality data set for machine learning.
A Multilingual Wikified Data Set of Educational Material (L18-1)

Copied to clipboard

Challenge: a crowdsourcing effort to annotate and link parallel texts has been unsuccessful . a data set of parallel texts in eleven languages is presented .
Approach: They present a wikified data set of English sentences linked to Wikipedia pages . they use crowdsourcing to annotate the texts and perform crowdsourcing for complex annotations .
Outcome: The proposed data set is valuable as it constitutes a rich resource . it includes annotated data of English sentences linked to translations in eleven languages .
Using Crowd Agreement for Wordnet Localization (L18-1)

Copied to clipboard

Challenge: Lexical-semantic resources like WordNet are a fundamental resource for many NLP and semantic applications.
Approach: They propose a crowdsourcing workflow that consists of synset localization and validation . they use inter-rater agreement metrics to estimate the precision of the results .
Outcome: The proposed method is cost-effective and provides a good trade-off between quality and speed of progress.
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)

Copied to clipboard

Challenge: a large corpus of online content has been developed via large-scale crowdsourcing.
Approach: They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform.
Outcome: The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines.
Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing (L18-1)

Copied to clipboard

Challenge: a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing .
Approach: They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners.
Outcome: a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy .
Chinese Relation Classification using Long Short Term Memory Networks (L18-1)

Copied to clipboard

Challenge: Relation classification is the task to predict semantic relations between pairs of entities in a given text.
Approach: They propose to extract relations between entities in Chinese text using a long-term memory network.
Outcome: The proposed system achieves state-of-the-art F-measure on ACE 2005 corpus . it predicts relations between head entity e h and tail entity t from sentence .
The UIR Uncertainty Corpus for Chinese: Annotating Chinese Microblog Corpus for Uncertainty Identification from Social Media (L18-1)

Copied to clipboard

Challenge: Uncertainty identification is an important semantic processing task, critical to the quality of information in terms of factuality in many NLP techniques and applications.
Approach: They propose to annotate Chinese microblogs with an open uncertainty corpus . they propose to use contextual uncertain semantics rather than traditional cue-phrases to identify uncertainty .
Outcome: The proposed corpus can be used to identify uncertainty in social media texts.
EventWiki: A Knowledge Base of Major Events (L18-1)

Copied to clipboard

Challenge: Existing knowledge bases focus on static entities such as people, locations and organizations.
Approach: They propose a new knowledge base resource called EventWiki which concentrates on major events . they show that EventWiki is a very useful resource for information extraction regarding events in NLP .
Outcome: The proposed resource is the first knowledge base resource of major events.
Annotating Spin in Biomedical Scientific Publications : the case of Random Controlled Trials (RCTs) (L18-1)

Copied to clipboard

Challenge: Fig. 6: Annotation of biomedical abstracts for automatic detection of inadequate claims (spin) spin is a misleading presentation of scientific results in randomized controlled trials, an important type of clinical trial.
Approach: They propose an algorithm for automatic detection of inadequate claims (spin) they propose to use a corpus of biomedical articles for the task .
Outcome: The proposed algorithm can detect inadequate claims in biomedical abstracts without requiring any prior knowledge of the literature.
Visualization of the occurrence trend of infectious diseases using Twitter (L18-1)

Copied to clipboard

Challenge: Existing methods for visualizing epidemics of infectious diseases are limited . a system to obtain how many people are infected by the given disease is created from tweets .
Approach: They propose a system for visualizing the epidemics of infectious diseases . they apply factuality analysis to disease detection and location estimation .
Outcome: The proposed method performs well on various diseases.
Reusable workflows for gender prediction (L18-1)

Copied to clipboard

Challenge: Existing systems for author profiling (AP) modeling require extensive feature engineering and testing.
Approach: They propose to implement a system for author profiling (AP) modeling that reduces the complexity and time of building a sophisticated model for a number of different AP tasks.
Outcome: The proposed model achieves comparable results to state of the art models for cross-genre gender prediction, but lags when genre of test set is different from genre of train set.
Knowing the Author by the Company His Words Keep (L18-1)

Copied to clipboard

Challenge: In traditional linguistics, there exists a famous saying that one should know a word by the company it keeps.
Approach: They propose a method which uses word embeddings to identify pairwise relational features in the context of authorship attribution.
Outcome: The proposed method is based on three literary corpora and shows that word similarity is a key feature in the authorship attribution task.
Towards a Gold Standard Corpus for Variable Detection and Linking in Social Science Publications (L18-1)

Copied to clipboard

Challenge: a new corpus for detecting and linking survey variables is being developed . the corpus is multilingual and includes manually curated word and phrase alignments .
Approach: They propose to create a corpus for the evaluation of detecting and linking survey variables in social science publications.
Outcome: The proposed corpus is the first gold standard for the variable detection and linking task.
KRAUTS: A German Temporally Annotated News Corpus (L18-1)

Copied to clipboard

Challenge: Temporal tagging is an important task towards improved natural language understanding.
Approach: They present a new German temporally annotated corpus with 192 documents with 1,140 annotations . they propose to make temporal tagging a viable research area .
Outcome: The proposed corpus contains 192 documents with 1,140 annotated temporal expressions.
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)

Copied to clipboard

Challenge: a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks.
Approach: They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community .
Outcome: The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges.
A Framework for the Needs of Different Types of Users in Multilingual Semantic Enrichment (L18-1)

Copied to clipboard

Challenge: FREME framework bridges Language Technologies (LT) and Linked Data (LD) core attributes of FREMe are usability, reusability and interoperability.
Approach: They define user types and user levels and describe how they influence design decisions in a LT and Linked Data processing framework.
Outcome: The proposed framework bridges Language Technologies (LT) and Linked Data (LD) it addresses common challenges that researchers and industry face when integrating LT and LD: interoperability, "silo" solutions and the lack of adequate tooling.
The LREC Workshops Map (L18-1)

Copied to clipboard

Challenge: a corpus of workshops titles and related presentations has been retrieved from the conference's website . data is used to analyze the research presented at the conferences over the years 1998-2016 .
Approach: a paper aims to present an overview of the research presented at the LREC workshops over the years 1998-2016.
Outcome: The aim of the present study is to shed light on the community represented by workshop participants over the years 1998-2016.
Preserving Workflow Reproducibility: The RePlay-DH Client as a Tool for Process Documentation (L18-1)

Copied to clipboard

Challenge: a tool for elicitation and management of process metadata is presented . detailed documentation of workflows is an arduous and neglected task .
Approach: They propose a software tool for elicitation and management of process metadata.
Outcome: The proposed tool minimizes the additional effort required for producing a sustainable workflow documentation.
The ACoLi CoNLL Libraries: Beyond Tab-Separated Values (L18-1)

Copied to clipboard

Challenge: a new set of Java archives facilitates advanced manipulations of corpora annotated in TSV formats.
Approach: They propose to use Java archives to facilitate advanced manipulations of corpora annotated in TSV formats.
Outcome: The proposed libraries support all members of the CoNLL format family.
What’s Wrong, Python? – A Visual Differ and Graph Library for NLP in Python (L18-1)

Copied to clipboard

Challenge: a library that allows the user to visualise and compare the output of a program with a well-known data format is needed.
Approach: They propose a supervised learning tool that allows users to visualise and compare program output . they use popular off-the-shelf visualisation programs to specify essential primitive functions .
Outcome: The proposed tool gives the user total control over visualisation and compares output of any program with a well-known data format.
ScholarGraph:a Chinese Knowledge Graph of Chinese Scholars (L18-1)

Copied to clipboard

Challenge: ScholarSpace integrates chinese academic information from chin scholars and science . data integration system needs to be focused on scholars, says dr. s. k. o. j. nielson .
Approach: a data integration system is built to integrate chinese academic information from chin scholars and science. a system can give you an academic portrait about a chinoise scholar with the form of a knowledge graph.
Outcome: a data integration system called ScholarSpace can integrate chinese academic information from chin scholars and science.
Enriching Frame Representations with Distributionally Induced Senses (L18-1)

Copied to clipboard

Challenge: lexical resource that enriches Framester knowledge graph with semantic features from text corpora . paves way for development of novel, deeper semantic-aware applications .
Approach: They propose a lexical resource that enriches the Framester knowledge graph with semantic features from text corpora.
Outcome: The proposed resource enables the development of deeper semantic-aware applications . it combines knowledge from text and symbolic representations of events and participants .
An Integrated Formal Representation for Terminological and Lexical Data included in Classification Schemes (L18-1)

Copied to clipboard

Challenge: e-lexicography is a field of study dealing with the automated creation of specialized multilingual dictionaries from structured data.
Approach: They propose to use a SKOS-XL vocabulary for modelling the multilingual terminological part of comparable taxonomies and OntoLex-Lemon for modelling multilingual lexical entries.
Outcome: The proposed model can be explicitly cross-linked in the context of the Linguistic Linked Open Data (LLOD).
One event, many representations. Mapping action concepts through visual features. (L18-1)

Copied to clipboard

Challenge: a proposed classification of general verbs is characterized by a high ambiguity and high frequency in the use.
Approach: They propose to use IMAGACT visual component as linkage point between resources . they propose automatic linking with BabelNet and manual linking with Praxicon .
Outcome: The proposed solution exploits the IMAGACT visual component as the linkage point among resources.
Tel(s)-Telle(s)-Signs: Highly Accurate Automatic Crosslingual Hypernym Discovery (L18-1)

Copied to clipboard

Challenge: a heuristic that exploits morphological cues in French to uniquely identify hypernyms is used in other languages . a recent study shows that this heuriistic is more informative than its English counterpart .
Approach: They propose a hypernym discovery heuristic that leverages morphological cues in French . they exploit morphology in the trigger phrase tel-1 que to uniquely identify the correct hypernaym .
Outcome: The proposed method exploits morphological cues in French to uniquely identify hypernyms . it can be used in other languages, and it is more accurate than its English counterpart .
Disambiguation of Verbal Shifters (L18-1)

Copied to clipboard

Challenge: Negation is a contextual phenomenon that needs to be addressed in sentiment analysis.
Approach: They propose a supervised learning approach to disambiguate verbal shifters using generalization features and a new lexicon.
Outcome: The proposed approach takes into account various features, particularly generalization features.
Bootstrapping Polar-Opposite Emotion Dimensions from Online Reviews (L18-1)

Copied to clipboard

Challenge: Existing bootstrapping methods for learning lexicons from unannotated online texts have important drawbacks.
Approach: They propose a bootstrapping approach that softly labels unlabeled terms for polar-opposite emotion dimension values from the Ortony/Clore/Collins model of emotions.
Outcome: The proposed approach achieves considerably better performance than several baselines.
Sentiment-Stance-Specificity (SSS) Dataset: Identifying Support-based Entailment among Opinions. (L18-1)

Copied to clipboard

Challenge: Argument mining is a method for extracting argument components and structures from natural language texts.
Approach: They propose to model arguments as a set of premises that either support each other or collectively support a conclusion.
Outcome: The proposed rules give an overall accuracy of 0.83 for the three datasets.
Resource Creation Towards Automated Sentiment Analysis in Telugu (a low resource language) and Integrating Multiple Domain Sources to Enhance Sentiment Prediction (L18-1)

Copied to clipboard

Challenge: Sentiment Analysis of text is an important task in many applications . but the task becomes challenging when it comes to low resource languages .
Approach: They propose to create a corpus of polarity-based sentiment classifiers in Telugu for different domains like movie reviews, song lyrics, product reviews and book reviews.
Outcome: The proposed model performs well in multiple domains and is compared with the previous models.
Multilingual Multi-class Sentiment Classification Using Convolutional Neural Networks (L18-1)

Copied to clipboard

Challenge: a new language-independent model for sentiment analysis is proposed for social media . a sentiment dictionary cannot list all the possible ways people can express their opinions .
Approach: They propose a language-independent model for multi-class sentiment analysis using a neural network architecture.
Outcome: The proposed model does not rely on language-specific features such as ontologies, dictionaries, or morphological or syntactic pre-processing.
A Large Self-Annotated Corpus for Sarcasm (L18-1)

Copied to clipboard

Challenge: Existing datasets for sarcasm detection have unbalanced and self-annotated labels, allowing for learning in both balanced and unbalanciated label regimes.
Approach: They introduce the Self-Annotated Reddit Corpus (SARC) which has 1.3 million sarcastic statements and many times more instances of non-sarcasm statements.
Outcome: The proposed corpus has 1.3 million sarcastic statements and many more instances of non-sarcasm statements, allowing for learning in both balanced and unbalanced label regimes.
HappyDB: A Corpus of 100,000 Crowdsourced Happy Moments (L18-1)

Copied to clipboard

Challenge: Recent research has focused on developing technologies that help users incorporate the findings of the science of happiness into their daily lives.
Approach: They crowd-sourced HappyDB, a corpus of 100,000 happy moments, and applied several state-of-the-art analysis techniques to analyze HappyDB.
Outcome: The proposed technology can understand how people express their happy moments in text and analyze them using state-of-the-art techniques.
MultiBooked: A Corpus of Basque and Catalan Hotel Reviews Annotated for Aspect-level Sentiment Classification (L18-1)

Copied to clipboard

Challenge: sentiment analysis research has focused on unsupervised or semi-supervised approaches, but these still require a large number of resources and do not reach the performance of supervised approaches.
Approach: They propose two datasets for supervised aspect-level sentiment analysis in Basque and Catalan.
Outcome: The proposed datasets are based on two under-resourced languages, basque and catalan.
BlogSet-BR: A Brazilian Portuguese Blog Corpus (L18-1)

Copied to clipboard

Challenge: Several efforts have been made to build a corpus based on user-generated content . however, there is still a lack of a large semi-structured corpus that also contains author profiles in Brazilian Portuguese.
Approach: They propose to build a Brazilian Portuguese corpus with 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs.
Outcome: The proposed corpus contains 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs.
SoMeWeTa: A Part-of-Speech Tagger for German Social Media and Web Texts (L18-1)

Copied to clipboard

Challenge: Off-the-shelf part-of-speech taggers perform poorly on web and social media data . this is due to the many unconventional spelling variants that occur in web and twitter texts and that result in a high proportion of out-of vocabulary words.
Approach: They propose to use TIGER corpus as a part-of-speech tagger to train a German part- of-speak tagger on the web and social media data of the EmpiriST 2015 shared task.
Outcome: The proposed tagger significantly improves on the state-of-the-art for both the web and social media data.
Collecting Code-Switched Data from Social Media (L18-1)

Copied to clipboard

Challenge: a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages .
Approach: They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets .
Outcome: The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets .
Classifying the Informative Behaviour of Emoji in Microblogs (L18-1)

Copied to clipboard

Challenge: Emoji are pictographs used in microblogs as emotion markers, but can also represent a wider range of concepts.
Approach: They analyze a corpus of tweets pairs and classify emoji with respect to redundancy . they propose to further investigate the informative behaviour of e-mails using eoji .
Outcome: The proposed model achieved an F-score of 0.7 for emoji use in 2475 tweets pairs.
A Taxonomy for In-depth Evaluation of Normalization for User Generated Content (L18-1)

Copied to clipboard

Challenge: Existing taxonomies for lexical normalization are not suitable for the task of normalization since the categories are substantially different.
Approach: They propose a taxonomy of error categories for lexical normalization . they annotate a recent normalization dataset and read a near-perfect agreement .
Outcome: The proposed taxonomy is based on a recent normalization dataset and it performs well.
Gaining and Losing Influence in Online Conversation (L18-1)

Copied to clipboard

Challenge: a study aimed to determine if people who are influential in online discussions retain influence when placed in a topic that is less familiar or perhaps not as interesting.
Approach: They conducted a study to determine if people who are highly influential retain influence when moving to a topic that is less familiar or perhaps not as interesting.
Outcome: The results show that people who are highly influential in group discussions lose influence when placed in a topic that is less familiar or perhaps not as interesting.
Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification (L18-1)

Copied to clipboard

Challenge: Existing corpus of Arabic textual data is limited to English or other European languages.
Approach: They present a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the arab world representing the major Arabic dialectal varieties.
Outcome: The provided corpus will enrich the limited set of available language resources for Arabic and be invaluable enabler for developing author profiling tools and NLP tools for Arabic.
Transc&Anno: A Graphical Tool for the Transcription and On-the-Fly Annotation of Handwritten Documents (L18-1)

Copied to clipboard

Challenge: Transc&Anno is a web-based collaboration tool for linguists to facilitate the transcription of text images and their shallow on-the-fly annotation.
Approach: They propose a web-based collaboration tool that allows the transcription of text images and their shallow on-the-fly annotation.
Outcome: The Transc&Anno tool can be used for any type of corpora requiring transcription and shallow on-the-fly annotation resulting in inline XML.
Correction of OCR Word Segmentation Errors in Articles from the ACL Collection through Neural Machine Translation Methods (L18-1)

Copied to clipboard

Challenge: Optical Character Recognition (OCR) can produce a range of errors depending on the quality of the original document.
Approach: They applied a sequence-to-sequence machine translation system to correct word-single-word OCR errors in scientific texts from the ACL collection with an estimated precision and recall above 0.95 on test data.
Outcome: The proposed system corrects word-segmentation OCR errors with an estimated precision and recall above 0.95 on test data.
From Manuscripts to Archetypes through Iterative Clustering (L18-1)

Copied to clipboard

Challenge: philologists have since the beginnings of the age of print attempted to provide one textual representation of this variety using stemmata . stemmats are trees depicting the copy history (manuscripts = nodes, Copy processes = edges)
Approach: They propose to use stemmata to extract the most likely common ancestor of all observed variants and then iteratively cluster them to create a single textual representation.
Outcome: The proposed method uses trees depicting the copy history to find the base text which is most likely the latest common ancestor of all observed variants.
Building A Handwritten Cuneiform Character Imageset (L18-1)

Copied to clipboard

Challenge: Currently, the digitization process is laborious due to the huge scale of the documents and no trustful (semi-)automatic method has been established.
Approach: They propose to build a handwritten cuneiform character imageset from handwritten documents to support OCR system.
Outcome: The proposed image processing method will support the development of handwritten cuneiform OCR system.
PDF-to-Text Reanalysis for Linguistic Data Mining (L18-1)

Copied to clipboard

Challenge: In the 1990s, extracting semistructured text from scientific writing in PDF files was largely a computer vision and OCR problem.
Approach: They propose a system for the reanalysis of PDF-extracted text that performs block detection, respacing, and tabular data analysis for linguistic data mining.
Outcome: The proposed system eliminates the extreme verbosity of XML output while leaving important positional information available for downstream processes.
Crowdsourced Multimodal Corpora Collection Tool (L18-1)

Copied to clipboard

Challenge: a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab.
Approach: They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion.
Outcome: The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus.
Expert Evaluation of a Spoken Dialogue System in a Clinical Operating Room (L18-1)

Copied to clipboard

Challenge: With the emergence of new technologies, the surgical working environment becomes increasingly complex and comprises many medical devices which have to be monitored and controlled.
Approach: They propose to use natural spoken language to control surgical operating rooms to reduce the amount of staff needed during a procedure.
Outcome: The proposed system can control the operating room using natural spoken language and is evaluated by experts in the field of minimally invasive surgery.
JAIST Annotated Corpus of Free Conversation (L18-1)

Copied to clipboard

Challenge: Annotated corpus of free conversations in Japanese is the first publicly available one.
Approach: They propose to annotate free conversations in Japanese with dialog act and sympathy tags . they report how to construct the corpus and its statistics .
Outcome: The proposed corpus is the first annotated corpus of free conversations in Japanese . it consists of 92,031 utterances in 97 dialogs.
The Metalogue Debate Trainee Corpus: Data Collection and Annotations (L18-1)

Copied to clipboard

Challenge: Argumentation is an important component of human intelligence and is used to train lawyers and citizens in legal domains.
Approach: They describe the Metalogue Debate Trainee Corpus (DTC) which contains data on motion and speech capture devices and semantic annotations.
Outcome: The metalogue Debate Trainee Corpus (DTC) was developed to facilitate the design of instructional and interactive models for the Virtual Debate Coach application.
Towards Continuous Dialogue Corpus Creation: writing to corpus and generating from it (L18-1)

Copied to clipboard

Challenge: Existing methods to create dialogue corpora annotated with interoperable semantic information are based on ISO standard data models and tools.
Approach: They propose to use a corpus as a shared repository for analysis and modelling of interactive dialogue behaviour and for implementation, integration and evaluation of dialogue system components.
Outcome: The proposed method is applied to the design of two multimodal interactive applications - the Virtual Negotiation Coach and the Virtual Debate Coach.
MYCanCor: A Video Corpus of spoken Malaysian Cantonese (L18-1)

Copied to clipboard

Challenge: The corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
Approach: the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
Outcome: the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
KTH Tangrams: A Dataset for Research on Alignment and Conceptual Pacts in Task-Oriented Dialogue (L18-1)

Copied to clipboard

Challenge: Existing studies on instructor-manipulator dialogue use disparate but similar datasets . a recent study examined the alignment of referring expressions (RL) in situated dialogue .
Approach: They propose to use a corpus of referring expressions in a relatively free dialogue with physical features generated in simulated situations to study alignment in referring language.
Outcome: The proposed datasets facilitate analysis of dialogic linguistic phenomena regarding alignment in the formation of referring expressions known as conceptual pacts.
On the Vector Representation of Utterances in Dialogue Context (L18-1)

Copied to clipboard

Challenge: In recent years, the representation of words as vectors in a vector space has gained a high degree of attention in the research community.
Approach: They introduce a new language resource that represents dialogue utterances in vector space and captures the semantic meaning of those utterrances in the dialogue context.
Outcome: The proposed model captures relevant semantic information by comparing them to manually annotated dialogue acts.
ES-Port: a Spontaneous Spoken Human-Human Technical Support Corpus for Dialogue Research in Spanish (L18-1)

Copied to clipboard

Challenge: ES-Port is a spontaneous spoken human-human dialogue corpus in Spanish that consists of 1170 dialogues from calls to the technical support department of a telecommunications provider.
Approach: They describe the compilation process from transcription to anonymisation of sensitive data contained in the transcriptions.
Outcome: The ES-Port corpus is a human-human dialogue corpus that consists of 1170 calls to the technical support department of a telecommunications provider.
From analysis to modeling of engagement as sequences of multimodal behaviors (L18-1)

Copied to clipboard

Challenge: Embodied Conversational Agents (ECAs) are virtual characters that can interact with a user.
Approach: They propose to endow an Embodied Conversational Agent with engagement capabilities . they use a corpus of expert-novice interactions to analyze user's engagement level .
Outcome: The proposed approach analyzes user's engagement level and controls agent's behavior.
A corpus of German political speeches from the 21st century (L18-1)

Copied to clipboard

Challenge: a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level .
Approach: a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level .
Outcome: The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts .
Building Literary Corpora for Computational Literary Analysis - A Prototype to Bridge the Gap between CL and DH (L18-1)

Copied to clipboard

Challenge: Literature analysis using corpus-based literary analysis is slow, says aaron s. e. . literary studies researchers should focus on the research practices of literary studies, he says .
Approach: et al. show litText can extract text from a 20 million word corpus using SPARQL queries.
Outcome: The proposed method uses a 20 million word corpus from English, German, Spanish, French and Italian texts and an example query to identify texts where animals behave like humans as it is the case in fables.
Towards faithfully visualizing global linguistic diversity (L18-1)

Copied to clipboard

Challenge: Currently, visualizations of worldwide linguistic diversity are limited by point symbology .
Approach: They propose a method to visualize linguistic diversity using point symbology instead of Mercator . instead of languages-as-points, they use Voronoi/Thiessen tessellations to model linguistic areas .
Outcome: The proposed method is based on an Eckert IV projection instead of Mercator . instead of languages-as-points, it uses Voronoi/Thiessen tessellations to model linguistic areas .
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
Identifying Speakers and Addressees in Dialogues Extracted from Literary Fiction (L18-1)

Copied to clipboard

Challenge: Using a sequence labeling approach, it is possible to identify speakers and addressees in dialogues extracted from literary fiction using a small amount of training data.
Approach: They propose to use a sequence labeling approach applied to a given set of characters to identify speakers and addressees in dialogues extracted from literary fiction.
Outcome: The proposed method allows for enriched search facilities and construction of social networks from the corpora.
Word Embedding Evaluation Datasets and Wikipedia Title Embedding for Chinese (L18-1)

Copied to clipboard

Challenge: Existing evaluation sets for word embeddings in English are limited.
Approach: They propose to translate existing evaluation sets from English to Chinese to evaluate Chinese word embeddings.
Outcome: The proposed evaluation sets are based on translations of popular evaluation sets from English to Chinese and human rating from Amazon Mechanical Turk workers.
An Automatic Learning of an Algerian Dialect Lexicon by using Multilingual Word Embeddings (L18-1)

Copied to clipboard

Challenge: a study on the Algerian Arabic dialect aims to build a lexicon of words written in Arabic or Latin script . multilinguality of the corpus is due to the fact that people use several languages to post comments . stretched letters, misspelled words, emoticons, condensed writing are among the problems .
Approach: They propose to build automatically from a social network an Algerian dialect lexicon.
Outcome: The proposed method leads to a score of 73% on a test lexicon . the study is based on analyzing a lexical corpus of an Algerian dialect .
Candidate Ranking for Maintenance of an Online Dictionary (L18-1)

Copied to clipboard

Challenge: lexicographers have traditionally identified a lexical item to add to a dictionary . but in the modern age of online dictionaries, queries for lexicals are indistinguishable from a larger list of misspellings . a system that uses machine learning techniques to assign "misspells" a probability of being a novel or missing entry is developed .
Approach: They develop a system that uses machine learning techniques to assign "misspells" a probability of being a novel or missing entry.
Outcome: The proposed system assigns "misspells" a probability of being a novel or missing entry . it uses signals from orthography, usage by trusted online sources, and dictionary query patterns .
Language adaptation experiments via cross-lingual embeddings for related languages (L18-1)

Copied to clipboard

Challenge: Language Adaptation is a general approach to extend existing resources from a better resourced language to a lesser resourced one.
Approach: They propose to exploit lexical and grammatical similarity between languages when they are related by using orthographic similarity.
Outcome: The proposed method improves the state of the art in induction of bilingual lexicons . it also improves induction performance in the Named-Entity Recognition task .
Tools for Building an Interlinked Synonym Lexicon Network (L18-1)

Copied to clipboard

Challenge: a new lexicon is being developed for cross-lingual (Czech and English) synonyms based on their syntactic and semantic behavior in (bilingual) context.
Approach: They propose to build a new interlinked verbal synonym lexicon called CzEngClass using a tool that helps to keep cross-lingual synonym classes consistent.
Outcome: The proposed lexicon captures cross-lingual (Czech and English) synonyms . the tool, called Synonym Class Editor -SynEd, is customized to build and edit entries .
Very Large-Scale Lexical Resources to Enhance Chinese and Japanese Machine Translation (L18-1)

Copied to clipboard

Challenge: A major issue in machine translation applications is the recognition and translation of named entities.
Approach: They propose to integrate Very Large-Scale Lexical Resources (VLSLR) with lexicons to improve machine translation accuracy.
Outcome: The proposed lexical resources can enhance the quality of MT in general and NMT systems, which currently don't use lexicons.
Combining Concepts and Their Translations from Structured Dictionaries of Uralic Minority Languages (L18-1)

Copied to clipboard

Challenge: a new method to expand the knowledge in existing dictionaries is proposed . small Uralic languages are facing a problem of limited language resources .
Approach: They propose to combine conceptually divided translations from multilingual dictionaries for small Uralic languages into a single lexical entry.
Outcome: The proposed method can be used to expand existing dictionaries and provide translations when adding new entries.
Transfer of Frames from English FrameNet to Construct Chinese FrameNet: A Bilingual Corpus-Based Approach (L18-1)

Copied to clipboard

Challenge: Current publicly available Chinese FrameNet has a relatively low coverage of frames and lexical units compared with other languages.
Approach: They propose an automatic way to construct Chinese FrameNet using a sentence-aligned English-Chinese bilingual corpus.
Outcome: The proposed resource can provide frame recommendations acceptable by annotators.
EFLLex: A Graded Lexical Resource for Learners of English as a Foreign Language (L18-1)

Copied to clipboard

Challenge: EFLLex describes the use of 15,280 English words in pedagogical materials across proficiency levels.
Approach: They propose to use a part-of-speech tagger and a robust estimator to compute frequency and do manual post-editing work to improve the resource.
Outcome: The proposed resource describes the use of 15,280 English words across proficiency levels of the European Framework of Reference for Languages.
English-Basque Statistical and Neural Machine Translation (L18-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) requires large training corpora, which is problematic for low-resource languages.
Approach: They propose to use an open-domain and an IT-domain corpora to train machine translations in English-Basque.
Outcome: The proposed systems outperform OpenNMT, Moses SMT and Google Translate in English-Basque translation.
TQ-AutoTest – An Automated Test Suite for (Machine) Translation Quality (L18-1)

Copied to clipboard

Challenge: Especially the trend towards neural MT has renewed peoples' interest in better and more analytical diagnostic methods for MT quality.
Approach: They propose a framework that supports a linguistic evaluation of machine translations using test suites.
Outcome: The proposed framework supports linguistic evaluation of (machine) translations using test suites.
Exploiting Pre-Ordering for Neural Machine Translation (L18-1)

Copied to clipboard

Challenge: Existing studies have shown that Neural Machine Translation suffers from the problems that some source words are mistakenly translated for multiple times .
Approach: They propose a pre-ordering approach to solve the under-translation problem by pre-ordnanced source sentences and position embedding to enhance monotone translation.
Outcome: The proposed method significantly improves translation quality by 2.43 BLEU points on Chinese-to-English translation.
Improving a Multi-Source Neural Machine Translation Model with Corpus Extension for Low-Resource Languages (L18-1)

Copied to clipboard

Challenge: In machine translation, we often try to collect resources to improve performance.
Approach: They propose to use synthetic methods to extend low-resource corpus to create target sentences using synthetic methods.
Outcome: The proposed method improves translation performance for low-resource language pairs.
Dynamic Oracle for Neural Machine Translation in Decoding Phase (L18-1)

Copied to clipboard

Challenge: Existing methods to improve NMT performance but there is a discrepancy between training and inference when decoding.
Approach: They propose to use Scheduled Sampling to reduce the discrepancy between training and inference in NMT when decoding to mitigate the discrépancy.
Outcome: The proposed methods improve translation quality over standard NMT system.
One Sentence One Model for Neural Machine Translation (L18-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) is a new state of the art that can produce better results than traditional statistical machine translation.
Approach: They propose a dynamic neural network which learns a general network as usual and fine-tunes it for each test sentence.
Outcome: The proposed method improves translation performance when similar sentences are available.
A Parallel Corpus of Arabic-Japanese News Articles (L18-1)

Copied to clipboard

Challenge: a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs.
Approach: They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus.
Outcome: The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it .
Examining the Tip of the Iceberg: A Data Set for Idiom Translation (L18-1)

Copied to clipboard

Challenge: Neural Machine Translation (NMT) has been widely used in recent years with significant improvements for many language pairs.
Approach: They propose to use a large-scale data set to evaluate idiom translation in GermanEnglish.
Outcome: The proposed dataset is used to perform preliminary NMT experiments on idiom translation in GermanEnglish.
Automatic Enrichment of Terminological Resources: the IATE RDF Example (L18-1)

Copied to clipboard

Challenge: a recent paper aims to automate the maintenance of terminological resources.
Approach: They propose automatic approaches to maintain and increase lexical coverage of knowledge bases by using machine translation and multilingual word sense disambiguation.
Outcome: The proposed approach outperforms the existing methods with random sentences in most languages .
A Comparative Study of Extremely Low-Resource Transliteration of the World’s Languages (L18-1)

Copied to clipboard

Challenge: a phrase-based MT system performs better than other methods for transliterating Bible names . combining data and training a single neural system yields significant gains .
Approach: They compare several machine translation methods for transliterating Bible names . they find a phrase-based MT system performs better than other methods .
Outcome: The phrase-based MT system performs better than other methods, the study finds . but, the single-language system outperforms the phrase-backed MT systems .
Translating Web Search Queries into Natural Language Questions (L18-1)

Copied to clipboard

Challenge: a new method to generate natural language questions from keyword-based queries is proposed . a synergy between query-to-question problem and standard machine translation (MT) model is found .
Approach: They propose a method to generate well-formed natural language questions from keyword-based queries.
Outcome: The proposed method is well-formed natural language question generated from keyword-based query.
Construction of a Japanese Word Similarity Dataset (L18-1)

Copied to clipboard

Challenge: evaluating distributed word representations in languages that do not have such resources is difficult . et al., 2015: distributed word represent a sparse vector indicating the word itself or the context of the word.
Approach: They constructed a Japanese word similarity dataset to evaluate distributed representations in Japanese.
Outcome: a Japanese word similarity dataset is the first resource that can be used to evaluate distributed representations in Japanese . the dataset contains various parts of speech and includes rare words in addition to common words .
Acquiring Verb Classes Through Bottom-Up Semantic Verb Clustering (L18-1)

Copied to clipboard

Challenge: Existing methods for creating verbal classifications are limited or non-existent in most languages . a range of automatic verb classification approaches have been proposed, but high-quality resources are needed .
Approach: They propose to use top-up semantic clustering to extract syntactic and semantic information from verbs in English, Polish and Croatian.
Outcome: The proposed classifications in English, Polish and Croatian are compared with other languages.
Constructing High Quality Sense-specific Corpus and Word Embedding via Unsupervised Elimination of Pseudo Multi-sense (L18-1)

Copied to clipboard

Challenge: Existing word embedding frameworks distinguish different senses of words by their contexts.
Approach: They propose a framework for unsupervised corpus sense tagging which trains multi-sense word embeddings on a given corpus.
Outcome: The proposed framework detects pseudo multi-senses without extra language resources without additional language resources.
Urdu Word Embeddings (L18-1)

Copied to clipboard

Challenge: Recent advances in distributional semantics have led to the rise of neural network-based models that use unsupervised learning to represent words as dense, distributed vectors, called 'word embeddings' embedders hold key to improving natural language processing for low-resource languages, since they require significant time and manpower.
Approach: They train a skip-gram model on 140 million Urdu words to create the first large-scale word embeddings for the Urdu language.
Outcome: The proposed models capture high degree of syntactic and semantic similarity between words and are able to generalize well on the Urdu translation task.
Social Image Tags as a Source of Word Embeddings: A Task-oriented Evaluation (L18-1)

Copied to clipboard

Challenge: Distributional hypothesis-based word representations lack perceptual and empirical knowledge.
Approach: They evaluate the effectiveness of social image tags in generating word embeddings . they find that generated word embeds exhibit somewhat different behaviors from corpus-originated representations - authors .
Outcome: The generated word embeddings exhibit comparable performance with corpus-originated representations.
Towards AMR-BR: A SemBank for Brazilian Portuguese Language (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area.
Approach: They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach.
Outcome: The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues.
Towards a Welsh Semantic Annotation System (L18-1)

Copied to clipboard

Challenge: Automatic semantic annotation of natural language data is an important task in Natural Language Processing.
Approach: They develop a Welsh semantic annotation tool that can be used to analyze Welsh text . it uses Lancaster's USAS semantic classification scheme to tag words with semantic tags .
Outcome: The proposed tool can cover up to 91.78% of words in Welsh text.
Semantic Frame Parsing for Information Extraction : the CALOR corpus (L18-1)

Copied to clipboard

Challenge: a recent study compares the semantic parsing of encyclopedic history texts with the Berkeley FrameNet project.
Approach: They propose to use Berkeley FrameNet to parse encyclopedic history texts . they use a sequence labeling model which optimizes frame identification and role segmentation .
Outcome: The proposed approach leverages the manual annotation of larger corpora than full text parsing.
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)

Copied to clipboard

Challenge: specialized corpora on a variety of topics are available online, but online corporates are scarce.
Approach: They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists .
Outcome: The proposed corpus contains more than six million speeches in English and Chinese and is available for free online.
A Multi- versus a Single-classifier Approach for the Identification of Modality in the Portuguese Language (L18-1)

Copied to clipboard

Challenge: Comparative study of two different approaches to build an automatic classification system for Modality values in the Portuguese language.
Approach: They propose to use a single multi-class classifier with the full Portuguese language dataset that includes eleven modal verbs and a weighted average approach to build different classifiers for each verb.
Outcome: The proposed system is based on a Portuguese language dataset with 11 modal verbs and two different classifiers, one for each verb.
All-words Word Sense Disambiguation Using Concept Embeddings (L18-1)

Copied to clipboard

Challenge: Existing work on all-words word sense disambiguation (all-word WSD) uses word embeddings to identify the senses of words in documents.
Approach: They propose a new concept embedding method to predict target word senses . concept embeds are constructed from concept tag sequences created from previous predictions .
Outcome: The proposed concept embeddings improve Japanese all-words word sense disambiguation task.
Enhancing Modern Supervised Word Sense Disambiguation Models by Semantic Lexical Resources (L18-1)

Copied to clipboard

Challenge: Existing supervised models for Word Sense Disambiguation (WSD) are limited to knowledge-based approaches.
Approach: They propose to use WordNet and WordNet Domains to enhance supervised WSD models by introducing semantic features into the classifiers and using the SLR structure to augment training data.
Outcome: The proposed model improves the state-of-the-art in Word Sense Disambiguation (WSD) The proposed approach is compared with the state of the art in the most popular benchmarks.
An Unsupervised Word Sense Disambiguation System for Under-Resourced Languages (L18-1)

Copied to clipboard

Challenge: Existing systems for word sense disambiguation are limited to the Russian language and lack of resources to address the problem.
Approach: They propose an unsupervised system for word sense disambiguation that uses a traditional vector space model to estimate the most similar word sense corresponding to its context.
Outcome: The proposed system outperforms the sparse mode on all datasets according to the adjusted Rand index.
Unsupervised Korean Word Sense Disambiguation using CoreNet (L18-1)

Copied to clipboard

Challenge: Unsupervised learning based Korean word sense disambiguation is needed to distinguish between sense candidates.
Approach: They investigated unsupervised Korean word sense disambiguation using CoreNet, a Korean lexical semantic network.
Outcome: The proposed method exhibited an 80.9% accuracy on the datasets constructed and proved to be effective for practical applications.
UFSAC: Unification of Sense Annotated Corpora and Tools (L18-1)

Copied to clipboard

Challenge: a dozen sense annotated English corpora are used in Word Sense Disambiguation (WSD) a new format of corpus is proposed that can be used for training or testing a disambiguation system .
Approach: They propose a format of corpus that can be used for training or testing a disambiguation system . they provide the source code and a complete Java API for manipulating corpora in this format .
Outcome: The proposed format of corpus can be used for training or testing a disambiguation system . the source code and a complete Java API are provided for building the corpus .
Retrofitting Word Representations for Unsupervised Sense Aware Word Similarities (L18-1)

Copied to clipboard

Challenge: Standard word embeddings lack the ability to distinguish senses of a word by projecting them to exactly one vector.
Approach: They propose to retrofit standard word embeddings to produce sense-aware embeddable vectors using external resources as sense inventories.
Outcome: The proposed method improves word similarity and relatedness scores on multiple word embeddings and established word similarities, sometimes up to an impressive margin of +0.15 Spearman correlation score.
FastSense: An Efficient Word Sense Disambiguation Classifier (L18-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a task that is often overlooked by NLP pipelines because of its complexity and complexity.
Approach: They propose a neural network-based tool for word sense disambiguation called fastSense.
Outcome: The proposed tool can process huge amounts of data quickly and surpasses state-of-the-art tools in terms of F-measure.
Text Annotation Graphs: Annotating Complex Natural Language Phenomena (L18-1)

Copied to clipboard

Challenge: Text Annotation Graphs is a web-based tool for annotating text . it provides functionality for representing complex relationships between words and word phrases .
Approach: They introduce a web-based tool for annotating text, Text Annotation Graphs, or TAG . it provides functionality for representing complex relationships between words and word phrases .
Outcome: The proposed software can represent complex relationships between words and words . it can also be used to find similar structures within the current document or external annotated documents.
Manzanilla: An Image Annotation Tool for TKB Building (L18-1)

Copied to clipboard

Challenge: Currently, images are stored in the TKB as a whole and are only linked to the concept itself.
Approach: They propose to use images as representations of concepts in EcoLexicon to improve image reusability and consistency.
Outcome: The tool was created to enhance the consistency of knowledge representation through images with the conceptual knowledge in EcoLexicon and to improve image reusability.
Tools for The Production of Analogical Grids and a Resource of N-gram Analogical Grids in 11 Languages (L18-1)

Copied to clipboard

Challenge: a Python module implements several previously presented algorithms to build analogical grids from words contained in a corpus.
Approach: They propose to release a Python module which implements several previously presented algorithms to build analogical grids from words contained in a corpus.
Outcome: The tools were built on vocabularies contained in 1,000 lines of the 11 different language versions of the Europarl corpus v.3 and are language-independent, allowing their use with any language and any writing system.
The Automatic Annotation of the Semiotic Type of Hand Gestures in Obama’ s Humorous Speeches (L18-1)

Copied to clipboard

Challenge: Existing studies on hand gestures from video-recorded speeches have not identified them.
Approach: They annotated and analysed hand gestures produced by Barack Obama . they trained machine learning algorithms to classify the semiotic type of hand gesture .
Outcome: The proposed method can be used to classify hand gestures on video-recorded speeches and in advanced multimodal interactive systems.
WASA: A Web Application for Sequence Annotation (L18-1)

Copied to clipboard

Challenge: a major barrier to research on CS has been the lack of large multilingual, multi-genre CS-annotated corpora.
Approach: They propose a web-based annotation system that manages large-scale CS data annotation.
Outcome: The proposed system can manage large-scale multilingual code switching (CS) data annotation.
Annotation and Quantitative Analysis of Speaker Information in Novel Conversation Sentences in Japanese (L18-1)

Copied to clipboard

Challenge: a qualitative lexicological analysis of conversation sentences in novels is performed . gender and age of conversation sentence information is not used as actual speech .
Approach: They performed a quantitative lexicological analysis using attributed speaker information . they also examined the differences between Japanese novels and translations of foreign novels .
Outcome: The results show that conversation sentences in novels are representative of spoken language . the authors conclude that conversation sentence data are not useful as speech materials .
PDFAnno: a Web-based Linguistic Annotation Tool for PDF Documents (L18-1)

Copied to clipboard

Challenge: Currently, linguistic annotation tools for PDF documents focus on plain-text documents.
Approach: They propose a web-based linguistic annotation tool for PDF documents . it offers functions for various types of linguistic annotations directly on PDF .
Outcome: The proposed tool can annotate on PDF documents with named entity, dependency relation, and coreference chain.
A Lightweight Modeling Middleware for Corpus Processing (L18-1)

Copied to clipboard

Challenge: Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources.
Approach: They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools.
Outcome: The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface.
An Annotation Language for Semantic Search of Legal Sources (L18-1)

Copied to clipboard

Challenge: formalizing legal sources is an important challenge, but the generation of a formal representation from legal texts has been less considered and requires considerable expertise.
Approach: They propose to experiment with annotations and the annotation process to improve uniformity and efficiency of legal annotation.
Outcome: The proposed method improves the richness and efficiency of legal annotations.
Resource Interoperability for Sustainable Benchmarking: The Case of Events (L18-1)

Copied to clipboard

Challenge: Despite efforts to improve interoperability, there are still problems with benchmark corpora that are hampered by too laborious conversion steps.
Approach: They assess aspects of interoperability at the document-level across 20 annotated corpora and compare their compatibility and consistency across the corpors.
Outcome: The proposed framework enables the analysis of document intersections between the corpora and shows their compatibility and consistency across the corpus.
Parsivar: A Language Processing Toolkit for Persian (L18-1)

Copied to clipboard

Challenge: a preprocessing step is required to convert text into a standard format for NLP tasks.
Approach: They propose a Persian preprocessing toolkit that performs various kinds of activities . they use a plagiarism detection system to exploit the proposed toolkit .
Outcome: The proposed tool outperforms available Persian preprocessing tools by about 8 percent in terms of F1 . the proposed toolkit performs normalization, space correction, tokenization, stemming, parts of speech tagging and shallow parsing tasks.
Multilingual Word Segmentation: Training Many Language-Specific Tokenizers Smoothly Thanks to the Universal Dependencies Corpus (L18-1)

Copied to clipboard

Challenge: Towards language scalability, major progress has been achieved in multilingual language technology in recent years.
Approach: They propose a tokenizer that can be trained from any Universal Dependencies corpus dataset . they argue that tokenization should be seen as a supervised task and scalability requires a software engineering process across languages.
Outcome: The proposed tokenizer can be trained from any dataset in the corpus UD2 . the proposed software tool relies on elephant to perform the training .
Build Fast and Accurate Lemmatization for Arabic (L18-1)

Copied to clipboard

Challenge: Lemmatization is the process of finding the base form (lemma) of a word by considering its inflected forms.
Approach: They propose a lemmatizer for Arabic with a dataset that can be used to test lemma accuracy.
Outcome: The proposed algorithm outperforms state-of-the-art Arabic lemmatization in accuracy and speed.
JESC: Japanese-English Subtitle Corpus (L18-1)

Copied to clipboard

Challenge: Existing data on Japanese-English subtitles are limited due to the high cost of manual construction.
Approach: They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web.
Outcome: The JESC dataset covers the underrepresented domain of conversational dialogue.
Building a Corpus for Personality-dependent Natural Language Understanding and Generation (L18-1)

Copied to clipboard

Challenge: The computational treatment of human personality is central to the development of NLP applications.
Approach: They propose to use the b5 corpus to generate controlled and free (non-topic specific) texts . preliminary results of personality recognition from text are presented .
Outcome: The proposed corpus is the largest resource of this kind to be made available for research purposes in the Brazilian Portuguese language.
Linguistic and Sociolinguistic Annotation of 17th Century Dutch Letters (L18-1)

Copied to clipboard

Challenge: In the late 16th and 17th century, the Dutch language was subject to a series of standardization and modernization developments.
Approach: They propose to annotate a 17th century Dutch letter corpus manually with parts-of-speech, document segmentation and sociolinguistic metadata.
Outcome: The proposed corpus is annotated with parts-of-speech, document segmentation and sociolinguistic metadata.
Simplified Corpus with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a study has found that simple Japanese is more accessible to foreigners than English.
Approach: They have constructed a simplified corpus for the Japanese language and selected the core vocabulary.
Outcome: The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa.
A Pragmatic Approach for Classical Chinese Word Segmentation (L18-1)

Copied to clipboard

Challenge: Classical Chinese word segmentation is largely neglected due to its obsoleteness . a new approach to segmentation using a marked-up corpus is needed .
Approach: They propose a pragmatic approach to deal with Classical Chinese word segmentation without any marked-up corpus.
Outcome: The proposed method makes the CCWS without any marked-up corpus more accurate compared with collocation-based methods.
ASAP++: Enriching the ASAP Automated Essay Grading Dataset with Essay Attribute Scores (L18-1)

Copied to clipboard

Challenge: Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP).
Approach: They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays.
Outcome: The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc.
MirasText: An Automatically Generated Text Corpus for Persian (L18-1)

Copied to clipboard

Challenge: Natural language processing is one of the most important fields of artificial intelligence.
Approach: They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens .
Outcome: The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles .
The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)

Copied to clipboard

Challenge: a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania .
Approach: a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results .
Outcome: a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages .
A Corpus of Drug Usage Guidelines Annotated with Type of Advice (L18-1)

Copied to clipboard

Challenge: Current research indicates patients are often unaware of such critical information / advice related to their prescription drugs due to lack of communication with their doctors and/or pharmacists.
Approach: They propose an annotation scheme for annotating safety critical advice from drug usage guidelines and an annotated dataset containing drug usage guideline data.
Outcome: The proposed dataset will accelerate further release of annotated drug usage guideline datasets and research on automatically filtering safety critical information from these documents.
BioRo: The Biomedical Corpus for the Romanian Language (L18-1)

Copied to clipboard

Challenge: Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward.
Approach: They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining.
Outcome: The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language .
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)

Copied to clipboard

Challenge: a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes .
Approach: They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones .
Outcome: The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions.
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)

Copied to clipboard

Challenge: a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text .
Approach: They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags.
Outcome: The proposed method detects humor in code-mixed tweets in English-Hindi.
Dialogue Scenario Collection of Persuasive Dialogue with Emotional Expressions via Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Existing methods for data collection and annotation are costly and prevent launching new dialogue systems.
Approach: They asked crowd workers to create persuasive dialogue systems using emotional expressions . they annotated emotional states and users' acceptance for system persuasion .
Outcome: The proposed system has sufficient agreement even without training, the researchers found . the experiment showed that the collected data are comparable to real-world dialogue recording methods .
SentiArabic: A Sentiment Analyzer for Standard Arabic (L18-1)

Copied to clipboard

Challenge: Sentiment analysis is a process of applying computational approaches to identify attitudes, emotions and opinions in text, speech and visual data.
Approach: They propose a sentiment analyzer that identifies the overall contextual polarity for Arabic text.
Outcome: The proposed system achieves an F-score of 76.5% when evaluated on a blind test set.
Contextual Dependencies in Time-Continuous Multidimensional Affect Recognition (L18-1)

Copied to clipboard

Challenge: despite of the research done in this area there is still no agreement on this issue.
Approach: a paper compares the amount of context used in a model and performance of a time-continuous labelled spontaneous interaction.
Outcome: a new study shows that the amount of context used in a model and performance is similar across models . the results show that knowledge about an appropriate context can reduce complexity and flexibility .
WikiArt Emotions: An Annotated Dataset of Emotions Evoked by Art (L18-1)

Copied to clipboard

Challenge: a dataset of 4,000 pieces of art has annotations for emotions evoked in the observer . the dataset can help answer questions about what makes art evocative, how does art convey different emotions, what attributes of a painting make it well liked, and how much does the title impact the affectual response to art.
Approach: They create a dataset of 4,000 western art pieces that has annotations for emotions . they use crowdsourcing to annotate the art for one or more of twenty emotion categories . fear, happiness, love, sadness were the dominant emotions that obtained consistent annotations .
Outcome: The dataset shows that the most popular emotions are fear, happiness, love and sadness . the dataset can be used to develop systems that detect emotions evoked by art .
Arabic Data Science Toolkit: An API for Arabic Language Feature Extraction (L18-1)

Copied to clipboard

Challenge: Data scientists unfamiliar with Arabic or natural language processing prefer statistical methods because they are language-independent.
Approach: They propose a framework for Arabic feature extraction that leverages Arabic-specific linguistic and stylistic features to enhance their systems.
Outcome: The Arabic Data Science Toolkit (ADST) is a framework for Arabic language feature extraction.
Sentence and Clause Level Emotion Annotation, Detection, and Classification in a Multi-Genre Corpus (L18-1)

Copied to clipboard

Challenge: Existing methods for predicting emotion categories are limited due to their multi-label nature . e.g. anger, joy, sadness are difficult to predict due to inherent multi-genre nature - a problem that is often overlooked in single-genrete text.
Approach: They propose to expand existing annotated data to include 8 emotions from Plutchik's Wheel of Emotions . they explore the effectiveness of clause annotation in sentence-level emotion detection and classification .
Outcome: The proposed system is the first to target the clause level and provides emotion classification for movie reviews datasets.
A Swedish Cookie-Theft Corpus (L18-1)

Copied to clipboard

Challenge: Language disturbances can be a diagnostic marker for neurodegenerative diseases, such as Alzheimer's disease, at earlier stages.
Approach: They develop a corpus of audio recordings of the Cookie-theft, a standardized test that has been used in studies in the past.
Outcome: The proposed corpus is based on audio recordings of the Cookie-theft . it provides a rich resource for future research and experimentation in many areas .
Sharing Copies of Synthetic Clinical Corpora without Physical Distribution — A Case Study to Get Around IPRs and Privacy Constraints Featuring the German JSYNCC Corpus (L18-1)

Copied to clipboard

Challenge: eu legal culture imposes unsurmountable hurdles to exploit copyright protected language data . legal constraints have seriously hampered progress in resource-greedy NLP research . authors propose a new approach for the creation and re-use of clinical corpora .
Approach: They propose a method for the creation and re-use of clinical corpora based on a two-step workflow . they substitute authentic clinical documents by synthetic ones, i.e., made-up reports and case studies .
Outcome: a new approach replaces authentic clinical documents by synthetic ones, i.e., made-up reports and case studies published in medical e-textbooks.
A Legal Perspective on Training Models for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: a significant concern in processing natural language data is the unclear legal status of the input and output data/resources.
Approach: They examine which legal rules apply at relevant steps and how they affect the legal status of the results.
Outcome: The proposed model training process is based on three scenarios . the analysis focuses on which legal rules apply and how they affect the legal status of the results .
LREMap, a Song of Resources and Evaluation (L18-1)

Copied to clipboard

Challenge: a new map of Language Resources was introduced at LREC 2010 . the map aims to shed light on the vast amount of resources that represent the background of the research presented at LRE .
Approach: They revisit the LRE Map of Language Resources, introduced at LREC 2010 . the map was designed to shed light on the vast amount of resources that represent the background of the research presented at LRE .
Outcome: The paper analyzes the LRE Map of Language Resources, introduced at LREC 2010, from many different perspectives.
Metadata Collection Records for Language Resources (L18-1)

Copied to clipboard

Challenge: a pilot project aimed at bringing metadata records to the CLARIN context has been conducted . a virtual language observatory (VLO) was developed to provide an entry point to the language resources available in the infrastructure.
Approach: They propose to implement a CMDI profile for Dutch language resources . they propose an interface for creating, editing, listing, copying and exporting metadata records .
Outcome: The proposed interface is validated in a pilot with 45 Dutch language resources . the proposed interface provides a user interface for creating, editing, listing, copying and exporting descriptions of metadata collection records.
Managing Public Sector Data for Multilingual Applications Development (L18-1)

Copied to clipboard

Challenge: eTranslation is a digital service that enables multilingual communication across public administrations in 30 European countries.
Approach: They propose to develop a repository infrastructure specifically tailored to the needs of the eTranslation service of the European Commission.
Outcome: The ELRC-SHARE repository is designed and developed specifically for the eTranslation service of the European Commission.
Bridging the LAPPS Grid and CLARIN (L18-1)

Copied to clipboard

Challenge: The LAPPS-CLARIN project is creating a "trust network" between the Language Applications Grid and WebLicht workflow engine . the goal is to allow users on one side of the bridge to gain appropriately authenticated access to the other .
Approach: The LAPPS-CLARIN project is creating a "trust network" between the Language Applications Grid and WebLicht workflow engine hosted by the CLARIN-D Center in Tübingen.
Outcome: The LAPPS-CLARIN project is creating a "trust network" between the Language Applications (LAPPS) Grid and the WebLicht workflow engine hosted by the CLARIN-D Center in Tübingen.
Fluid Annotation: A Granularity-aware Annotation Tool for Chinese Word Fluidity (L18-1)

Copied to clipboard

Challenge: Using word segmentation, we propose a wordhood annotation framework for Chinese language . word segmentations have been used for years in preprocessing NLP tasks for languages without explicit word delimiter.
Approach: They propose a word-granularity-aware annotation framework for Chinese language . they argue that word segmentation is fluid in nature and that it rearranges the boundary of word segmentations and linguistic annotation.
Outcome: The proposed framework rearranges the boundary between word segmentation and linguistic annotation and supports flexible annotation tasks for various linguistic and affective phenomena.
E-magyar – A Digital Language Processing System (L18-1)

Copied to clipboard

Challenge: e-magyar is a free, open, modular text processing pipeline for Hungarian . existing tools were overhauled to operate in the pipeline with a uniform encoding and run in the same Java platform.
Approach: e-magyar is a free, open, modular text processing pipeline for Hungarian . it was created by a collaborative effort by the language technology community . the system is aimed at a broad range of users, from language developers to researchers .
Outcome: The proposed tool is open source and available for download on the HFST framework.
ILCM - A Virtual Research Infrastructure for Large-Scale Qualitative Data (L18-1)

Copied to clipboard

Challenge: iLCM project develops integrated research environment for qualitative data analysis . text mining and text mining tools are extended by "Open Research Computing"
Approach: iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture.
Outcome: iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture.
CLARIN’s Key Resource Families (L18-1)

Copied to clipboard

Challenge: CLARIN is a European Research Infrastructure that supports the accessibility of language resources and technologies to researchers from the Digital Humanities and Social Sciences.
Approach: They propose to present key resource families in a uniform way for researchers to use in their research using the CLARIN infrastructure.
Outcome: The key resource families are newspaper, parliamentary, CMC (computer-mediated communication), and parallel corpora.
Indra: A Word Embedding and Semantic Relatedness Server (L18-1)

Copied to clipboard

Challenge: Word embedding/distributional semantic models are a fundamental component in many natural language processing (NLP) architectures.
Approach: They propose a multi-lingual word embedding/distributional semantics framework which supports creation, use and evaluation of word embedded models.
Outcome: The proposed tool supports the creation, use and evaluation of word embedding models.
A UIMA Database Interface for Managing NLP-related Text Annotations (L18-1)

Copied to clipboard

Challenge: despite the use of UIMA as a document-based schema, it does not provide native database support.
Approach: They develop a database interface to allow generic use of UIMA documents in database systems.
Outcome: The framework is evaluated in relation to file system-based storage and provides data protection.
European Language Resource Coordination: Collecting Language Resources for Public Sector Multilingual Information Management (L18-1)

Copied to clipboard

Challenge: European Language Resource Coordination (ELRC) initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries.
Approach: They propose to initiate actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries.
Outcome: The European Language Resource Coordination (ELRC) consortium initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries.
Tilde MT Platform for Developing Client Specific MT Solutions (L18-1)

Copied to clipboard

Challenge: a growing demand for translations and multilingual content is surpassing the supply of professional translation services.
Approach: They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality.
Outcome: The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities.
Improving homograph disambiguation with supervised machine learning (L18-1)

Copied to clipboard

Challenge: a new system for text-to-speech synthesis uses rule-based homograph disambiguation . a simple application of machine learning produces significant improvements in homograph ambiguity .
Approach: They propose a rule-based homograph disambiguation system for text-to-speech synthesis at Google . they compare it to a new system which performs disambiguations using classifiers trained on labeled data .
Outcome: The proposed system is more accurate than hand-written rules or machine learning alone.
Text Normalization Infrastructure that Scales to Hundreds of Language Varieties (L18-1)

Copied to clipboard

Challenge: a multi-language text normalization infrastructure is used to train language models for keyboards and speech recognition systems.
Approach: They describe a multi-language text normalization infrastructure that prepares textual data to train language models used in Google's keyboards and speech recognition systems.
Outcome: The proposed system can normalize training data across hundreds of languages . it can detect errors in training data and detect corruption issues .
DeModify: A Dataset for Analyzing Contextual Constraints on Modifier Deletion (L18-1)

Copied to clipboard

Challenge: a text fragment is discarded when it has a smaller context, causing it to acquire a new meaning or even become false.
Approach: They build a dataset to study the effect of modifiers on the larger context . they focus on single-word modifiers, the smallest unit that can be considered disposable .
Outcome: The proposed dataset aims to determine whether modifiers can be removed without undesirable consequences.
Open Subtitles Paraphrase Corpus for Six Languages (L18-1)

Copied to clipboard

Challenge: Opusparcus is a new corpus of paraphrases for six European languages . it is based on movie and TV subtitles, which are colloquial and informal .
Approach: They propose to use opensubtitles2016 paraphrase corpus for six European languages . they extract paraphrases from movie and TV subtitles from the corpus .
Outcome: The new corpus is available in German, English, Finnish, French, Russian, and Swedish . it is extracted from the OpenSubtitles2016 corpus, which contains subtitles from movies and TV shows .
Fine-grained Semantic Textual Similarity for Serbian (L18-1)

Copied to clipboard

Challenge: Semantic textual similarity (STS) is a task of assigning a numerical score to short texts based on the level of semantic equivalence between them.
Approach: They propose to annotate Serbian STS dataset with fine-grained similarity scores . they propose a supervised bag-of-words model that combines part-of speech weighting with term frequency weighting .
Outcome: The proposed model outperforms existing models on the Serbian STS News Corpus . the proposed model is based on a new morphologically rich language .
SPADE: Evaluation Dataset for Monolingual Phrase Alignment (L18-1)

Copied to clipboard

Challenge: Existing studies on sentential paraphrase detection focus on finer grained paraphrases, i.e., phrasal paraphrase.
Approach: They propose to use the SPADE to evaluate syntactic phrase alignment in paraphrasal sentences.
Outcome: The proposed method is compared with humans and provides benchmarks to show its performance.
ETPC - A Paraphrase Identification Corpus Annotated with Extended Paraphrase Typology and Negation (L18-1)

Copied to clipboard

Challenge: Extended Paraphrase Typology addresses limitations of existing typologies . extended typology provides better means for evaluation and error analysis .
Approach: a new typology copes with non-paraphrase pairs in paraphrase identification corpora, a paper proposes . a large corpus annotated with atomic paraphrase types is the largest to date .
Outcome: The Extended Paraphrase Typology (EPT) and the Extended Typology Paraphrase Corpus (ETPC) address practical limitations of existing paraphrase typologies.
Introducing a Lexicon of Verbal Polarity Shifters for English (L18-1)

Copied to clipboard

Challenge: Negation words can change the sentiment polarity of a phrase, but there are more than 1200 other polarities.
Approach: They propose a lexicon of verbal polarity shifters that covers the entirety of verbs found in WordNet.
Outcome: The proposed lexicon covers the entirety of verbs found in WordNet.
JFCKB: Japanese Feature Change Knowledge Base (L18-1)

Copied to clipboard

Challenge: constructing commonsense knowledge including connotational meanings is challenging . a recent study focused on denotation and connotations, but few studies focused on connotating meanings .
Approach: They propose to construct a Japanese knowledge base where arguments in event sentences are associated with feature changes caused by events.
Outcome: The proposed knowledge base is able to generate anaphora resolution tasks in Japanese . it is useful for computers to understand texts, but it is difficult to acquire it .
Quantifying Qualitative Data for Understanding Controversial Issues (L18-1)

Copied to clipboard

Challenge: 'Controversy' is a state of sustained public debate on a topic or issue that evokes conflicting opinions, beliefs, claims, arguments, and points of view.
Approach: They propose a crowdsourced approach to quantifying qualitative information on controversial issues by analyzing crowdsourced assertions in social media.
Outcome: The proposed dataset consists of over 2,000 assertions on 16 controversial issues.
Distribution of Emotional Reactions to News Articles in Twitter (L18-1)

Copied to clipboard

Challenge: Social networks have created datasets of opinions of users that focus on the writers' perspective, which does not consider the source that provokes those opinions.
Approach: They propose to analyze opinions of Twitter users' after reading a news article and use it to predict the distribution of emotions.
Outcome: The proposed dataset aims to explore how the six emotions are expressed by Twitter users' after reading a news article.
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)

Copied to clipboard

Challenge: a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people.
Approach: They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook.
Outcome: The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.
Creating a Verb Synonym Lexicon Based on a Parallel Corpus (L18-1)

Copied to clipboard

Challenge: a new lexical resource called CzEngClass is being built to help define synonyms in a bilingual context.
Approach: They propose to group verb senses into bilingual verbal synonym groups and use a parallel dependency corpus to explore semantic 'equivalence' they argue that existence of core argument mappings and adjunct mappings to a common set of semantic roles is a suitable criterion for a reasonable verb synonymy definition .
Outcome: The proposed resource will be available by mid-2018 .
Evaluation of Domain-specific Word Embeddings using Knowledge Resources (L18-1)

Copied to clipboard

Challenge: Existing word embeddings capture a range of semantic relations relevant to the interpretation of lexical items, but domain-specific terms are difficult to evaluate because of a lack of statistical clues in the underlying corpus.
Approach: They conduct intrinsic and extrinsic evaluations of both general and domain-specific embeddings and adapt embeddment enhancement methods to provide vector representations for infrequent and unseen terms.
Outcome: The proposed model improves both in the intrinsic evaluation and extrinsic evaluation of the embedding models and their representations of infrequent and unseen terms.
Automatic Thesaurus Construction for Modern Hebrew (L18-1)

Copied to clipboard

Challenge: Modern Hebrew lacks lexical resources fundamental to many natural language processing tools.
Approach: They propose a method for generating a cooccurrence based thesaurus in a MRL and a distributional similarity method for Hebrew.
Outcome: The proposed method is not optimal for modern Hebrew, the authors show . they used Hebrew WordNet as their gold standard for the analysis .
Automatic Wordnet Mapping: from CoreNet to Princeton WordNet (L18-1)

Copied to clipboard

Challenge: Existing mappings focus on identifying the semantic categories of CoreNet, but not the word senses.
Approach: They propose to map the word senses of CoreNet into Princeton WordNet synsets by lexical relations by a taxonomy.
Outcome: The proposed mapping bridging the gap between CoreNet and WordNet shows that the word senses of CoreNet are mapped with precision of 91.2%.
The New Propbank: Aligning Propbank with AMR through POS Unification (L18-1)

Copied to clipboard

Challenge: Existing Propbank corpus converts sense labels to a format which is more compatible with AMR and more robust to sparsity.
Approach: They propose a corpus which converts existing Propbank sense labels to a new unified format which is more compatible with AMR and more robust to sparsity.
Outcome: The proposed format is more compatible with AMR and robust to sparsity.
The Boarnsterhim Corpus: A Bilingual Frisian-Dutch Panel and Trend Study (L18-1)

Copied to clipboard

Challenge: a corpus of 250 hours of speech in both west frisian and Dutch is being developed . the corpus is a sociolinguistic corpus based on the recordings of four generations of bilingual speakers .
Approach: This paper describes the Boarnsterhim Corpus project which started in 2016 . it aims to make available 250 hours of speech in both west frisian and Dutch by same speakers .
Outcome: The Boarnsterhim Corpus is a sociolinguistic corpus of west frisian and Dutch speakers . it spans four generations and includes panel and trend data .
The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)

Copied to clipboard

Challenge: The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody .
Approach: They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language .
Outcome: The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies.
Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing project in the field of Swiss German dialects and Swiss French accents collects linguistic data.
Approach: a gamified crowdsourcing platform was set up to collect linguistic data on Swiss German and Swiss French accents.
Outcome: a gamified crowdsourcing platform collects linguistic data on Swiss German and Swiss French accents . the platform has provided 470,000 localizations, with 7,500 registered users and 30,000 anonymous visitors .
Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition (L18-1)

Copied to clipboard

Challenge: a phonetic balance in code-mixed Hindi-English corpus has been created . code-switching is a common phenomenon in multilingual and bilingual communities .
Approach: They propose to create a phonetically balanced read speech corpus of code-mixed Hindi-English . they use a method to select sentences that contain triphones lower in frequency than a threshold .
Outcome: The proposed corpus is phonetically balanced with a large code-mixed reference corpus.
Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts (L18-1)

Copied to clipboard

Challenge: Chinese and Portuguese are very populous languages, but there is not much parallel corpora in the Chinese-Portuguese language pair.
Approach: They propose to curate Chinese-Portuguese parallel corpora and evaluate their quality . they extract bilingual data from government websites and use Phrased-Based Machine Translation (PBMT) and Neural Machine Translation models to build large corpus.
Outcome: The proposed method can be used as a benchmark for future Chinese-Portuguese MT systems.
Evaluating the WordsEye Text-to-Scene System: Imaginative and Realistic Sentences (L18-1)

Copied to clipboard

Challenge: WordsEye is a system for automatically converting natural language text into 3D scenes representing the meaning of that text.
Approach: They evaluate WordsEye's output vs. simple search methods to find a picture to illustrate a sentence.
Outcome: The WordsEye system produced imaginative sentences and realistic sentences . the results show that wordsEyed produced better results than standard image search engines .
Computer-assisted Speaker Diarization: How to Evaluate Human Corrections (L18-1)

Copied to clipboard

Challenge: a framework to evaluate the human corrections of a speaker diarization is presented for the French National Audiovisual Institute (INA) the speaker diaarization task is a necessary pre-processing step for speaker identification and speech transcription.
Approach: They propose a framework to evaluate the human corrections of a speaker diarization . they propose four elementary actions to correct the diarized speaker and an automaton to simulate the correction sequence.
Outcome: The proposed framework copes with the needs of the French National Audiovisual Institute (INA) due to the increasing number of documents and the limited number of annotators, many documents remain undocumented or only partly documented.
Performance Impact Caused by Hidden Bias of Training Data for Recognizing Textual Entailment (L18-1)

Copied to clipboard

Challenge: a method to improve the quality of training data is needed . annotation errors of dialog act corpus mislead learning results of Bayesian network .
Approach: They propose to introduce a null hypothesis for predictability of textual entailment labels and test it using a Naive Bayes model.
Outcome: The proposed method does not reject the null hypothesis, but it improves on the existing models.
Evaluation of Croatian Word Embeddings (L18-1)

Copied to clipboard

Challenge: Currently, research is focusing mostly on English.
Approach: They propose to use word analogy datasets to evaluate word similarities in Croatian . they use Word2Vec and FastText to create word analogies from highdimensional space .
Outcome: The proposed datasets show that word embeddings are able to capture the syntactic and semantic relationship between words.
C-HTS: A Concept-based Hierarchical Text Segmentation approach (L18-1)

Copied to clipboard

Challenge: Existing approaches to hierarchical text segmentation use lexical and/or syntactic similarity to identify the coherent segments of text.
Approach: They propose a Concept-based Hierarchical Text Segmentation approach that uses the semantic relatedness between text constituents to represent meaning.
Outcome: The proposed method performs well on two publicly available datasets.
Semantic Supersenses for English Possessives (L18-1)

Copied to clipboard

Challenge: Existing semantic categories for possessive constructions are limited to nominals and s-genitives.
Approach: They propose to use a supersense inventory to annotate English possessives . they show existing supersensor categories are readily applicable to possessives.
Outcome: The proposed annotations are applied to English possessives in a corpus of web reviews.
A Corpus of Metaphor Novelty Scores for Syntactically-Related Word Pairs (L18-1)

Copied to clipboard

Challenge: Existing data on metaphor novelty are limited, making it difficult to perform research on this topic.
Approach: They propose to release a corpus of metaphor novelty scores for syntactically related word pairs . they establish a performance benchmark to which future researchers can compare .
Outcome: The proposed corpus of metaphor novelty scores is compared to other datasets . it performs better than chance or nave strategies, the authors show .
Improving Hypernymy Extraction with Distributional Semantic Classes (L18-1)

Copied to clipboard

Challenge: Existing methods for extracting hypernyms focus on the acquisition of binary hypernies .
Approach: They propose a distributionally-induced semantic class for extracting hypernyms . they also use distributional semantics to induce sense-aware semantic classes .
Outcome: The proposed method improves the quality of the hypernymy extraction in terms of precision and recall.
Laying the Groundwork for Knowledge Base Population: Nine Years of Linguistic Resources for TAC KBP (L18-1)

Copied to clipboard

Challenge: Knowledge Base Population (KBP) evaluations target information extraction technologies for knowledge bases comprised of entities, relations, and events.
Approach: They describe the linguistic resources provided by Linguistic Data Consortium for TAC KBP since 2009 . they highlight changes made to support evolving evaluation requirements .
Outcome: The evaluations have targeted information extraction technologies for the population of knowledge bases comprised of entities, relations, and events.
A Dataset for Inter-Sentence Relation Extraction using Distant Supervision (L18-1)

Copied to clipboard

Challenge: Existing methods for intra-sentence relation extraction use a distance supervision method to extract relations between entities.
Approach: They propose a benchmark dataset for the task of inter-sentence relation extraction using relations previously used for intra-sentent relation extraction.
Outcome: The proposed dataset is compared with baseline models and recurrent neural network models on the developed dataset.
Diacritics Restoration Using Neural Networks (L18-1)

Copied to clipboard

Challenge: a novel combination of character-level recurrent neural network and language model is proposed . people often replace characters with diacritics with their ASCII counterparts .
Approach: They propose a character-level recurrent neural network-based model and a language model for diacritics restoration.
Outcome: The proposed model reduces error of current best systems by 20% to 64% on four languages . it is also able to restore diacritical marks on a number of languages using the same model .
Ensemble Romanian Dependency Parsing with Neural Networks (L18-1)

Copied to clipboard

Challenge: SSPR is a Python 3.5 application based on the Microsoft Cognitive Toolkit 2.0 Python API.
Approach: a Python 3.5 application is based on the Microsoft Cognitive Toolkit 2.0 Python API.
Outcome: SSPR outperforms the best individual parser at the CONLL 2017 dependency parsing shared task.
Classifying Sluice Occurrences in Dialogue (L18-1)

Copied to clipboard

Challenge: Ellipsis is an important challenge for natural language processing systems, says a new paper . previous work on ellipsis focused on news data, but sluicing presents a challenge for dialogue systems .
Approach: They describe a corpus of 4100 sluice occurrences from the NYTimes Gigaword corpus . they build a classifier model to automatically classify slujce .
Outcome: The proposed corpus contains 4100 sluice occurrences, with an accuracy of 67% . the work will support empirical research into slujcing in dialogue systems .
Collection of Multimodal Dialog Data and Analysis of the Result of Annotation of Users’ Interest Level (L18-1)

Copied to clipboard

Challenge: a group of researchers is building a corpus for evaluating elements of multimodal dialogue systems.
Approach: They propose to build a corpus for evaluating elements of the multimodal dialogue system . they use the Wizard of Oz method to record chat dialogue data between a human and a virtual agent .
Outcome: The proposed method annotates chat dialogue data between a human and a virtual agent and measures their interest level in the data.
Recognizing Behavioral Factors while Driving: A Real-World Multimodal Corpus to Monitor the Driver’s Affective State (L18-1)

Copied to clipboard

Challenge: Existing studies on the induced emotional states of the driver in a car have not been published.
Approach: They used three sensor systems to collect emotional multimodal data while driving . they defined neutral, positive, frustrated and anxious states of the driver .
Outcome: The collected data were analyzed using a Wizard-of-Oz technique . the participants were asked to fill out questionnaires and annotate the data .
EmotionLines: An Emotion Corpus of Multi-Party Conversations (L18-1)

Copied to clipboard

Challenge: Emotion is a critical characteristic to distinguish people from machines.
Approach: They propose a dataset with emotions labeling on all utterances in each dialogue . they use Friends TV scripts and Facebook messenger dialogues to collect the data .
Outcome: The proposed dataset is the first with emotions labeling on all utterances in each dialogue based on their textual content.
Academic-Industrial Perspective on the Development and Deployment of a Moderation System for a Newspaper Website (L18-1)

Copied to clipboard

Challenge: a system that supports the moderation of user comments on a large newspaper website is described in this paper.
Approach: They describe an approach and experiences from the development, deployment and usability testing of a natural language processing and information retrieval system that supports the moderation of user comments on a large newspaper website.
Outcome: The proposed system supports the moderation of user comments on a large newspaper website.
Community-Driven Crowdsourcing: Data Collection with Local Developers (L18-1)

Copied to clipboard

Challenge: a community-driven approach to annotation applications and crowdsourcing programs is feasible, says a new study.
Approach: They propose to partner with local developers to create custom annotation applications . they recruit and motivate crowd contributors from their communities to perform an annotation task .
Outcome: The proposed approach combines local developers' knowledge of their social networks to collect labeled data.
Building Open Javanese and Sundanese Corpora for Multilingual Text-to-Speech (L18-1)

Copied to clipboard

Challenge: Using multi-speaker text-to-speech systems, we build systems for Javanese and Sundanese . progress in this direction is difficult because languages in the long tail of the distribution of the majority of the world's languages lack adequate linguistic resources .
Approach: They present multi-speaker text-to-speech corpora for Javanese and Sundanese . they use mixed-gender recordings to build multi-language text-based systems .
Outcome: The proposed multi-speaker text-to-speech systems outperform the systems constructed from a single language.
An Integrated Representation of Linguistic and Social Functions of Code-Switching (L18-1)

Copied to clipboard

Challenge: Linguistic studies on code-switching focus on the "how" and "why" of CS . a new model aims to derive CS functions from local and global properties of the code-witched discourse .
Approach: They propose a model that integrates CS phenomena and modalities into a representation that includes local and global properties of the code-switched discourse.
Outcome: The proposed model simplifies the analysis of English/Hindi CS datasets and provides a flexible framework for further studies.
A Corpus of eRulemaking User Comments for Measuring Evaluability of Arguments (L18-1)

Copied to clipboard

Challenge: eRulemaking is a way for government agencies to directly reach citizens to solicit their opinions and experiences regarding newly proposed rules.
Approach: They propose an argument mining corpus annotated with argumentative structure information capturing the evaluability of arguments.
Outcome: The proposed corpus contains 731 user comments on consumer debt collection practices rule by the Consumer Financial Protection Bureau.
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)

Copied to clipboard

Challenge: Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data.
Approach: They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes.
Outcome: The proposed corpus includes 112 argumentative microtexts and a new annotation tool.
Discourse Coherence Through the Lens of an Annotated Text Corpus: A Case Study (L18-1)

Copied to clipboard

Challenge: a corpus-based study of local coherence as established by anaphoric links in discourses is presented . the study uses Czech data present in the Prague Dependency Treebank 3.0 .
Approach: They propose to look at local coherence as established by anaphoric links between elements in the thematic and rhematic parts of sentences.
Outcome: The proposed method takes into account syntactic relations, contextual boundness and coreference and bridging relations.
Automatic Prediction of Discourse Connectives (L18-1)

Copied to clipboard

Challenge: Discourse connectives are used to bind together and explicate the relation between pieces of text.
Approach: They propose to use a dataset of 2.9M sentence pairs separated by discourse connectives to test their accuracy.
Outcome: The proposed model outperforms the human model in the prediction task . the proposed model has a higher F1 under specific conditions .
Handling Rare Word Problem using Synthetic Training Data for Sinhala and Tamil Neural Machine Translation (L18-1)

Copied to clipboard

Challenge: Lack of parallel training data influences rare word problem in Neural Machine Translation systems, especially for underresourced languages.
Approach: They propose to use Parts of Speech tagging and morphological analysis as syntactic features to prune generated synthetic sentence pairs that do not adhere to language syntax.
Outcome: The proposed methods show that they can prune sentences that do not adhere to language syntax over Sinhala to Tamil and Tamil to Sinhalak translation systems.
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .
Creating a Translation Matrix of the Bible’s Names Across 591 Languages (L18-1)

Copied to clipboard

Challenge: In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed .
Approach: They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output.
Outcome: The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages.
Building a Word Segmenter for Sanskrit Overnight (L18-1)

Copied to clipboard

Challenge: Sanskrit word segmentation is challenging due to the issue of Sandhi . digitisation efforts have made the manuscripts available in the public domain .
Approach: They propose a deep sequence to sequence model that takes only the sandhied string as input and predicts the unsandhized string.
Outcome: The proposed model improves on the current state of the art by 16.79% . the system can be trained "overnight" and be used for production .
Simple Semantic Annotation and Situation Frames: Two Approaches to Basic Text Understanding in LORELEI (L18-1)

Copied to clipboard

Challenge: Existing annotations for low resource languages are under-resourced for human language technology, but lack of resources does not correlate with lack of need for such technologies.
Approach: They propose two types of semantic annotation for the DARPA Low Resource Languages for Emerging Incidents program: Simple Semantic Annotation (SSA) and Situation Frames (SF).
Outcome: The proposed approaches are aimed at labeling basic semantic information relevant to humanitarian aid and disaster relief scenarios.
Abstract Meaning Representation of Constructions: The More We Include, the Better the Representation (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) uses a flexible pattern or template of multiple lexical items to provide semantic representation of certain constructions.
Approach: They propose to expand the AMR project's lexicon of predicate senses to include entries for a growing set of constructions.
Outcome: The proposed approach provides coverage for the annotation of certain types of constructions.
Evaluating Scoped Meaning Representations (L18-1)

Copied to clipboard

Challenge: Semantic parsing offers many opportunities to improve natural language understanding . current research on open-domain semantic parsers focuses on supervised learning methods .
Approach: They propose a semantically annotated parallel corpus for English, German, Italian, and Dutch . they use a matching tool to evaluate scoped meaning representations to match clauses .
Outcome: The proposed method captures the semantics of negation, modals, quantification, and presupposition triggers . it compares scoped meaning representations to gold standard parsers and finds improvements .
Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation (L18-1)

Copied to clipboard

Challenge: Word Sense Disambiguation is a crucial task in Natural Language Processing . supervised systems need to be trained on word-by-word basis, a problem that is beyond reach for resource-rich languages like English.
Approach: They release six large-scale sense-annotated datasets in multiple languages to pave the way for supervised multilingual Word Sense Disambiguation.
Outcome: The results show that large-scale sense annotations can be used as training sets for supervised systems.
SentEval: An Evaluation Toolkit for Universal Sentence Representations (L18-1)

Copied to clipboard

Challenge: a toolkit for evaluating the quality of universal sentence representations is available for download and preprocessing . word embeddings are not trained to perform well on one specific task, but their value lies in their transferability . evaluation of general-purpose word and sentence embeddables has been problematic .
Approach: They propose a toolkit to evaluate the quality of universal sentence representations.
Outcome: The proposed toolkit includes scripts to download and preprocess datasets and an easy interface to evaluate sentence encoders.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition (L18-1)

Copied to clipboard

Challenge: a new general framework for sign recognition from monocular video is presented . the framework exploits state-of-the-art learning methods while incorporating features based on what we know about the linguistic composition of lexical signs.
Approach: They propose a general framework for sign recognition from monocular video . they exploit state-of-the-art learning methods while incorporating features from linguistic information .
Outcome: The proposed framework exploits state-of-the-art learning methods while incorporating features based on what we know about linguistic composition of lexical signs.
CONDUCT: An Expressive Conducting Gesture Dataset for Sound Control (L18-1)

Copied to clipboard

Challenge: Recent studies on music-gesture relationship focus on sound variations and expressiveness of gestures.
Approach: They propose to use a database to create a set of expressive gestures for orchestral conductors . they assume that the gestures convey some meaning shared by most conductor .
Outcome: The proposed database will be used to train a gesture recognition system for live sound control and modulation.
Neural Caption Generation for News Images (L18-1)

Copied to clipboard

Challenge: Existing methods for automatic caption generation of images are lacking in the field of image-related applications.
Approach: They propose a method for automatically generating captions for news images . they propose several deep neural network architectures built upon Recurrent Neural Networks .
Outcome: The proposed method outperforms a traditional method on a BBC News dataset using automatic evaluation and human evaluation.
MPST: A Corpus of Movie Plot Synopses with Tags (L18-1)

Copied to clipboard

Challenge: a corpus of movie plot synopses and tags can be used to build automatic tagging systems . a method to collect these tags allows us to learn to predict tags from plot synoopsis .
Approach: They propose to collect a corpus of movie plot synopses and 70 tags to analyze their properties.
Outcome: The proposed method can be used to predict movie tags from plot synopses.
OpenSubtitles2018: Statistical Rescoring of Sentence Alignments in Large, Noisy Parallel Corpora (L18-1)

Copied to clipboard

Challenge: Movie and TV subtitles are a valuable resource for the compilation of parallel corpora . however, the quality of the resulting sentence alignments is often lower than for other parallel corpoora.
Approach: They propose to use movie and TV subtitles to extract parallel corpora from 3.7 million subtitles spread over 60 languages to obtain explicit quality scores for each sentence alignment.
Outcome: The proposed model predicts translation probabilities with a root mean square error of 0.07 . the results show that the model can prune out low-quality alignments .
Building an Ellipsis-aware Chinese Dependency Treebank for Web Text (L18-1)

Copied to clipboard

Challenge: ellipsis is a common linguistic phenomenon that some words are left out as they are understood from the context, especially in oral utterance.
Approach: They propose to use a Chinese dependency treebank to facilitate the parsing of web text . they propose to restore omissions and reserve contexts in the web text to improve dependency parsers .
Outcome: The proposed framework enables the parsing of web text from online microblogs.
EuroGames16: Evaluating Change Detection in Online Conversation (L18-1)

Copied to clipboard

Challenge: a new method for detecting salient changes from on-line conversations is needed . linguistic preprocessing and time series are used to build a time series .
Approach: They propose a framework for detecting salient changes from on-line conversations . they use linguistic preprocessing to build a time series and change point detection algorithms to detect salient change.
Outcome: The proposed method can detect salient changes in on-line conversations with high accuracy.
A Deep Neural Network based Approach for Entity Extraction in Code-Mixed Indian Social Media Text (L18-1)

Copied to clipboard

Challenge: a huge number of people use social media to express and exchange information in their own languages.
Approach: They propose to use a code-mixed environment to extract higher level features from text . they use 'gadget' algorithm that automatically discovers higher level feature from text.
Outcome: The proposed approach is generic and does not make use of handcrafted features or rules.
PoSTWITA-UD: an Italian Twitter Treebank in Universal Dependencies (L18-1)

Copied to clipboard

Challenge: Various approaches and ad hoc resources are needed to provide proper coverage of specific linguistic phenomena.
Approach: They propose to annotate tweets using a well-known dependency-based annotation format . they propose to use the tweets for training NLP systems to improve their performance .
Outcome: The proposed resource can be used for training of NLP systems on social media texts.
Annotating If the Authors of a Tweet are Located at the Locations They Tweet About (L18-1)

Copied to clipboard

Challenge: a tweet's locations do not always indicate spatial information involving the author of the tweet . a corpus of 1,062 tweets contains 1,200 location named entities .
Approach: They propose a corpus annotating whether tweet authors are located in locations . they use temporal tags centered around tweet timestamps to temporally anchor this information .
Outcome: The proposed method annotates whether authors are located in tweet locations . it shows that no spatial relationship can be inferred in 21% of instances .
MOCCA: Measure of Confidence for Corpus Analysis - Automatic Reliability Check of Transcript and Automatic Segmentation (L18-1)

Copied to clipboard

Challenge: The production of speech corpora typically involves manual labor to verify and correct the output of automatic transcription/segmentation processes.
Approach: They propose to use Support Vector Machine/Support Vector Regression and Random Forest to predict transcription errors in an annotated speech corpus.
Outcome: The proposed methods can be implemented as free-to-use common language and resources and technology infrastucture web services.
Towards an ISO Standard for the Annotation of Quantification (L18-1)

Copied to clipboard

Challenge: Quantification occurs in every sentence of written text or spoken discourse because application of a predicate to one or more sets of objects gives rise to questions of relative scope, of cardinality, and of distribution (or 'distributivity') of the predicacy over the sets of arguments.
Approach: They propose an approach to the annotation of quantification that is being developed as part of an effort by the International Organisation for Standardisation ISO to define interoperable semantic annotation schemes.
Outcome: The proposed scheme includes both count and mass NP quantifiers, as well as NPs with syntactically and semantically complex heads with internal quantification and scoping structures.
Lightweight Grammatical Annotation in the TEI: New Perspectives (L18-1)

Copied to clipboard

Challenge: a small set of descriptive devices have been made available for lightweight linguistic annotation . merit of a predefined TEI tagset is the homogeneity of tagging and better interoperability of simple linguistic resources encoded in the TE.
Approach: They propose a new attribute class that would gather token-level attributes facilitating simple linguistic annotation.
Outcome: The proposed attribute class addresses community feedback on the lack of a specific tagset for lightweight linguistic annotation within the TEI.
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.
Handling Big Data and Sensitive Data Using EUDAT’s Generic Execution Framework and the WebLicht Workflow Engine. (L18-1)

Copied to clipboard

Challenge: a new workflow engine for web-based tools and workflow engines can be used to process big data and data with restrictive property rights.
Approach: They propose to bring WebLicht workflow engine with EUDAT-based Generic Execution Framework to address this issue.
Outcome: The proposed workflow engine can handle large data sets with restrictive property rights . the EUDAT project is developing the Generic Execution Framework (GEF)
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)

Copied to clipboard

Challenge: DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing .
Approach: They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus .
Outcome: The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset.
Universal Dependencies Version 2 for Japanese (L18-1)

Copied to clipboard

Challenge: UD Japanese resources are built on automatic conversion from several treebanks.
Approach: They propose to port the word delimitation, POS, and syntactic relations of existing treebanks to UD Japanese . they discuss the issues of the UD scheme found through porting of the Japanese language .
Outcome: The proposed UD Japanese resources are based on automatic conversion from treebanks.
Developing the Bangla RST Discourse Treebank (L18-1)

Copied to clipboard

Challenge: a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla .
Approach: They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus .
Outcome: The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications .
A New Version of the Składnica Treebank of Polish Harmonised with the Walenty Valency Dictionary (L18-1)

Copied to clipboard

Challenge: wigra parser was used to generate the treebank, but the differences between the resources made it necessary to manually correct some parse trees.
Approach: They propose a procedure to update manually disambiguated trees of Skadnica due to the switch to the Walenty valency dictionary.
Outcome: The proposed method allows to check the consistency of the treebank and valence dictionary.
Parse Me if You Can: Artificial Treebanks for Parsing Experiments on Elliptical Constructions (L18-1)

Copied to clipboard

Challenge: ellipsis is a phenomenon present in many natural languages, but it complicates syntactic parsing of the content that is not omitted.
Approach: They analyze outputs of state-of-the-art parsers to learn about parsing accuracy and typical errors from the perspective of elliptical constructions.
Outcome: The proposed treebank is a semi-artificially constructed treebank of ellipsis.
Semi-Automatic Construction of Word-Formation Networks (for Polish and Spanish) (L18-1)

Copied to clipboard

Challenge: a semi-automatic method for the construction of derivational networks is proposed . the proposed method is general enough to be adopted for other languages .
Approach: They propose a semi-automatic method for the construction of derivational networks using a sequential pattern mining technique.
Outcome: The proposed method is general enough to be adopted for other languages.
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)

Copied to clipboard

Challenge: Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics .
Approach: They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative.
Outcome: The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy .
UniMorph 2.0: Universal Morphology (L18-1)

Copied to clipboard

Challenge: The Universal Morphology project is a collaborative effort to improve how NLP handles complex morphology across the world's languages.
Approach: They propose to use a universal tagset to annotate morphological data using a schema that includes a lemma and a bundle of morphology features.
Outcome: The project releases annotated morphological data using a universal tagset, the UniMorph schema.
A Computational Architecture for the Morphology of Upper Tanana (L18-1)

Copied to clipboard

Challenge: a computational model of Upper Tanana is described to model the Dene language . the model uses lexical-inflectional verb classes to predict possible derivations and their morphological behavior.
Approach: They propose a computational model of Upper Tanana, a highly endangered Dene language . the model parses and generates inflected Upper Tanans and uses a lexical-inflectional verb system to predict possible derivations and their morphological behavior.
Outcome: The proposed model parses and generates inflected Upper Tanana verb forms . it also uses the language's verb theme category system to predict possible derivations and their morphological behavior .
Expanding Abbreviations in a Strongly Inflected Language: Are Morphosyntactic Tags Sufficient? (L18-1)

Copied to clipboard

Challenge: In this paper, the problem of recovery of morphological information lost in abbreviated forms is addressed . correct inflected form of expanded abbrevation can be deduced from context words .
Approach: They propose a deep bidirectional LSTM network with tag embedding to predict abbreviated words . they train on 10 million words from the Polish Sejm Corpus and achieve 74.2% prediction accuracy .
Outcome: The proposed model achieves 74.2% accuracy on a smaller but more general corpus of Polish words.
A High-Quality Gold Standard for Citation-based Tasks (L18-1)

Copied to clipboard

Challenge: Citation recommendation tasks involve recommending citations within their specific contexts.
Approach: They propose to use arXiv.org's citation-dependent evaluation data set to evaluate citations . their data set is characterized by the fact that it exhibits almost zero noise in its extracted content .
Outcome: The proposed data set exhibits almost zero noise in extracted content and all citations are linked to their correct publications.
Measuring Innovation in Speech and Language Processing Publications. (L18-1)

Copied to clipboard

Challenge: The authors of this paper analyze the publications in the field of speech and language processing.
Approach: They propose to analyze publications in the field of speech and language processing to measure innovation . they use the corpus of papers published over 50 years and enlarge it to the SNLP corpus .
Outcome: The proposed method enlarges the corpus of 65,003 documents published over 50 years to the Speech and Language Processing (SNLP) corpus.
PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles (L18-1)

Copied to clipboard

Challenge: Existing tools to extract structured textual content from PDFs are essential to enable scientific text mining.
Approach: They propose a PDF-to-XML textual content extraction tool that extracts structured textual contents from scientific articles in PDF format.
Outcome: The proposed tool extracts structured textual content from scientific articles in PDF format while preserving both the textual contents and layout details of the input PDF document.
Automatic Identification of Research Fields in Scientific Papers (L18-1)

Copied to clipboard

Challenge: TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers.
Approach: TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers.
Outcome: The proposed method will help scientists identify geographical territories areas from scientific papers available in digital versions within and outside the ISTEX library.
A «Portrait» Approach to Multichannel Discourse (L18-1)

Copied to clipboard

Challenge: a new study examines the multichannel discourse analysis of human communication . we examine the individual variation in multichannel behavior .
Approach: They propose to use a multichannel resource to study multichannel discourses . they propose to analyze verbal structure, prosody, gesticulation, facial expression, eye gaze .
Outcome: The proposed method is crucially important for fine-grained annotation procedures and statistical analyses of multichannel data.
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)

Copied to clipboard

Challenge: Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications.
Approach: They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages.
Outcome: The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages.
Building a Macro Chinese Discourse Treebank (L18-1)

Copied to clipboard

Challenge: Discourse structure analysis is an important research topic in natural language processing.
Approach: They propose to construct a macro discourse structure framework and annotate 147 Newswire articles.
Outcome: The proposed framework can lay the foundation for further analysis of macro discourse structure.
Enhancing the AI2 Diagrams Dataset Using Rhetorical Structure Theory (L18-1)

Copied to clipboard

Challenge: Existing annotation schemas for diagrams are based on Rhetorical Structure Theory (RST) paper documents proposed schema, reports on inter-annotator agreement for this task, and discusses use of AI2D-RST for research on multimodality and artificial intelligence.
Approach: They propose to replace the annotation of semantic relations between diagram elements by building on Rhetorical Structure Theory (RST) the paper documents the proposed annotation schema, describes challenges in applying RST to diagrams, and reports on inter-annotator agreement for this task.
Outcome: The proposed schema is based on Rhetorical Structure Theory, which has been used to describe the multimodal structure of diagrams and documents.
QUD-Based Annotation of Discourse Structure and Information Structure: Tool and Evaluation (L18-1)

Copied to clipboard

Challenge: a new annotation scheme and discourse-analytic method is developed for information structure annotation.
Approach: They propose a new annotation scheme and a discourse-analytic method based on Questions under Discussion . they introduce a tool which enables the analyst to semi-automatically segment texts and enhance them with QUDs .
Outcome: The proposed method achieves good inter-annotator scores and good agreement with discourse annotations.
The Spot the Difference corpus: a multi-modal corpus of spontaneous task oriented spoken interactions (L18-1)

Copied to clipboard

Challenge: a new corpus of task-oriented spontaneous dialogues is available for free . we study the structure of non task-orientated dialogues and how they evolve over time .
Approach: They describe a corpus of 54 interactions between pairs of subjects interacting to find differences in two very similar scenes.
Outcome: The proposed dataset contains 54 interactions between subjects to find differences in two very similar scenes.
Attention for Implicit Discourse Relation Recognition (L18-1)

Copied to clipboard

Challenge: Existing approaches to implicit discourse relation recognition reach F1 scores of 9.95% to 37.67% . a neural network exploits the strong correlation between pairs of words that implicitly signal a discourse relation.
Approach: They propose a neural network which exploits strong correlation between pairs of words . they use an encoder-decoder model with attention to detect a latent discourse relation .
Outcome: The proposed model outperforms state-of-the-art models on fine-grained classification and fine-granular classification while computing parameters without pooling and fully connected layers.
A Context-based Approach for Dialogue Act Recognition using Simple Recurrent Neural Networks (L18-1)

Copied to clipboard

Challenge: Existing models of dialogue act classification work on the utterance-level and only very few consider context.
Approach: They propose to use a character-level language model to classify dialogue acts without context . they find that the preceding utterances are a context of the current utterant .
Outcome: The proposed method improves on the Switchboard Dialogue Act corpus . it includes context and leads to 3% higher accuracy .
TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations (L18-1)

Copied to clipboard

Challenge: TREEANNOTATOR is a browser-based tool for annotating tree-like structures . it provides a wider range of formats and provides graphical annotations .
Approach: They evaluate TREEANNOTATOR, a browser-based tool for annotating tree-like structures, in particular structures that jointly map dependency relations and inclusion hierarchies, as used by Rhetorical Structure Theory.
Outcome: The GUI interface is user-friendly and provides two visualization modes.
Chats and Chunks: Annotation and Analysis of Multiparty Long Casual Conversations (L18-1)

Copied to clipboard

Challenge: dyadic conversations are attracting more interest with attempts to build more friendly and natural spoken dialog systems.
Approach: They describe the collection, organization, and annotation of structural chat and chunk phases in three existing corpora and analyse their preliminary results to find that chunk dominates as conversations get longer.
Outcome: The results show that chunk dominates conversations as they get longer .
Extending the gold standard for a lexical substitution task: is it worth it? (L18-1)

Copied to clipboard

Challenge: a lexical substitution task requires systems to identify words that are semantically close to the target and to select among candidates those that best fit the context.
Approach: They propose to use a lexical substitution task to evaluate systems' performance . they use 300 sentences containing a target word and a second dataset based on the same data .
Outcome: The proposed model is based on a set of 300 sentences containing a target word . the proposed model has not been evaluated to our knowledge .
Lexical and Semantic Features for Cross-lingual Text Reuse Classification: an Experiment in English and Latin Paraphrases (L18-1)

Copied to clipboard

Challenge: Analyzing historical languages is challenging because they lack primary material for certain time periods . under-resourced languages such as Ancient Greek and Latin lack advanced natural-language processing (NLP) techniques .
Approach: They propose to use machine learning to detect and classify paraphrastic text reuse in historical texts.
Outcome: The proposed method improves the accuracy of paraphrastic text reuse detection in historical languages.
Investigating the Influence of Bilingual MWU on Trainee Translation Quality (L18-1)

Copied to clipboard

Challenge: a method for automatic extraction of bilingual multiword units (BMWUs) from a parallel corpus has been shown to be useful for estimating human translation quality.
Approach: They applied a method for automatic extraction of bilingual multiword units from a parallel corpus in order to investigate their contribution to translation quality in terms of adequacy and fluency.
Outcome: The method is based on generalized additive modelling and it shows that normalized BMWU ratios can be useful for estimating human translation quality.
Evaluation of Dictionary Creating Methods for Finno-Ugric Minority Languages (L18-1)

Copied to clipboard

Challenge: a project aims to provide linguistically based support for small Finno-Ugric (FU) digital communities to generate online content and revitalize the digital functions of some FU minority languages.
Approach: They evaluate bilingual dictionary building methods for six small fino-ugric minority languages . they use Wikipedia title pairs extracted via inter-language links and Wiktionary-based methods .
Outcome: The proposed methods proved that standard lexicon building methods are low for under-resourced languages.
Dysarthric speech evaluation: automatic and perceptual approaches (L18-1)

Copied to clipboard

Challenge: Perceptual evaluation is still the most common method in clinical practice for the diagnosis and monitoring of the condition progression of people suffering from dysarthria.
Approach: They propose an automatic approach for anomaly detection at the phone level for dysarthric speech . they propose a perceptual evaluation protocol that uses annotated french corpora to analyze the system behavior.
Outcome: The proposed method was validated on different corpora and speech styles.
Towards an Automatic Assessment of Crowdsourced Data for NLU (L18-1)

Copied to clipboard

Challenge: Recent development of spoken dialog systems aims at allowing a natural input style.
Approach: They investigate how crowdsourced data can be assessed with respect to its naturalness and usefulness by using a word based language model to identify valid data.
Outcome: The proposed methods show that valid data can be identified with the help of a word based language model.
Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning (L18-1)

Copied to clipboard

Challenge: Existing methods for evaluating plausibility of events are focused on measuring causal dependency between events or actions.
Approach: They propose a task to identify the more plausible alternative with their commonsense causal context.
Outcome: The proposed task is based on a visual COPA dataset with 380 questions and over 1K images with various topics.
Is it worth it? Budget-related evaluation metrics for model selection (L18-1)

Copied to clipboard

Challenge: linguistic resources can be labor-intensive, requiring great amounts of work-hours and expert annotation.
Approach: They propose a machine learning model that pre-annotates or filters content before annotating it . they argue that the model with the highest F-score may not have best separation .
Outcome: a case study shows that the model with the highest F-score does not yield the highest profits . the model that has the highest score does not produce the highest profit, the study shows .
Automated Evaluation of Out-of-Context Errors (L18-1)

Copied to clipboard

Challenge: Existing methods to modify text understanding systems use only one sentence at a time . however, considering a larger context can improve performance for text understanding tasks.
Approach: They propose to modify existing text data to insert out-of-context errors . they use a 2016 TEDTalk corpus to evaluate computational models for text understanding .
Outcome: The proposed method targets real-world problems of transcription and translation systems by inserting authentic out-of-context errors.
Matics Software Suite: New Tools for Evaluation and Data Exploration (L18-1)

Copied to clipboard

Challenge: Numerous works propose interfaces or frameworks to build, explore and visualize corpora of annotated data.
Approach: Matics proposes a dataframe data model for exploring annotated data and evaluation results.
Outcome: The tools already run on several Natural Language Processing tasks and standard annotation formats, and are under on-going development.
MGAD: Multilingual Generation of Analogy Datasets (L18-1)

Copied to clipboard

Challenge: Existing methods for word embedding evaluation are computationally expensive and task-specific.
Approach: They propose a minimally supervised method for generating word embedding evaluation datasets for a large number of languages using existing dependency treebanks and parsers.
Outcome: The proposed method evaluates three popular word embedding algorithms against these datasets and shows that their performance varies between syntactic categories.
MIsA: Multilingual “IsA” Extraction from Corpora (L18-1)

Copied to clipboard

Challenge: In this paper, we present a collection of hypernymy relations extracted from the Wikipedia corpus in five languages.
Approach: They present a collection of hypernymy relations extracted from the Wikipedia corpus in five languages . they use existing or newly defined lexico-syntactic patterns to extract hyperniyms .
Outcome: The proposed tool is based on a dictionary extracted from the full Wikipedia corpus.
Biomedical term normalization of EHRs with UMLS (L18-1)

Copied to clipboard

Challenge: Currently, there is no tool for this language and for this specific purpose.
Approach: They propose a multilingual and cross-lingual tool that searches for biomedical terms in clinical texts with the Unified Medical Language System (UMLS) Metathesaurus.
Outcome: The proposed tool performs biomedical term normalization in clinical texts with the unified medical language system (UMLS) Metathesaurus . it is based on Apache Lucene TM and is available on-line 2 .
Revisiting the Task of Scoring Open IE Relations (L18-1)

Copied to clipboard

Challenge: Recent Open Information Extraction systems allow us to extract ever larger (yet incomplete) open-domain Knowledge Bases from text.
Approach: They propose a baseline model which gives competitive results in a previously defined protocol and provides an independent source of signal to judge arbitrary fact plausibility.
Outcome: The proposed model gives competitive results in the previously defined protocol and provides an independent source of signal to judge arbitrary fact plausibility.
A supervised approach to taxonomy extraction using word embeddings (L18-1)

Copied to clipboard

Challenge: a recent evaluation of a method for organizing texts into a hierarchy showed that it did not outperform a baseline.
Approach: They propose a method that uses supervised learning to combine multiple features with a support vector machine classifier including the baseline features.
Outcome: The proposed method outperforms the baseline method and provides stronger method for identifying taxonomic relations than previous methods.
A Chinese Dataset with Negative Full Forms for General Abbreviation Prediction (L18-1)

Copied to clipboard

Challenge: a common phenomenon across languages is abbreviation, but it's not always possible to predict it accurately.
Approach: They build a dataset for general Chinese abbreviation prediction using a negative full form . they find that abbrevation prediction can improve the performance of abbreviation recognition .
Outcome: The proposed dataset evaluates models on abbreviation prediction in Chinese . it shows that abbrevation prediction improves performance in language processing tasks .
Korean TimeBank Including Relative Temporal Information (L18-1)

Copied to clipboard

Challenge: Temporal information extraction is one of the important research fields in natural language processing.
Approach: They propose a concept of relative temporal information and supplement a Korean annotation language to represent new relative expressions and extend an annotated dataset through the revised language.
Outcome: The proposed language can be used to represent relative temporal information and extend an annotated dataset, Korean TimeBank, through the revised language.
Mining Biomedical Publications With The LAPPS Grid (L18-1)

Copied to clipboard

Challenge: Natural language processing (NLP) text mining can increase productivity and innovation in the sciences by orders of magnitude.
Approach: The Language Applications Grid is an infrastructure for rapid development of natural language processing applications (NLP) it provides an intuitive and easy-to-use platform for users to exploit NLP tools and resources . the Grid integrates the services and resources provided by PubAnnotation to greatly enhance the user's ability to annotate scientific publications .
Outcome: The Language Applications (LAPPS) Grid is an infrastructure for rapid development of natural language processing applications (NLP) it integrates services and resources provided by PubAnnotation to greatly enhance user's ability to annotate scientific publications and share the results.
An Initial Test Collection for Ranked Retrieval of SMS Conversations (L18-1)

Copied to clipboard

Challenge: a test collection for retrieving SMS content is described . the collection contains 31 topics, which are considered too few for reliable statistical significance tests.
Approach: They describe a test collection for evaluating systems that search SMS conversations . the collection is built from 120,000 text messages .
Outcome: The proposed test collection can be used to compare some alternative retrieval systems.
FrNewsLink : a corpus linking TV Broadcast News Segments and Press Articles (L18-1)

Copied to clipboard

Challenge: a corpus of TV Broadcast News resources is proposed to address several applicative tasks.
Approach: They propose to use a corpus to address several applicative tasks that are made public . they propose to gather TVBN shows and press articles and use them to study semantic similarity .
Outcome: The proposed corpus is based on 112 TVBN shows and press articles . it allows to study semantic similarity and multimedia News linking .
PyRATA, Python Rule-based feAture sTructure Analysis (L18-1)

Copied to clipboard

Challenge: a new Python module supports rules-based analysis on structured data . the module is available under the Apache V2 license .
Approach: They propose a Python module which supports rules-based analysis on structured data.
Outcome: The proposed module supports rules-based analysis on structured data.
Towards Processing of the Oral History Interviews and Related Printed Documents (L18-1)

Copied to clipboard

Challenge: a project aims to create an integrated archive of the recordings, scanned documents and photographs from totalitarian regimes in Czechoslovakia . the archive will be accessible online and provide multifaceted search capabilities .
Approach: They propose to use automatic speech recognition and optical character recognition to build an archive of the interviews, scanned documents and photographs.
Outcome: The proposed archive will be accessible online and provide multifaceted search capabilities.
Multi Modal Distance - An Approach to Stemma Generation With Weighting (L18-1)

Copied to clipboard

Challenge: Stemma generation is a task where manuscripts are copied and copied from each other and from M. Existing methods to generate stemma using unweighted token similarity weighting have been used.
Approach: They propose to use a distance model to weight the texts of M1 and M2 to estimate the most likely tree from a series of mapping processes.
Outcome: The proposed method is small in the experimental scenario(s) it is based on psycholinguistically gained distance matrices of letters in three modalities: vision, audition and motorics.
A Corpus of Natural Multimodal Spatial Scene Descriptions (L18-1)

Copied to clipboard

Challenge: Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions.
Approach: They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks.
Outcome: The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces.
The Effects of Unimodal Representation Choices on Multimodal Learning (L18-1)

Copied to clipboard

Challenge: In the real world, multiple modes of information are gathered to create knowledge in a way humans can understand.
Approach: They propose to combine unimodal representations to map multiple modes of information to a single space . they argue that the way they are combined can affect performance and classification metrics .
Outcome: The proposed model can be used to correlate words in a textual description of an object with multimodal representations.
An Evaluation Framework for Multimodal Interaction (L18-1)

Copied to clipboard

Challenge: a framework for evaluating multimodal interactions is presented . it leverages the semantics of language and gesture to assess mutual understanding . consistent evaluation is required to test areas where the system needs improvement .
Approach: They propose a framework for evaluating interactions between human and virtual agent . they use VoxML as a platform to model interactions using natural language and gesture .
Outcome: The proposed framework assesses the level of mutual understanding and ease of communication between human and computer agents in a blocks world scenario.
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)

Copied to clipboard

Challenge: Using the corpus, we study the characteristics of interpreters' work and train machine translation systems.
Approach: They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work.
Outcome: The proposed corpus can be used for teaching interpreters and to train machine translation systems.
Polish Corpus of Annotated Descriptions of Images (L18-1)

Copied to clipboard

Challenge: a new dataset of image descriptions is presented in Polish . the dataset is too small for training a sophisticated language-vision system.
Approach: They propose to use a Polish dataset to analyze image descriptions . the descriptions are morphosyntactically analysed and annotated by human annotators .
Outcome: The proposed model learns about the inter-modal correspondences between language and vision.
Action Verb Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of 390 simple actions is based on multimodal data of 12 humans . the dataset is annotated with orthographic transcriptions of utterances and part-of-speech tags .
Approach: They present a multimodal corpus of 12 humans performing 390 simple actions . they also propose an algorithm for segmenting words into utterances and aligning visual information and speech .
Outcome: The presented dataset includes 390 simple actions performed by 12 humans . it includes transcriptions of utterances, part-of-speech tags, lemmata, and hand touches .
EMO&LY (EMOtion and AnomaLY) : A new corpus for anomaly detection in an audiovisual stream with emotional context. (L18-1)

Copied to clipboard

Challenge: Anomalies in discourse are induced or acted by a machine learning algorithm.
Approach: They propose to use facial and speech video to create a corpus that contains controlled anomalies.
Outcome: The proposed corpus contains controlled anomalies in speech and facial video recordings of subjects.
Development of an Annotated Multimodal Dataset for the Investigation of Classification and Summarisation of Presentations using High-Level Paralinguistic Features (L18-1)

Copied to clipboard

Challenge: Existing summarisation methods take no account of multimodal high-level paralinguistic features which form part of audio-visual presentations.
Approach: They propose to use audiovisual recordings to extract paralinguistic features from audio recordings . they use manual annotations to help users find relevant material .
Outcome: The proposed method can identify the most important or emphasised material within a presentation.
BKTreebank: Building a Vietnamese Dependency Treebank (L18-1)

Copied to clipboard

Challenge: In this paper, we present the building of a dependency treebank for Vietnamese .
Approach: They propose to build a Vietnamese dependency treebank using automatic taggers and automatic tagging.
Outcome: The proposed treebank is a useful resource for Vietnamese language processing.
GeCoTagger: Annotation of German Verb Complements with Conditional Random Fields (L18-1)

Copied to clipboard

Challenge: Complement phrases are essential for constructing well-formed sentences in German.
Approach: They propose an algorithm which can identify and classify complement phrases of any German verb in any written sentence context.
Outcome: The proposed algorithm can identify and classify complement phrases of any German verb in any written sentence context.
AET: Web-based Adjective Exploration Tool for German (L18-1)

Copied to clipboard

Challenge: AET enables research on the modificational behavior of German adjectives and adverbs . currently available online corpus query tools for German do not lend themselves specifically to research on adjectives - e.g., syntactic relationships or morphological properties.
Approach: They propose a web-based corpus query tool that can be used to query German corpus . they extracted modifiers and modifiees from a print media corpus and stored them in a database .
Outcome: The proposed tool can be transferred to other languages and modification phenomena.
ZAP: An Open-Source Multilingual Annotation Projection Framework (L18-1)

Copied to clipboard

Challenge: Existing frameworks for annotation projection in parallel corpora limit reproducibility and comparison of experiments.
Approach: They propose an open-source framework for annotation projection in parallel corpora . framework is Java-based and includes methods for preprocessing corpors, computations and visualization .
Outcome: The proposed framework is designed for ease-of-use with lightweight APIs.
Palmyra: A Platform Independent Dependency Annotation Tool for Morphologically Rich Languages (L18-1)

Copied to clipboard

Challenge: PALMYRA is an annotation tool designed to help with syntactic annotation of morphologically rich languages.
Approach: They present PALMYRA, a platform independent graphical dependency tree visualization and editing software.
Outcome: PALMYRA is an graphical dependency tree visualization and editing software designed to support syntactic annotation of morphologically rich languages.
A Web-based System for Crowd-in-the-Loop Dependency Treebanking (L18-1)

Copied to clipboard

Challenge: Existing treebanks are limited in size, genre, and topic coverage, making manual annotation time-consuming and expensive.
Approach: They propose a web-based interactive tool for editing dependency trees that uses machine learning to accelerate annotation.
Outcome: CROWDTREE is a web-based interactive tool for editing dependency trees . it can train a parsing model during the annotation process and can even be compatible with Mechanical Turk.
Building Universal Dependency Treebanks in Korean (L18-1)

Copied to clipboard

Challenge: Several treebanks were introduced for Korean, all of which comprised annotation of morphemes and phrase structure trees, each following its own set of guidelines.
Approach: They propose to use Korean treebanks as dependency trees and to analyze their performance using morpheme-level annotations.
Outcome: The Korean portion of the Google UD Treebank, the Penn Korean Treebank and the KAIST Treebank are re-tokenized and assessed for errors.
Moving TIGER beyond Sentence-Level (L18-1)

Copied to clipboard

Challenge: TIGER 2.2-doc is a new set of annotations for the German TIger corpus.
Approach: They propose a new set of annotations for the German TIGER corpus . they introduce new document-level annotations: authors and their gender.
Outcome: The new annotations improve the TIGER corpus and its structure and authors and gender.
Spanish HPSG Treebank based on the AnCora Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of HPSG annotated trees for Spanish contains morphosyntactic information, annotations for semantic roles, clitic pronouns and relative clauses.
Approach: They propose to build a Spanish HPSG annotated corpus based on the Spanish corpus AnCora and an HTML format for visualizing the trees in a browser.
Outcome: The proposed corpus contains syntactic and morphological information, semantic roles, clitic pronouns and relative clauses, and has CFG style annotations.
Universal Dependencies for Amharic (L18-1)

Copied to clipboard

Challenge: Amharic is a morphologically rich language with a dependency relation between orthographic words and lexical categories.
Approach: They propose to create an Amharic Dependency Treebank by POS tagging, morphological information and dependency relations.
Outcome: The proposed treebanks are based on 1,096 sentences and are able to parse Amharic.
A Parser for LTAG and Frame Semantics (L18-1)

Copied to clipboard

Challenge: Existing parsers for Lexicalized Tree Adjoining Grammars and frame semantics are difficult to use due to the size of the resources to develop.
Approach: They propose a parser which uses Lexicalized Tree Adjoining Grammars and frame semantics to combine them.
Outcome: The proposed grammars are based on Lexicalized Tree Adjoining Grammars and frame semantics.
Multilingual Dependency Parsing for Low-Resource Languages: Case Studies on North Saami and Komi-Zyrian (L18-1)

Copied to clipboard

Challenge: Developing systems for low-resource languages is a crucial issue for Natural Language Processing (NLP).
Approach: They propose a method for parsing low-resource languages with very small training corpora using multilingual word embeddings and annotated corporata of larger languages.
Outcome: The proposed method improves dependency parsing for low-resource languages with very small training corpora compared to previous work . it also explores whether contemporary contact languages or genetically related languages would be the most fruitful starting point for multilingual parsers.
FonBund: A Library for Combining Cross-lingual Phonological Segment Data (L18-1)

Copied to clipboard

Challenge: Speech and language technology is currently only available for a tiny fraction of the world's languages.
Approach: They propose to map phonetic segments in International Phonetic Alphabet into multiple articulatory feature representations using a free-source library.
Outcome: The proposed library can be easily modified to support new phonological segment inventories.
Voice Builder: A Tool for Building Text-To-Speech Voices (L18-1)

Copied to clipboard

Challenge: a text-to-speech voice building tool is available for low-resourced languages . the tool allows researchers to run voice training experiments and listen to the resulting voice .
Approach: They propose an opensource text-to-speech (TTS) voice building tool that focuses on simplicity, flexibility, and collaboration.
Outcome: The proposed tool can help improve TTS research especially for low-resourced languages . it can be used to run voice training experiments and listen to the resulting synthesized voice .
Sudachi: a Japanese Tokenizer for Business (L18-1)

Copied to clipboard

Challenge: Lack of token unit compatibility is one of the critical problems of Japanese language resources.
Approach: They develop a Japanese tokenizer called Sudachi and its accompanying dictionary . they use multi-granular output and normalization of notation variations to improve tokenization .
Outcome: The proposed tokenizer and dictionary improve tokenization in Japanese for business use.
Chemical Compounds Knowledge Visualization with Natural Language Processing and Linked Data (L18-1)

Copied to clipboard

Challenge: Existing systems for chemical compounds extraction and registration depend on human labor . CAS databases are being created, but information written in other languages is not exploited well .
Approach: They propose a visualization system for chemical compounds extracted from Japanese texts and chemical compound databases represented as Linked Data (LD) system integrates extracted results with existing chemical compound knowledge to provide different views of chemical compounds.
Outcome: The proposed system integrates extraction results with existing chemical compound knowledge to provide different views of chemical compounds.
Using Discourse Information for Education with a Spanish-Chinese Parallel Corpus (L18-1)

Copied to clipboard

Challenge: Discourse information is crucial for many NLP tasks due to the great distance that spans between the two languages.
Approach: They propose to use a Spanish-Chinese parallel corpus with annotated discourse information to serve for bilingual language education.
Outcome: The proposed corpus is composed of 100 Spanish-Chinese parallel texts, and all the discourse markers (DM) have been annotated to form the education source.
A 2nd Longitudinal Corpus for Children’s Writing with Enhanced Output for Specific Spelling Patterns (L18-1)

Copied to clipboard

Challenge: IQB study looks at reading, mathematics and spelling ability across different states.
Approach: They collect three longitudinal corpora of German school children's weekly writing in German and transcribe them into a corpus for research via Linguistic Data Consortium.
Outcome: The corpus of German school children's weekly writing in German was collected and transcribed.
Development of a Mobile Observation Support System for Students: FishWatchr Mini (L18-1)

Copied to clipboard

Challenge: Several video annotation tools have been developed to observe educational activities, but they are not suitable for students' real-time annotation and group reflection.
Approach: They propose a system called FishWatchr Mini which supports students' observation and reflection in the classroom.
Outcome: The proposed system allows students to examine annotation data through reflection, by providing functions such as visualization.
The AnnCor CHILDES Treebank (L18-1)

Copied to clipboard

Challenge: Using the AnnCor CHILDES Treebank, we assign adult grammar syntactic structures to children's utterances.
Approach: They propose a partially manually verified treebank for Dutch CHILDES corpora . they argue that human annotation and automatic checks on this annotation must go hand in hand .
Outcome: The AnnCor CHILDES Treebank is the first partially manually verified treebank for Dutch CHILdes corpora.
BabyCloud, a Technological Platform for Parents and Researchers (L18-1)

Copied to clipboard

Challenge: a platform for capturing, storing and analyzing day-long audio recordings and photos of children's linguistic environments is proposed . the proposed platform connects families and academics, with strong innovation potential for each type of users.
Approach: They propose a platform for capturing, storing and analyzing audio recordings and photos of children's linguistic environments.
Outcome: The proposed platform connects families and academics with strong innovation potential for each type of users.
Infant Word Comprehension-to-Production Index Applied to Investigation of Noun Learning Predominance Using Cross-lingual CDI database (L18-1)

Copied to clipboard

Challenge: Existing theories suggest nouns should predominate verbs in children's word learning .
Approach: They define a measure called the comprehension-to-production index to investigate whether nouns have predominance over verbs in children's word learning.
Outcome: The proposed measure indicates noun predominance in word learning by children . it could provide clues for engineering solutions for teaching words to computers .
Building a TOCFL Learner Corpus for Chinese Grammatical Error Diagnosis (L18-1)

Copied to clipboard

Challenge: Annotated learner corpus is valuable for research in second language acquisition, foreign language teaching, and contrastive interlanguage analysis.
Approach: They construct a TOCFL learner corpus and use it for Chinese grammatical error diagnosis.
Outcome: The constructed corpus is available to the public and will be used for shared tasks on Chinese grammatical error diagnosis.
MIAPARLE: Online training for the discrimination of stress contrasts (L18-1)

Copied to clipboard

Challenge: Second language learners tend to imprint the prosody of their mother language onto the second language (L2) . this can hamper communication between learners and natives, and can also affect the credibility of learners and how they are evaluated by others.
Approach: They propose a tool that focuses on stress perception for speakers whose L1 is a fixed-stress language, such as French.
Outcome: The tool is particularly useful for speakers whose L1 is a fixed-stress language, such as French.
ESCRITO - An NLP-Enhanced Educational Scoring Toolkit (L18-1)

Copied to clipboard

Challenge: Existing implementations are very specific to specific use cases and datasets.
Approach: ESCRITO is a toolkit for scoring student writings using NLP techniques . authors propose teachers and NLP researchers to use APIs for scoring pipelines .
Outcome: ESCRITO is a toolkit for scoring student writings using NLP techniques . it addresses two main user groups: teachers and NLP researchers .
A Leveled Reading Corpus of Modern Standard Arabic (L18-1)

Copied to clipboard

Challenge: Using a reading corpus in Modern Standard Arabic, we explore the lexical coverage of textbooks and unabridged works of fiction.
Approach: They propose to use textbooks from the United Arab Emirates curriculum and a reading corpus in Modern Standard Arabic to enrich the sparse collection of resources available for educational applications.
Outcome: The corpus spans all 12 grades and contains 129 unabridged works of fiction spanning grades 1-12 . lexical coverage is compared to other genres, and the results show that the two sub-corpora are similar to each other to measure their genres.
Developing New Linguistic Resources and Tools for the Galician Language (L18-1)

Copied to clipboard

Challenge: Existing resources and tools for the Galician language are lacking for other less-resourced languages, such as statistical tools for lemmatization and Named Entity Recognition.
Approach: They propose to develop a manually revised corpus for POS tagging and lemmatization, and a new manually annotated corpus to train existing statistical tools for the Galician language.
Outcome: The proposed resources include a new corpus for POS tagging and lemmatization, and a manually annotated corpus to handle Named Entity recognition.
Modeling Northern Haida Verb Morphology (L18-1)

Copied to clipboard

Challenge: a computational model of the verbal morphology of Northern Haida is being developed . the model is capable of handling complex affixation patterns and morphophonological alternations .
Approach: They propose a computational model of the verbal morphology of Northern Haida based on finite state machines with a focus on verbs.
Outcome: The proposed model can handle complex affixation patterns and morphophonological alternations in the native language.
Low-resource Post Processing of Noisy OCR Output for Historical Corpus Digitisation (L18-1)

Copied to clipboard

Challenge: 7.6% of the words in the original OCR text contain an error; fully manual correction would take thousands of hours due to the size of the corpus.
Approach: They propose a post-processing system to efficiently correct OCR errors in a 2.7 million word Faroese corpus.
Outcome: The proposed method reduces the word error rate to 1.3% with around 65 hours of human annotator work.
Introducing the CLARIN Knowledge Centre for Linguistic Diversity and Language Documentation (L18-1)

Copied to clipboard

Challenge: Knowledge Centres comprise physical institutions with particular expertise in certain areas and are committed to providing their expertise in the form of reliable knowledge-sharing services.
Approach: They propose to build a Knowledge Sharing Infrastructure (KSI) to ensure existing knowledge and expertise is easily available for the CLARIN community and for the humanities research communities for which CLARINS is being developed.
Outcome: The CLARIN Knowledge Centre for Linguistic Diversity and Language Documentation (CKLD) is a virtual distributed centre comprising institutions at the Universities of London, Cologne and Hamburg.
Low Resource Methods for Medieval Document Sections Analysis (L18-1)

Copied to clipboard

Challenge: a small but unique collection of medieval Latin charters has been digitized and annotated . sections of these documents were manually annotating for deeper analysis of the structure of issued charters .
Approach: They propose to use a digitized collection of medieval Latin charters to analyze their structure . they propose to provide manual annotations of common structure of charters and general methods for automatic detection .
Outcome: The proposed methods can be applied to documents with partially repetitive character . they provide annotated sections of charters and tools for automatic recognition .
SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)

Copied to clipboard

Challenge: Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences.
Approach: They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats.
Outcome: The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral.
Universal Dependencies for Ainu (L18-1)

Copied to clipboard

Challenge: a task is underway to create a dependency tree bank for the Ainu language in the scheme of Universal Dependencies (UD).
Approach: They propose to create a dependency tree bank for the Ainu language in the scheme of Universal Dependencies (UD) their mini-lexicon is encoded under the W3C OntoLex specification with UD and UniMorph features with the system-friendly JSON-LD format and is bearable to future extensions.
Outcome: The proposed tree bank contains 10,000 word tokens and is small enough to be used as a base annotation for the next step.
Signbank: Software to Support Web Based Dictionaries of Sign Language (L18-1)

Copied to clipboard

Challenge: Auslan Signbank is an on-line dictionary for Australian Sign Language (Auslan) it was originally built to support the Auslan signbank web dictionary, but was re-implemented using Microsoft SQL Server.
Approach: This paper describes the overall architecture of the Auslan Signbank system and its representation of lexical entries and associated entities.
Outcome: The current version of Auslan Signbank is an open-source re-implementation of the original website, with features added to allow updates to the database by researchers.
J-MeDic: A Japanese Disease Name Dictionary based on Real Clinical Usage (L18-1)

Copied to clipboard

Challenge: a study finds that medical texts are written mostly in natural language, requiring NLP for medical texts.
Approach: They develop a Japanese disease name dictionary to fill the gap between medical names and clinical words . they allocated the standard disease code to the names using manual, semi-automatic or automatic methods .
Outcome: The Japanese disease name dictionary fills the gap between standard medical names and real clinical words . the study found that 55.3% of the names covered by the dictionary were SDNs .
Building a List of Synonymous Words and Phrases of Japanese Compound Verbs (L18-1)

Copied to clipboard

Challenge: Japanese is rich in compound verbs consisting of two verbs joined together.
Approach: They built a database of Japanese "Verb + Verb" compounds semi-automatically . they extracted Japanese compound verbs from corpus and found suitable clusters .
Outcome: The proposed database extracts synonymous expressions of Japanese compound verbs from corpus . it then links the results to the "Compound Verb Lexicon"
Evaluating EcoLexiCAT: a Terminology-Enhanced CAT Tool (L18-1)

Copied to clipboard

Challenge: EcoLexiCAT is a web-based tool for terminology-enhanced translation of environmental texts . most terminological modules in CAT tools do not go beyond a simple glossary of source and target terms .
Approach: They propose to integrate terminology-enhanced translation into a web-based tool . EcoLexiCAT is a terminology-enriched CAT tool for the English-Spanish-English translation .
Outcome: The EcoLexiCAT tool is an open-source version of the CAT tool MateCat . it enriches a source text with information from a multimodal and multilingual terminological knowledge base on the environment .
A Danish FrameNet Lexicon and an Annotated Corpus Used for Training and Evaluating a Semantic Frame Classifier (L18-1)

Copied to clipboard

Challenge: a Danish FrameNet is a lexicon based on the Danish Thesaurus . it is significantly faster than building a new one from scratch .
Approach: They propose a way to efficiently compile a Danish FrameNet based on the Danish Thesaurus . they present the corresponding corpus annotations of frames and roles and show how this can be used for a semantic frame classifier .
Outcome: The proposed approach is faster than building a lexicon from scratch.
SLIDE - a Sentiment Lexicon of Common Idioms (L18-1)

Copied to clipboard

Challenge: Compositional solutions for phrase sentiment are not able to handle idioms because their sentiment is not derived from the sentiment of the individual words.
Approach: They propose a crowdsourcing approach for collecting sentiment annotations of idiomatic expressions using crowdsourcing.
Outcome: The proposed approach is able to capture sentiment strength and ambiguity in idiomatic expressions using crowdsourcing.
PronouncUR: An Urdu Pronunciation Lexicon Generator (L18-1)

Copied to clipboard

Challenge: acoustic modeling, large text data and a pronunciation lexicon are the bottlenecks for speech recognition systems for resource scarce languages.
Approach: They propose a grapheme-to-phoneme conversion tool that generates a pronunciation lexicon from a list of Urdu words.
Outcome: The proposed tool predicts pronunciation of words using a LSTM-based model trained on a handcrafted expert lexicon of around 39,000 words and shows an accuracy of 64% upon internal evaluation.
SimLex-999 for Polish (L18-1)

Copied to clipboard

Challenge: a linguistic evaluation tool is needed for distributional semantics tasks.
Approach: They extend the Polish version of SimLex-999 to include measurement of similarity and relatedness.
Outcome: The proposed model is compared with distributional semantics models for other languages.
Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese (L18-1)

Copied to clipboard

Challenge: A distributional semantics model is instrumental to improve the performance of many applications and processing tasks for any language.
Approach: They propose to develop an advanced distributional model for Portuguese with the largest vocabulary and best evaluation scores published so far.
Outcome: The proposed model has the largest vocabulary and the best evaluation scores published so far.
Teanga: A Linked Data based platform for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: Using linked data, we can use many NLP services from a single interface . integrating components within a development model is endemic to software development .
Approach: They propose a linked data based platform for natural language processing that uses linked data to define the types of services input and output.
Outcome: The proposed platform is easy to install and run, easy to use and able to run multiple NLP tasks from one interface.
Automatic and Manual Web Annotations in an Infrastructure to handle Fake News and other Online Media Phenomena (L18-1)

Copied to clipboard

Challenge: a growing number of people consume news online, but there are different types of "fake news" many online news outlets use the same journalistic principles that have been in use for newspapers for decades, especially factchecking.
Approach: They propose a metadata scheme to enable users to handle "fake news" they also propose 'filter bubble' effect and abuse language .
Outcome: The proposed metadata scheme enables standardisation of these phenomena in online media.
The LODeXporter: Flexible Generation of Linked Open Data Triples from NLP Frameworks for Automatic Knowledge Base Construction (L18-1)

Copied to clipboard

Challenge: Linked Open Data (LOD) principles are used to export natural language processing (NLP) results to graph-based knowledge base.
Approach: They propose a method for exporting NLP results to a graph-based knowledge base using Linked Open Data principles.
Outcome: The proposed method is available as an open source component for the GATE framework and is available on GitHub.
LiDo RDF: From a Relational Database to a Linked Data Graph of Linguistic Terms and Bibliographic Data (L18-1)

Copied to clipboard

Challenge: linguists and researchers benefit from the data by looking it up on the Web . a new approach allows the direct use and reuse of the data for scientific research and machine processing .
Approach: They propose to convert LiDo TBD database into Linked Data graph using Semantic Web . goal is to enable direct use and reuse of data for scientific research community .
Outcome: The proposed dataset is based on the framework developed by linguist Dr. Christian Lehmann 40 years ago and is available on the LiDo website since 2006.
Towards a Linked Open Data Edition of Sumerian Corpora (L18-1)

Copied to clipboard

Challenge: Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity .
Approach: They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts .
Outcome: The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts .
A Bird’s-eye View of Language Processing Projects at the Romanian Academy (L18-1)

Copied to clipboard

Challenge: a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects.
Approach: a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications .
Outcome: a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications .
PMKI: an European Commission action for the interoperability, maintainability and sustainability of Language Resources (L18-1)

Copied to clipboard

Challenge: Public Multilingual Knowledge Management Infrastructure (PMKI) is a project launched by the European Commission to promote the digital single market in the EU.
Approach: The paper presents the Public Multilingual Knowledge Management Infrastructure (PMKI) action launched by the European Commission to promote the Digital Single Market in the European Union.
Outcome: The proposed public multilingual knowledge management infrastructure (PMKI) is a pilot project launched by the European Commission to promote the digital single market in the European Union.
The Abkhaz National Corpus (L18-1)

Copied to clipboard

Challenge: Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language .
Approach: They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus .
Outcome: The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives.
Collecting Language Resources from Public Administrations in the Nordic and Baltic Countries (L18-1)

Copied to clipboard

Challenge: Several large-scale projects and initiatives have been undertaken in this century to collect language resources and create LR repositories and infrastructures on a pan-European scale.
Approach: They present the work of Tilde on collecting language resources from government institutions and other public administrations in the Nordic and Baltic countries.
Outcome: The results of the European Language Resources Coordination (ELRC) action in the Nordic and Baltic countries are presented.
LIdioms: A Multilingual Linked Idioms Data Set (L18-1)

Copied to clipboard

Challenge: Recent studies have focused on linguistic data sets that are bilingual on the Linguistic Linked Open Data (LLOD) 1 .
Approach: They describe a multilingual RDF representation of idioms currently containing five languages . they use a model to structure the data and a method to link the data to well-known multilingual data sets such as BabelNet.
Outcome: The proposed model complies with best practices according to Linguistic Linked Open Data Community.
Annotating Modality Expressions and Event Factuality for a Japanese Chess Commentary Corpus (L18-1)

Copied to clipboard

Challenge: In recent years, there has been a surge of interest in the natural language processing related to the real world . shogi commentaries are an interesting testbed for these tasks, but can be grounded in the game tree .
Approach: They propose to augment shogi commentaries with game states to generate a game commentary generator.
Outcome: The proposed system can be used to ground symbols and events with factuality . it can be compared with other systems to find out if a commentator is a human .
Annotating Chinese Light Verb Constructions according to PARSEME guidelines (L18-1)

Copied to clipboard

Challenge: Using existing resources, we can annotate Chinese multiword expressions using PARSEME guidelines.
Approach: They propose to use an existing resource containing Chinese light verbs to make an annotation of a Chinese UD treebank in two steps.
Outcome: The proposed annotations are based on an existing treebank containing Chinese light verbs and are consistent with the proposed guidelines.
Using English Baits to Catch Serbian Multi-Word Terminology (L18-1)

Copied to clipboard

Challenge: a new method for bilingual terminology extraction is proposed for a source language and a target language.
Approach: They propose to use a bilingual terminology extraction approach for a source language and a target language to extract the terminology for sri lanka.
Outcome: The proposed method extracts terminology for a source language and a target language from it.
Construction of Large-scale English Verbal Multiword Expression Annotated Corpus (L18-1)

Copied to clipboard

Challenge: In this paper, we focus on verbal MWEs, whose accurate recognition is challenging because they could be discontinuous.
Approach: They conduct large-scale annotations of VMWEs on the Wall Street Journal portion of Ontonotes . they first construct a VMwe dictionary based on the english-language Wiktionary .
Outcome: The proposed resource annotates 7,833 VMWE instances belonging to various categories . the authors hope the results will help to develop models for MWE recognition and dependency parsing .
Konbitzul: an MWE-specific database for Spanish-Basque (L18-1)

Copied to clipboard

Challenge: Multiword Expressions (MWEs) are combinations of words which express a single meaning.
Approach: They present an online database of verb+noun MWEs in Spanish and Basque.
Outcome: The proposed database helps to identify occurrences of MWEs in multiple morphosyntactic variants and improve translation quality in rule-based MT.
A Multilingual Test Collection for the Semantic Search of Entity Categories (L18-1)

Copied to clipboard

Challenge: Despite the high popularity of entity search, entity categories have not received equal attention.
Approach: They propose to make public a multilingual test collection comprehending English, Portuguese and German to meet the demands of the entity search community.
Outcome: The proposed test collection comprehends English, Portuguese and German and provides comparative baselines and an analysis of the results.
Towards the Inference of Semantic Relations in Complex Nominals: a Pilot Study (L18-1)

Copied to clipboard

Challenge: Complex nominals (CNs) show similar external forms but encode different semantic relations because of noun packing.
Approach: They propose to use paraphrases to convey conceptual content of english two-term CNs in the domain of environmental science to disambiguate the semantic relation between constituents of CN.
Outcome: The proposed method disambiguates the semantic relation between constituents of the CN and infers the semantic relations in these multi-word terms.
Generation of a Spanish Artificial Collocation Error Corpus (L18-1)

Copied to clipboard

Challenge: collocations are combinations of two elements where one (the base) is freely chosen, despite the limitations of the other (collocate) current tools for collocation error detection and correction focus on collocation validation and identification of miscollocations .
Approach: They propose an algorithm for automatic generation of an artificial collocation error corpus of american English learners of Spanish that includes 17 different types of collocation errors.
Outcome: The proposed algorithm can detect and classify collocation errors in learners' writings . collocation error detection and correction has not received the attention it deserves .
Improving a Neural-based Tagger for Multiword Expressions Identification (L18-1)

Copied to clipboard

Challenge: MUMULS tagger for automatic detection of verbal multiword expressions is based on neural networks . character-level embeddings can improve the performance, reducing out-of-vocabulary rate . multiword Expressions are viewed by computational linguists as a "pain in the neck of NLP"
Approach: They propose to improve MUMULS, a tagger for automatic detection of verbal multiword expressions.
Outcome: The proposed tagger performed better on Czech language than the previous taggers.
Designing a Russian Idiom-Annotated Corpus (L18-1)

Copied to clipboard

Challenge: a pilot experiment using the idiom-annotated corpus of Russian is described . corpora that could be used for training idiomatic classifiers are scarce, especially if one turns to other languages.
Approach: They describe the development of an idiom-annotated corpus of Russian . the corpus is compiled from freely available online resources .
Outcome: The proposed corpus is based on an online corpus of Russian texts . it is available for research purposes and can be used for linguistic studies and pedagogy .
DeepTC – An Extension of DKPro Text Classification for Fostering Reproducibility of Deep Learning Experiments (L18-1)

Copied to clipboard

Challenge: a current state of DKPro TC does not allow integration of deep learning . we integrate Keras, DyNet, and DeepLearning4J as proof-of-concept .
Approach: They propose a deep learning extension for the multi-purpose text classification framework DKPro Text Classification.
Outcome: The proposed extension improves readability and reduces redundant source code.
Improving Hate Speech Detection with Deep Learning Ensembles (L18-1)

Copied to clipboard

Challenge: censorship is a potential risk when addressing these issues with automated text classification methods.
Approach: They propose to use a neural network-based ensemble method to better classify hate speech using a publicly available embedding model and a popular sentiment dataset.
Outcome: The proposed method improves by 5 points on a hate speech corpus from Twitter and a popular sentiment dataset.
Distributional Term Set Expansion (L18-1)

Copied to clipboard

Challenge: Iterative term set expansion methods for distributional semantic models are used to label terms belonging to a sought after term set.
Approach: They compare iterative term set expansion methods for distributional semantic models to the Simple Margin method, an active learning approach to classification using Support Vector Machines.
Outcome: The proposed methods outperform centrality and classification based methods for distributional semantic models over five different term sets.
Can Domain Adaptation be Handled as Analogies? (L18-1)

Copied to clipboard

Challenge: Aspect identification in user generated texts might suffer degradation when changing to other domains than the one used for training.
Approach: They propose to use offset method to handle domain shifts when there is no available labeled data in a new target domain for an aspect classifier to be retrained.
Outcome: The proposed method found analogues in the new domain for the initial features but did not deliver the expected results.
Author Profiling from Facebook Corpora (L18-1)

Copied to clipboard

Challenge: Existing studies on author profiling focus on age and gender, and use only English text.
Approach: They propose to model author profiling from a Brazilian Portuguese corpus using standard gender and age prediction tasks and two less-known alternatives: predicting an author's degree of religiosity and IT background status.
Outcome: The proposed tasks are based on a Brazilian Portuguese corpus and are compared with other languages and tasks.
Semantic Relatedness of Wikipedia Concepts – Benchmark Data and a Working Solution (L18-1)

Copied to clipboard

Challenge: Existing methods to measure relatedness between Wikipedia concepts are lacking.
Approach: They propose a new type of concept relatedness dataset, WORD, which is annotated by a human . they use this dataset to assess relatedness between Wikipedia concepts using supervised methods.
Outcome: The proposed dataset outperforms existing methods for measuring relatedness between Wikipedia concepts.
Experiments with Convolutional Neural Networks for Multi-Label Authorship Attribution (L18-1)

Copied to clipboard

Challenge: Existing methods for authorship attribution tasks are difficult, but they are effective.
Approach: They propose a CNN that averaging author probability distributions at sentence level for longer documents and treating smaller documents as sentences adapts to single-label datasets and various document sizes.
Outcome: The proposed method outperforms state-of-the-art models on a single-label AA benchmark dataset.
A Fast and Accurate Vietnamese Word Segmenter (L18-1)

Copied to clipboard

Challenge: Experimental results show that our approach outperforms previous state-of-the-art approaches in terms of accuracy and performance speed.
Approach: They propose a method where rules are stored in an exception structure and new rules are only added to correct segmentation errors.
Outcome: The proposed approach outperforms existing methods on Vietnamese treebank benchmarks.
Finite-state morphological analysis for Gagauz (L18-1)

Copied to clipboard

Challenge: a finite-state approach to morphological analysis and generation of Gagauz is used . the model has a reasonable coverage over a range of freely-available corpora .
Approach: They propose a finite-state approach to morphological analysis and generation of Gagauz . they explicitly handle orthographic errors and variance, in addition to loan words .
Outcome: The proposed approach has a reasonable coverage over a range of freely-available corpora.
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)

Copied to clipboard

Challenge: a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part.
Approach: They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers.
Outcome: The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets .
Morphology Injection for English-Malayalam Statistical Machine Translation (L18-1)

Copied to clipboard

Challenge: Statistical Machine Translation fails to handle the rich morphology when translating into morphologically rich language.
Approach: They propose a method to generate unseen morphological forms from the parallel corpus . they propose morphology injection method to enrich the corpus with generated morphologies .
Outcome: The proposed method improves the quality of the translation in English-Malayalam.
The Morpho-syntactic Annotation of Animacy for a Dependency Parser (L18-1)

Copied to clipboard

Challenge: Animacy is a feature found in nouns such as 'gender', 'number' and 'case' that improves parser accuracy.
Approach: They propose an annotation scheme and parser results for the animacy feature in Russian and Arabic, morphologically rich languages, using the universal dependency framework.
Outcome: The proposed scheme and parser improve on the animacy feature in Russian and Arabic, and the results show that the feature is more accurate than other features found in nouns, namely, 'gender', , and 'number'
MADARi: A Web Interface for Joint Arabic Morphological Annotation and Spelling Correction (L18-1)

Copied to clipboard

Challenge: Standard Arabic morphology is rich, but Arabic dialects introduce more complexity.
Approach: They propose a joint morphological annotation and spelling correction system for Arabic texts . they propose morphology tools that can be used to help with productivity .
Outcome: The proposed system is based on a standard and dialectal Arabic text.
A Morphological Analyzer for St. Lawrence Island / Central Siberian Yupik (L18-1)

Copied to clipboard

Challenge: St. Lawrence Island / Central Siberian Yupik is an endangered language . it exhibits pervasive agglutinative and polysynthetic properties .
Approach: They propose to implement a finite-state morphological analyzer for the endangered language . it cyclically interweaves morphology and phonology to account for the language's intricate morphophonological system.
Outcome: The proposed method cyclically interweaves morphology and phonology to account for the language's intricate morphophonological system.
Universal Morphologies for the Caucasus region (L18-1)

Copied to clipboard

Challenge: Caucasus region is famed for its rich and diverse arrays of languages and language families . authors describe efforts to improve the coverage of Universal Morphologies for languages of the region .
Approach: They propose to improve the coverage of Universal Morphologies for Caucasus languages . they propose to complement the Universal Dependencies which focus on morphosyntax and syntax.
Outcome: The proposed framework improves the coverage of languages of the Caucasus region . the proposed framework criticizes the UniMorph TSV format for its limited expressiveness .
EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text (L18-1)

Copied to clipboard

Challenge: Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters.
Approach: They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators.
Outcome: The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators.
Complex and Precise Movie and Book Annotations in French Language for Aspect Based Sentiment Analysis (L18-1)

Copied to clipboard

Challenge: Aspect Based Sentiment Analysis (ABSA) aims at collecting detailed opinion information according to products and their features.
Approach: They propose to use linguistics tools to enhance text classification with aspect-based sentiment analysis.
Outcome: The proposed method is based on two French online reviews datasets.
Lingmotif-lex: a Wide-coverage, State-of-the-art Lexicon for Sentiment Analysis (L18-1)

Copied to clipboard

Challenge: Sentiment Analysis is a subtask of Natural Language Processing.
Approach: They propose a new, domain-neutral lexicon for sentiment analysis in English . they test it on two publicly available sentiment analysis datasets .
Outcome: The proposed lexicon performs better than existing sentiment lexiconics on two publicly available datasets.
A Japanese Corpus for Analyzing Customer Loyalty Information (L18-1)

Copied to clipboard

Challenge: a corpus of customer loyalty information is used to analyze customer's voice . a variety of studies have focused on analyzing attitudes, opinions, sentiments of text data .
Approach: They present a corpus for analyzing customer loyalty information . they use voice of customer to capture customer's behaviors, needs and feedbacks .
Outcome: The proposed corpus analyzes customer loyalty by analyzing their voice . the study is based on a corpus of customer loyalty information .
FooTweets: A Bilingual Parallel Corpus of World Cup Tweets (L18-1)

Copied to clipboard

Challenge: a new study analyzes the nature of twitter data and compares it with other social networking websites.
Approach: They develop a parallel corpus of tweets for an English-German pair that can be translated into German using a machine translation tool.
Outcome: The proposed method can be used to translate tweets from English to German using a parallel corpus of 4, 000 tweets.
The SSIX Corpora: Three Gold Standard Corpora for Sentiment Analysis in English, Spanish and German Financial Microblogs (L18-1)

Copied to clipboard

Challenge: SSIX corpora provide annotated data for supervised learning methods . polarity annotation is performed on two financial microblog platforms .
Approach: They propose three SSIX corpora for sentiment analysis which provide annotated data for supervised learning methods.
Outcome: The proposed corpora are in English, German and Spanish.
Sarcasm Target Identification: Dataset and An Introductory Approach (L18-1)

Copied to clipboard

Challenge: Past work on sarcasm detection has focused on identifying the sarcasm target of ridicule in a sarkastic text.
Approach: They propose a task of extracting the sarcastic target of ridicule from a sarcastical text using a manually annotated dataset and an automatic approach.
Outcome: The proposed approach establishes the viability of sarcasm target identification and will serve as a baseline for future work.
Annotating Opinions and Opinion Targets in Student Course Feedback (L18-1)

Copied to clipboard

Challenge: a student feedback corpus is a novel resource for opinion target extraction and sentiment analysis.
Approach: They propose to annotate student feedback corpus with an opinion target extraction method and an annotation scheme for sentiment analysis.
Outcome: The proposed corpus summarises student feedback on undergraduate courses . the method is difficult, and the results are presented in a tee .
Generating a Gold Standard for a Swedish Sentiment Lexicon (L18-1)

Copied to clipboard

Challenge: Existing sentiment lexicons are compiled by (machine) translation from English resources, obscuring language-specific characteristics of sentiment-loaded vocabulary.
Approach: They propose a gold standard for sentiment annotation of Swedish terms using the SALDO lexicon and the Gigaword corpus.
Outcome: The proposed model is based on the free SALDO lexicon and the Gigaword corpus and is compared with existing models using human annotations.
WordKit: a Python Package for Orthographic and Phonological Featurization (L18-1)

Copied to clipboard

Challenge: wordkit is a python package that allows users to switch between feature sets and featurizers with a uniform API . wordkit integrates orthographic and phonological featurizers in a single package .
Approach: They present a python package which allows users to switch between feature sets and featurizers with a uniform API.
Outcome: The proposed package is compatible with scikit-learn and extensible . it allows users to switch between feature sets and featurizers with a uniform API .
Pronunciation Variants and ASR of Colloquial Speech: A Case Study on Czech (L18-1)

Copied to clipboard

Challenge: a standard speech recognition system uses a pronunciation component that maps tokens in the transcripts to their phonetic representations.
Approach: They propose to use a pronunciation dictionary to map tokens in speech transcripts to phonetic representations.
Outcome: The proposed pronunciation dictionary performs better than a standard rule-based pronunciation component.
Epitran: Precision G2P for Many Languages (L18-1)

Copied to clipboard

Challenge: Epitran is a multilingual, multi-back-end system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under an MIT license .
Approach: Epitran is a multilingual back-end system for grapheme-to-phoneme transduction . it takes word tokens in the orthography of a language and outputs a phonemic representation . Epitran's efficacy has been demonstrated in multiple research projects .
Outcome: Epitran is a multilingual, multi-backend system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under MIT license .
A Multilingual Approach to Question Classification (L18-1)

Copied to clipboard

Challenge: Existing work on questions has focused on understanding the structure of questions per se . a few approaches explicitly focus on information-seeking questions, but this work is either based on big data or crowdsourcing.
Approach: They propose a dependency-parsed, parallel multilingual corpus of information-seeking and non-information-seeing questions . they employ a linguistically motivated rule-based system that uses linguistic cues from one language to help classify questions across other languages.
Outcome: The proposed system correctly classifies questions in 79% of cases, compared to other systems.
Dataset for the First Evaluation on Chinese Machine Reading Comprehension (L18-1)

Copied to clipboard

Challenge: Existing reading comprehension datasets are mostly in English .
Approach: They propose a Chinese reading comprehension dataset to add diversity to existing reading comprehension data . proposed dataset contains cloze-style reading comprehension and user query reading comprehension .
Outcome: The proposed dataset is based on a Chinese reading comprehension dataset . it includes two types of cloze-style and user query reading comprehension . the proposed dataset hosted the 1st Evaluation on Chinese Machine Reading Comprehension (CMRC-2017)
A Multi-Domain Framework for Textual Similarity. A Case Study on Question-to-Question and Question-Answering Similarity Tasks (L18-1)

Copied to clipboard

Challenge: Community Question Answering websites are becoming popular and useful source of information for users.
Approach: They propose to use community question answering forum to detect similar questions . they use question-answering similarity task to provide correct answers .
Outcome: The proposed framework provides the first framework on the evaluation of similar questions and question-answering detection on a multi-domain corpora.
WorldTree: A Corpus of Explanation Graphs for Elementary Science Questions supporting Multi-hop Inference (L18-1)

Copied to clipboard

Challenge: Existing methods of automated inference do not provide enough gold explanations to train models . standardized science exams are a challenge task for question answering .
Approach: They propose to manually construct a corpus of explanations for standardized science exams . they also provide an explanation-centered tablestore that contains the knowledge to construct these explanations .
Outcome: The proposed model provides detailed explanations for standardized science exams . the authors show that the proposed model can be trained on the basis of gold explanations .
Analysis of Implicit Conditions in Database Search Dialogues (L18-1)

Copied to clipboard

Challenge: Annotators annotated 50 database search dialogues with database field tags . 10% of the utterances included non-database-field information, authors say .
Approach: They propose to annotate database search dialogues on real estate and analyse their utterances for database queries.
Outcome: The proposed method can extract the implicit conditions from user utterances and construct queries.
An Information-Providing Closed-Domain Human-Agent Interaction Corpus (L18-1)

Copied to clipboard

Challenge: a human-agent interaction corpus is a corpus of conversations between a user and an embodied conversational agent operated by a wizard of oz . data collected to create a 'corpus' with unexpected situations, such as misunderstandings, false information, and interruptions.
Approach: They propose a public corpus for Human-Agent Interaction where the agent is controlled by a Wizard of Oz.
Outcome: The proposed corpus is based on 15 conversations between users and a wizard of Oz agent . the data are used to create a corpus with unexpected situations, such as misunderstandings, false information, and interruptions.
Augmenting Image Question Answering Dataset by Exploiting Image Captions (L18-1)

Copied to clipboard

Challenge: Image question answering requires large amounts of human-annotated data to achieve optimal performance.
Approach: They propose a framework to augment training data by generating additional examples from unannotated pairs of an image and captions.
Outcome: The proposed framework augments training data by generating additional examples from unannotated pairs of an image and captions.
Semi-supervised Training Data Generation for Multilingual Question Answering (L18-1)

Copied to clipboard

Challenge: Existing datasets for question answering (QA) tasks mostly support only English . however, existing resources for these tasks are labor intensive .
Approach: They propose to combine Korean QA datasets with machine-translated English resources to build seed resources.
Outcome: The proposed approach leads to 71.50 F1 on Korean QA (comparable to 77.3 F1)
PhotoshopQuiA: A Corpus of Non-Factoid Questions and Answers for Why-Question Answering (L18-1)

Copied to clipboard

Challenge: Community Question Answering web sites are used for non-factoid question answering . however, there is a scarcity of available datasets for this task . cnn.com's john m. sutter is releasing a dataset for why-QA .
Approach: They propose a dataset of 2,854 why-question and answer(s) pairs related to Adobe Photoshop usage from five CQA web sites.
Outcome: The new dataset is the first English dataset for Why-QA that focuses on a product . it can be used to build Why-Q systems, evaluate approaches and develop new models .
BioRead: A New Dataset for Biomedical Reading Comprehension (L18-1)

Copied to clipboard

Challenge: BioRead is a publicly available cloze-style biomedical machine reading comprehension (MRC) dataset with 16.4 million passage-question instances.
Approach: They propose to build a cloze-style biomedical machine reading comprehension (MRC) dataset with 16.4 million passage-question instances.
Outcome: The proposed method outperforms baselines on bioReadLite and bioASQ, and is currently the best on BioReadLite.
MMQA: A Multi-domain Multi-lingual Question-Answering Framework for English and Hindi (L18-1)

Copied to clipboard

Challenge: Existing work on multi-domain, multi-lingual question answering is limited to the same language.
Approach: They curate 500 articles in six different domains from the web and create question-answer pairs . they develop a deep learning based model for classifying an input question into coarse and finer categories .
Outcome: The proposed model accuracies 90.12% and 80.30% for coarse and finer classes . the proposed model is the first attempt to create multi-domain, multi-lingual question answering evaluation involving English and Hindi.
The First 100 Days: A Corpus Of Political Agendas on Twitter (L18-1)

Copied to clipboard

Challenge: The first 100 days corpus is a curated corpus of the first 100 of the president and senators . political communication has changed dramatically over recent years .
Approach: They analyze the first 100 days of the president and the senators to see differences in their language usage.
Outcome: The corpus analyzes the first 100 days of the president and the senators to see the differences in their language usage.
Medical Sentiment Analysis using Social Media: Towards building a Patient Assisted System (L18-1)

Copied to clipboard

Challenge: a study conducted by the pew Internet & American Life Project 1 shows that almost 80 percent of Internet users have explored health-related topic online.
Approach: They propose to crawl medical forums with opinions about medical condition self narrated by users.
Outcome: The proposed system is based on opinions about medical condition self-narrated by users on medical forums.
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
A Large Multilingual and Multi-domain Dataset for Recommender Systems (L18-1)

Copied to clipboard

Challenge: Existing algorithms for recommending items are limited and focused on specific domains.
Approach: They propose a multi-domain interests dataset to train and test Recommender Systems . the english dataset includes an average of 90 preferences per user on music, books, movies, celebrities, sport, politics .
Outcome: The proposed method exploits popular services such as Spotify, Goodreads and others to extract preferences from Twitter messages in Italian and English.
RtGender: A Corpus for Studying Differential Responses to Gender (L18-1)

Copied to clipboard

Challenge: Prior work on linguistic gender difference and communications about gender has focused on language about or portraying persons of a particular gender.
Approach: They present a multi-genre corpus of 25M comments from five socially and topically diverse sources tagged for the gender of the addressee and 30k annotations for sentiment and relevance of these responses.
Outcome: The proposed dataset shows that responses to women are more emotive and about the speaker as an individual (rather than about the content being responded to).
A Neural Network Model for Part-Of-Speech Tagging of Social Media Texts (L18-1)

Copied to clipboard

Challenge: Recent approaches based on end-to-end Deep Neural Networks (DNNs) have shown promising results for Natural Language Processing (NLP).
Approach: They propose a neural network model for part-of-speech (POS) tagging of User-Generated Content (UGC) such as Twitter, Facebook and Web forums that uses character and word representations.
Outcome: The proposed model is end-to-end and uses character and word representations . it is compared with existing models on social media in English, german, french, italian and spanish .
Utilizing Large Twitter Corpora to Create Sentiment Lexica (L18-1)

Copied to clipboard

Challenge: Existing sentiment analysis systems only use word unigrams and bigrams, but lexicons using sentiment lexica are effective.
Approach: They describe an automatic Twitter sentiment lexicon creator and a lexico-based sentiment analysis system.
Outcome: The proposed system outperforms a manually annotated system in a comparison experiment.
The Nautilus Speaker Characterization Corpus: Speech Recordings and Labels of Speaker Characteristics and Voice Descriptions (L18-1)

Copied to clipboard

Challenge: The Nautilus Speaker Characterization corpus is a conversational microphone speech recording corpus from 300 speakers.
Approach: They present a speaker characterization corpus from 300 german speakers . they use four scripted and four semi-spontaneous dialogs to simulate telephone calls .
Outcome: The speaker characterization corpus is presented in the acoustically-isolated room Nautilus . it comprises conversational microphone speech recordings from 300 speakers . the data will be made freely available to the scientific community .
Evaluation of Automatic Formant Trackers (L18-1)

Copied to clipboard

Challenge: Formant trackers are widely used by speech scientists and speech engineers.
Approach: They propose to use four open source formant trackers to evaluate the quality of speech recognition algorithms on the same American English data set.
Outcome: The proposed formant trackers outperform LPC-based and Deep Learning on the American English data set VTR-TIMIT.
Design and Development of Speech Corpora for Air Traffic Control Training (L18-1)

Copied to clipboard

Challenge: The current state-of-the-art training procedures involve retired pilots that train as virtual plane pilots and process the spoken prompts to form that can be entered into software that simulates the plane movement on the radar screen.
Approach: They describe the process of creating domain-specific speech corpora containing air traffic control (ATC) communication prompts.
Outcome: The proposed system could be used for training air traffic controllers in the Czech Republic.
A First South African Corpus of Multilingual Code-switched Soap Opera Speech (L18-1)

Copied to clipboard

Challenge: a corpus of code-switched speech from soap operas is compiled from soaps . the corpus contains 14.3 hours of annotated and segmented speech .
Approach: They propose a speech corpus containing multilingual code-switching from soap operas . the corpus contains English, isiZulu, isisXhosa, Setswana and Sesotho speech .
Outcome: The corpus contains 14.3 hours of annotated and segmented speech from soap operas . the speech rate is 1.22 to 1.83 times higher than prompted speech in the same languages .
A Web Service for Pre-segmenting Very Long Transcribed Speech Recordings (L18-1)

Copied to clipboard

Challenge: a new algorithm that pre-segments long speech recordings into manageable chunks is proposed . the run time of classical text-to-speech alignment algorithms is quadratically growing with the length of the input .
Approach: They propose two algorithms that pre-segment long speech recordings into manageable chunks . first algorithm is fast but cannot guarantee short chunks on noisy recordings or erroneous transcriptions a second algorithm delivers short chunk but is less effective in terms of run time and chunk boundary accuracy .
Outcome: The proposed algorithms reduce the run time of the speech segmentation system to under real-time even on recordings that could not previously be processed.
A Real-life, French-accented Corpus of Air Traffic Control Communications (L18-1)

Copied to clipboard

Challenge: AIRBUS-ATC corpus is a real-life, french-accented speech corpus of air traffic control (ATC) communications . it is composed of 59 hours of transcribed English audio, along with linguistic and meta-data annotations.
Approach: They propose to use a real-life, French-accented speech corpus of ATC communications to build a robust ATC speech recognition engine.
Outcome: The AIRBUS-ATC corpus is composed of 59 hours of transcribed English audio, along with linguistic and meta-data annotations.
Creating Lithuanian and Latvian Speech Corpora from Inaccurately Annotated Web Data (L18-1)

Copied to clipboard

Challenge: Existing acoustic model training data for low resource languages is not enough for low-resource languages such as Lithuanian and Latvian.
Approach: They propose a method to align audio data from the Web with imprecise non-normalised transcripts for acoustic models.
Outcome: The proposed method significantly improves word error rate for Lithuanian from 40% to 23% and word error rates for Latvian from 19% to 17%.
Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach (L18-1)

Copied to clipboard

Challenge: Automated Speech Recognition systems degrade in performance when recognizing accents that are different from the ones in training data.
Approach: They propose to adapt Acoustic Models that are trained on one accent to a target accent by using a small amount of speech data in the target accent.
Outcome: The proposed model can be used to identify accents in Indian English and other languages.
Extending Search System based on Interactive Visualization for Speech Corpora (L18-1)

Copied to clipboard

Challenge: Speech corpora are indispensable to speech research; data centers are set up to meet this demand . it is difficult for speech corpus users to compare and select suitable corporata from the large variety of languages .
Approach: They propose a search system that allows users to search for speech corpora interactively and visually . they add specification attributes and items to a large-scale metadata database "SHACHI"
Outcome: The proposed system can search speech corpora interactively and visually . it would be easier for users to select suitable corporata from the large number of languages available .
German Radio Interviews: The GRAIN Release of the SFB732 Silver Standard Collection (L18-1)

Copied to clipboard

Challenge: GRAIN contains German radio interviews and is annotated on multiple linguistic layers.
Approach: They present GRAIN as part of the SFB732 Silver Standard Collection . GRAIN contains German radio interviews and is annotated on multiple linguistic layers .
Outcome: The GRAIN data set contains German radio interviews and is annotated on multiple linguistic layers.
Preparing Data from Psychotherapy for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: mental health care is a demanding occupation, resulting in a severe gap in patient-centered care . a recent study shows that natural language processing can extract certain aspects of human-human communication.
Approach: They propose to use data from psychotherapy sessions to help improve quality of care . they use feedback and cooperation annotations to assess quality of therapy sessions .
Outcome: The proposed method aims to analyse psychotherapy data and assess its quality . it aims at identifying what qualifies for good feedback or cooperation in therapy sessions .
MirasVoice: A bilingual (English-Persian) speech corpus (L18-1)

Copied to clipboard

Challenge: Existing research and development areas in speech recognition are focused on the language of speakers.
Approach: They propose to use a bilingual (English-Farsi) speech corpus to validate and explore speaker verification systems.
Outcome: The proposed corpus can be used in a variety of language dependent and independent applications.
Dialog Intent Structure: A Hierarchical Schema of Linked Dialog Acts (L18-1)

Copied to clipboard

Challenge: a schema for dialog representation captures the pragmatic intents of the conversation independently from any semantic representation.
Approach: They propose a hierarchical and extensible schema for dialog representation . schema captures pragmatic intents of conversation independently from any semantic representation based on semantic content .
Outcome: The proposed schema captures the pragmatic intents of the conversation independently from any semantic representation.
JDCFC: A Japanese Dialogue Corpus with Feature Changes (L18-1)

Copied to clipboard

Challenge: Existing corpora focus on emotional expressions in conversations, but there are no large-scale corpors focusing on the relationships between emotions and utterances.
Approach: They propose a Japanese Feature Change Knowledge Base (JFCKB) that focuses on emotional expressions in conversations.
Outcome: The proposed corpus can recognize reasonableness of a given conversation.
Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags (L18-1)

Copied to clipboard

Challenge: Large-scale conventional dialogue corpora are mainly built for specified tasks with specially designed dialogue states.
Approach: They propose to annotate large-scale dialogue data with an extended ISO-24617-2 dialogue act tag-set to model a natural conversation with machines.
Outcome: The proposed corpus covers a wider range of dialogue tasks than existing task-oriented systems or text-chat systems.
The Niki and Julie Corpus: Collaborative Multimodal Dialogues between Humans, Robots, and Virtual Agents (L18-1)

Copied to clipboard

Challenge: Niki and Julie corpus contains more than 600 dialogues between humans and robots . corpus includes audio and video recordings, results of ranking tasks, questionnaire responses .
Approach: the corpus contains more than 600 dialogues between human participants and a robot . the dialogues are part of a collaborative item-ranking task designed to measure influence .
Outcome: the corpus contains more than 600 dialogues between human participants and a robot or virtual agent . the dialogues contain conversational errors by the robot, which simulates typical of modern automated agents .
Constructing a Chinese Medical Conversation Corpus Annotated with Conversational Structures and Actions (L18-1)

Copied to clipboard

Challenge: Recent studies have found that patients' advocacy for antibiotic treatment is consequential on antibiotic over-prescribing.
Approach: They propose to analyze a manually transcribed corpus of medical dialogue in Chinese pediatric consultations with annotation of conversational structures and actions.
Outcome: The proposed corpus can shed light on ways to improve physician-patient communication in order to reduce antibiotic over-prescribing.
Predicting Nods by using Dialogue Acts in Dialogue (L18-1)

Copied to clipboard

Challenge: Existing studies have generated nods from the final morphemes at the end of an utterance.
Approach: They propose to generate head nods from Japanese dialogues using morphemes . they compile a corpus of 24 dialogues including utterance and nod information .
Outcome: The proposed model outperforms a model using morpheme information in the Japanese language and shows that dialog acts can predict nods.
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.
A Semi-autonomous System for Creating a Human-Machine Interaction Corpus in Virtual Reality: Application to the ACORFORMed System for Training Doctors to Break Bad News (L18-1)

Copied to clipboard

Challenge: Existing methods for training doctors to break bad news are expensive and time consuming.
Approach: They propose a method to collect a corpus of human-machine interactions and then construct a semi-autonomous system based on the collected corpus.
Outcome: The proposed system is based on a corpus-based method to analyze human-machine interactions and then develop fully autonomous prototype.
Construction of English-French Multimodal Affective Conversational Corpus from TV Dramas (L18-1)

Copied to clipboard

Challenge: Existing technologies for speech recognition and speech synthesis focus on non-verbal content and paralinguistic information.
Approach: They propose to construct a multimodal affective conversational corpus based on TV dramas . their data contain parallel English-French languages in lexical, acoustic, and facial features .
Outcome: The proposed corpus can be used to assess speech recognition, speech recognition and synthesis, linguistic, and paralinguistic speech-to-speech translation and multimodal dialog systems.
QUEST: A Natural Language Interface to Relational Databases (L18-1)

Copied to clipboard

Challenge: Current systems focus on simple queries but neglect nested queries . nesting is a problem in SQL, but there is no easy way to achieve it .
Approach: They propose a system which can handle nested logic queries without restrictions . they propose QUEST, which can cope with nesting queries without restriction .
Outcome: The proposed system outperforms a baseline system by 11% accuracy.
TF-LM: TensorFlow-based Language Modeling Toolkit (L18-1)

Copied to clipboard

Challenge: Existing deep learning tools offer building blocks but training and building models takes time and knowledge.
Approach: They propose to make available LSTM language models trained on Dutch texts and English benchmarks.
Outcome: The proposed model can be used to test the perplexity, predict the next word(s), re-score hypotheses or generate debugging files for interpolation with n-gram models.
Grapheme-level Awareness in Word Embeddings for Morphologically Rich Languages (L18-1)

Copied to clipboard

Challenge: a study of inflectional and non-alphabetic languages shows word vectors are sparse in data sparsity due to the morphological system of a language and its syllables.
Approach: They propose a grapheme-level coding procedure for neural word embedding that uses syllable characters to represent word-internal features.
Outcome: The proposed model is more capable of representing functional and semantic similarities than syllable-level and word-level models.
Building a Constraint Grammar Parser for Plains Cree Verbs and Arguments (L18-1)

Copied to clipboard

Challenge: a Constraint Grammar parser for Plains Cree is developed for its rich morphology and syntax . the focus of the parsers is to identify relationships between verbs and arguments .
Approach: This paper presents the development and application of a Constraint Grammar parser for Plains Cree . the parsers focus on the identification of relationships between verbs and arguments .
Outcome: The proposed parser is based on a Constraint Grammar formalism for the Plains Cree language . it allows for the identification of common word order patterns and relationships between word order and information structure .
BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages (L18-1)

Copied to clipboard

Challenge: In an evaluation using fine-grained entity typing as testbed, BPEmb performs competitively . pre-trained subword embeddings for BPE units are commonly available .
Approach: They present a collection of pre-trained subword embeddings in 275 languages . they use fine-grained entity typing as testbed to evaluate BPEmb .
Outcome: The proposed method performs better than other methods, but requires less resources and no tokenization.
Reference production in human-computer interaction: Issues for Corpus-based Referring Expression Generation (L18-1)

Copied to clipboard

Challenge: Referring Expression Generation studies often use web-based data collection tasks without a particular addressee in mind.
Approach: They developed a parallel corpus of monologue and dialogue referring expressions and an annotated corpus to compare instances produced in both modes of communication.
Outcome: The results suggest that human reference production may be affected by the presence of a second (specific) human participant as the receiver of the communication in a number of ways.
Definite Description Lexical Choice: taking Speaker’s Personality into account (L18-1)

Copied to clipboard

Challenge: Referring Expression Generation (REG) lexical choice is the subtask that provides words to express an input meaning representation.
Approach: They propose a personality-dependent lexical choice model for Referring Expression Generation (REG) that provides words to express a given input meaning representation.
Outcome: The proposed model outperforms a standard lexicalisation model based on meaning-to-text mappings and personality information.
Referring Expression Generation in time-constrained communication (L18-1)

Copied to clipboard

Challenge: In game-like applications, an underlying Natural Language Generation system may have to express urgency or other dynamic aspects of a fast-evolving situation as text.
Approach: They propose to use a task of Referring Expression Generation (REG) to generate natural language text from a visual context provided by a video game application.
Outcome: The proposed algorithm outperforms standard approaches to REG in time-constrained communication.
Incorporating Semantic Attention in Video Description Generation (L18-1)

Copied to clipboard

Challenge: Existing methods to generate video descriptions fail to mention objects and actions in videos .
Approach: They propose an LSTM-based sequence-to-sequence model with semantic attention mechanism for video description generation that includes external fine-grained visual information.
Outcome: The proposed model can selectively focus on external fine-grained visual information and have a better quality of video descriptions.
GenDR: A Generic Deep Realizer with Complex Lexicalization (L18-1)

Copied to clipboard

Challenge: Generic deep realizers are used for natural language generation, but they are not yet fully dominated by statistical or neuronal methods.
Approach: They propose a generic deep realizer that produces syntactic dependency structures in languages . they use a graph transducer to lexicalize multiword expressions and build on it .
Outcome: The proposed system produces syntactic dependency structures in English, French, Lithuanian and Persian . it is generic in that it is designed to operate across a wide range of languages and applications .
A Detailed Evaluation of Neural Sequence-to-Sequence Models for In-domain and Cross-domain Text Simplification (L18-1)

Copied to clipboard

Challenge: Xu et al., 2016) show that a simple neural architecture can be efficiently used for in-domain and cross-domain text simplification.
Approach: They evaluate neural sequence-to-sequence models for text simplification on Wikipedia and Newsela datasets.
Outcome: The proposed model can generalize across corpora and overcome challenges when tested on Wikipedia and Newsela datasets.
Don’t Annotate, but Validate: a Data-to-Text Method for Capturing Event Data (L18-1)

Copied to clipboard

Challenge: Existing methods to create event data are limited by ambiguity and variation in the data.
Approach: They propose a method to obtain large volumes of text corpora with event data . they use a tool to annotate texts and enrich the reference texts with event coreference annotations.
Outcome: The proposed method obtains large volumes of high-quality text corpora with event data . the data obtained with this method have high precision and at a large scale .
RDF2PT: Generating Brazilian Portuguese Texts from RDF Data (L18-1)

Copied to clipboard

Challenge: Existing approaches to generate natural language from RDF data have been proposed to generate texts in Brazilian Portuguese.
Approach: They propose a rule-based approach to verbalize RDF data to Brazilian Portuguese language.
Outcome: The proposed approach generates text similar to that generated by humans and can hence be easily understood.
Towards a music-language mapping (L18-1)

Copied to clipboard

Challenge: a novel research idea investigates the possibility of musical input to speech interaction systems.
Approach: They propose a musical language processing idea that investigates the possibility of musical input to speech interaction systems.
Outcome: The proposed method could be used to map musical pieces and dialogues based on frequency of musical patterns . the proposed method is universal among different languages and easy to learn for musicians .
Up-cycling Data for Natural Language Generation (L18-1)

Copied to clipboard

Challenge: Existing systems for creating adaptive texts from cultural heritage data require expert input . a number of research projects have focused on using NLG systems to create multilingual adaptive texts .
Approach: They propose automatic processes which aim to reduce the need for expert input . they normalize the dates and names which occur in the data and link to the Semantic Web .
Outcome: The proposed processes reduce the need for expert input during conversion and up-cycling process.
Neural Models of Selectional Preferences for Implicit Semantic Role Labeling (L18-1)

Copied to clipboard

Challenge: Existing studies on implicit semantic role labeling have been limited due to the lack of training data.
Approach: They propose to use more complex machine learning models trained on a large amount of explicit roles to recover implicit roles.
Outcome: The proposed models outperform baseline models on ON5V dataset, but have mostly negative results . they show that multi-way selectional preference improves results for predicting explicit semantic roles, but harms performance for implicit roles.
A database of German definitory contexts from selected web sources (L18-1)

Copied to clipboard

Challenge: a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts.
Approach: They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction .
Outcome: The proposed method is based on a web corpus and a robust pattern-based extraction method.
Annotating Abstract Meaning Representations for Spanish (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a semantic representation language for natural language processing.
Approach: They propose a method that would lay the groundwork for building a large semantic bank for Spanish . they propose to use a database to annotate AMRs for other languages .
Outcome: The proposed method would lay the groundwork for building a large semantic bank for Spanish and guide those who would like to implement it for other languages.
Browsing the Terminological Structure of a Specialized Domain: A Method Based on Lexical Functions and their Classification (L18-1)

Copied to clipboard

Challenge: a method for browsing relations between terms and unveiling terminological structure of a specialized domain is described.
Approach: They propose a method for browsing relations between terms and unveiling terminological structure of a specialized domain.
Outcome: The proposed method expands graph that takes as input relations encoded in a terminological resource called the DiCoEnviro . it provides an explicit and intuitive representation of a wide variety of relations while making some graphical choices to make their interpretation as clear as possible for end users.
Rollenwechsel-English: a large-scale semantic role corpus (L18-1)

Copied to clipboard

Challenge: The Rollenwechsel-English corpus is a large corpus of automatically-labelled semantic frames extracted from the ukWaC corpus and BNC using Propbank roles.
Approach: They present a large corpus of automatically-labelled semantic frames extracted from ukWaC and BNC using Propbank roles.
Outcome: The rollenwechsel-English corpus is a large corpus of automatically-labelled semantic frames extracted from the ukWaC corpus and BNC using Propbank roles.
Towards a Standardized Dataset for Noun Compound Interpretation (L18-1)

Copied to clipboard

Challenge: Noun compounds are interesting constructs in Natural Language Processing . lack of standardized set of relation inventories and annotated datasets hinders interpretation .
Approach: They propose a dataset that uses FrameNet as its semantic relation inventory to examine noun compounds.
Outcome: The proposed dataset is linguistically grounded and uses FrameNet as its semantic relation inventory.
Structured Interpretation of Temporal Relations (L18-1)

Copied to clipboard

Challenge: Temporal relations between events and time expressions are often modeled in an unstructured manner, resulting in inconsistent and incomplete annotation and computational modeling.
Approach: They propose an annotation approach where events and time expressions form a dependency tree in which each dependency relation corresponds to an instance of temporal anaphora.
Outcome: The proposed approach annotates 235 documents in news and narratives with 48 doubly annotated documents.
NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System (L18-1)

Copied to clipboard

Challenge: NL2Bash is a new semantic parsing problem for mapping English sentences to Bash commands.
Approach: They propose a dataset of English commands and expert-written Bash commands to map English sentences to Bash.
Outcome: The proposed methods are significantly larger (from two to ten times) than most existing benchmarks.
World Knowledge for Abstract Meaning Representation Parsing (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) parsers are based on annotated graphs, but there is still room for improvement .
Approach: They examine the role played by world knowledge in parsing errors in a state-of-the-art parser . they examine the effects of different types of world knowledge on parsers .
Outcome: The proposed model improves on multiple fine-grained metrics, including a 6% increase in named entity F-score, and provides insight into the potential of world knowledge for future work in Abstract Meaning Representation parsing.
Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research (L18-1)

Copied to clipboard

Challenge: Existing methods to improve transcription and indexing quality of Oral History interviews are not available.
Approach: They propose to use a German Oral History test-set to improve transcription and indexing quality . they propose to combine acoustic modeling techniques with sophisticated neural networks .
Outcome: The proposed system reduces word error rate by 28.3% on German Oral History test-set compared to baseline system . the Fraunhofer IAIS Audio Mining system can process long audio-files to automatically create time-aligned transcriptions.
Sound Signal Processing with Seq2Tree Network (L18-1)

Copied to clipboard

Challenge: Recent LSTM models have been used to model sequential data processing tasks because of their ability to preserve previous information weighted on distance.
Approach: They propose to use a tree-structured tree-based neural network architecture to solve the problem of unbalanced connections between data units inside and outside semantic groups.
Outcome: The proposed model outperforms the state-of-the-art Bidirectional LSTM model on a signal and noise separation task.
Open ASR for Icelandic: Resources and a Baseline System (L18-1)

Copied to clipboard

Challenge: Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed.
Approach: They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic.
Outcome: The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary.
Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models (L18-1)

Copied to clipboard

Challenge: Existing methods for speaker modeling are based on hand-crafted statistics and ad hoc to a certain application.
Approach: They propose to use speaker classification as a surrogate task for general speaker modeling and collect massive data to facilitate research in this direction.
Outcome: The proposed models outperform the existing models and are feasible with speaker identity information.
Discriminating between Similar Languages on Imbalanced Conversational Texts (L18-1)

Copied to clipboard

Challenge: Empirical results suggest that our system achieves an accuracy of 95.7% on our Uyghur and Kazakh dataset, which is higher than that of the CNN classifier.
Approach: They propose to build a balanced Uyghur and Kazakh corpus and build morphological classifiers to discriminate between the two languages.
Outcome: The proposed system outperforms the champions on both test sets B1 and B2.
Data-Driven Pronunciation Modeling of Swiss German Dialectal Speech for Automatic Speech Recognition (L18-1)

Copied to clipboard

Challenge: a Swiss German speech recognizer is trained using a standard German annotation model.
Approach: They propose to train a Swiss German speech recognition system using a standard German annotation model.
Outcome: The proposed system is based on a standard German annotation model and a grapheme-to-phoneme conversion model.
Simulating ASR errors for training SLU systems (L18-1)

Copied to clipboard

Challenge: Existing methods to simulate automatic speech recognition errors from manual transcriptions are not available during training of the SLU model.
Approach: They propose to use acoustic and linguistic word embeddings to define a similarity measure between words to predict ASR confusions.
Outcome: The proposed method significantly improves the performance of spoken language understanding systems.
Evaluation of Feature-Space Speaker Adaptation for End-to-End Acoustic Models (L18-1)

Copied to clipboard

Challenge: Existing speaker adaptation algorithms for BLSTM-CTC AMs are lacking . TED-LIUM corpus shows speaker adaptation provides 11-20% word error rate reduction over baseline model built on raw filter-bank features.
Approach: They propose to use feature-space adaptation techniques for bidirectional long short term memory (BLSTM) recurrent neural network based acoustic models trained with the connectionist temporal classification objective function to improve speaker adaptation.
Outcome: The proposed approach provides up to 11-20% of word error reduction over baseline models on the TED-LIUM corpus.
Creating New Language and Voice Components for the Updated MaryTTS Text-to-Speech Synthesis Platform (L18-1)

Copied to clipboard

Challenge: a reboot of the MaryTTS system became unavoidable due to the number of people who have contributed to its development over the years.
Approach: They propose a workflow to create components for the MaryTTS text-to-speech synthesis platform.
Outcome: The proposed workflow is compatible with the updated MaryTTS architecture, enabling new features and state-of-the-art paradigms such as synthesis based on deep neural networks (DNNs).
Speech Rate Calculations with Short Utterances: A Study from a Speech-to-Speech, Machine Translation Mediated Map Task (L18-1)

Copied to clipboard

Challenge: Computer mediated multi-lingual communication is becoming more frequent.
Approach: They propose a method to verify if an utterance within a corpus is pronounced at a fast or slow pace.
Outcome: The proposed method provides a value for the utterance speech rate in a corpus of short utterations.
Beyond Generic Summarization: A Multi-faceted Hierarchical Summarization Corpus of Large Heterogeneous Data (L18-1)

Copied to clipboard

Challenge: Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader.
Approach: They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically.
Outcome: The proposed method can be used to develop and evaluate hierarchical summarization systems.
A New Annotated Portuguese/Spanish Corpus for the Multi-Sentence Compression Task (L18-1)

Copied to clipboard

Challenge: Existing corpus for Multi-sentence Compression (MSC) tasks is limited to English . a dataset is available for MSC tasks in the French language .
Approach: They propose a new corpus for Multi-Sentence Compression task in Portuguese and Spanish.
Outcome: The proposed corpus is compared with two state-of-the-art systems in Portuguese and Spanish.
Live Blog Corpus for Summarization (L18-1)

Copied to clipboard

Challenge: Live blogs are increasingly popular news format to cover breaking news and live events.
Approach: They propose to collect corpora for automatic live blog summarization by a web-based system . they make the tools publicly available to encourage the research community .
Outcome: The proposed method improves the accuracy of live blog summarization by allowing for public access to the corpus.
TSix: A Human-involved-creation Dataset for Tweet Summarization (L18-1)

Copied to clipboard

Challenge: a new dataset for tweet summarization is available for free.
Approach: They propose a dataset for tweet summarization that uses human annotations to evaluate extractive summarizing methods.
Outcome: The proposed dataset includes six events collected from Twitter . human-annotated gold-standard references facilitate evaluation, the study shows .
A Workbench for Rapid Generation of Cross-Lingual Summaries (L18-1)

Copied to clipboard

Challenge: a tool for automating cross-lingual information access is needed in multilingual societies . current state of machine translation is not able to generate publishable articles from English .
Approach: They propose a web-based tool for human editing of cross-lingual summaries . it generates publishable summary in a number of Indian Languages for news articles originally published in english .
Outcome: The proposed tool can generate publishable summaries in multiple languages with minimal human effort and collect detailed logs on the process.
Annotation and Analysis of Extractive Summaries for the Kyutech Corpus (L18-1)

Copied to clipboard

Challenge: Summarization of multi-party conversation requires corpora to analyze characteristics of conversations and construct a method for summary generation.
Approach: They propose to annotate a Japanese conversation corpus for a decision-making task . they compare extractive summarization methods with the annotated extractive summary .
Outcome: The proposed corpus is the first annotated for conversation summarization tasks and freely available to anyone.
A Repository of Corpora for Summarization (L18-1)

Copied to clipboard

Challenge: Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task.
Approach: They propose a repository containing corpora available to train and evaluate automatic summarization systems.
Outcome: The proposed system is based on a repository of corpora available for summarization tasks.
Auto-hMDS: Automatic Construction of a Large Heterogeneous Multilingual Multi-Document Summarization Corpus (L18-1)

Copied to clipboard

Challenge: Existing datasets for automatic text summarization are small and focused on newswires.
Approach: They propose to automatically generate a large multilingual multi-document summarization corpus using Wikipedia articles as summaries and to automatically search for appropriate source documents.
Outcome: The proposed corpus contains 7,316 topics in English and German with different summary lengths and number of source documents.
PyrEval: An Automated Method for Summary Content Analysis (L18-1)

Copied to clipboard

Challenge: PyrEval automates manual summarization evaluation of abstractive summarizing systems . extractive summaries that select complete sentences have shifted in recent years .
Approach: They propose a method for automatic summarization evaluation that automates the manual pyramid method by using pre-trained vectors and a greedy algorithm to evaluate the pyramid content.
Outcome: The proposed method can be applied to human and machine summaries with no retraining and in excellent time.
Mapping Texts to Scripts: An Entailment Study (L18-1)

Copied to clipboard

Challenge: Script knowledge is crucial for text understanding systems, providing a basis for commonsense inference.
Approach: They propose to map event mentions in a text to script events using crowdsourced event descriptions.
Outcome: The proposed model improves the performance of text-to-script mapping systems by integrating paraphrase sets with crowdsourced event descriptions.
Semantic Equivalence Detection: Are Interrogatives Harder than Declaratives? (L18-1)

Copied to clipboard

Challenge: Semantic Text Similarity (STS) tasks are often not seen as similar to semantic equivalence detection tasks.
Approach: They propose to assess the performance of different approaches to STS over different types of textual segments.
Outcome: The proposed methods differ in performance over different types of textual segments, including declaratives and interrogatives, under conditions of comparability.
CEFR-based Lexical Simplification Dataset (L18-1)

Copied to clipboard

Challenge: Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective.
Approach: They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile .
Outcome: The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus.
CLARIN: Towards FAIR and Responsible Data Science Using Language Resources (L18-1)

Copied to clipboard

Challenge: CLARIN is a European Research Infrastructure providing access to language resources and tools for researchers in the humanities and social sciences.
Approach: This paper outlines the CLARIN vision and strategy . it explains how the design and implementation of CLARINS are compliant with the FAIR principles .
Outcome: The paper outlines the CLARIN vision and strategy and explains how it is compliant with the FAIR principles: findability, accessibility, interoperability and reusability of data.
From ‘Solved Problems’ to New Challenges: A Report on LDC Activities (L18-1)

Copied to clipboard

Challenge: This paper reports on the activities of the Linguistic Data Consortium .
Approach: This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report .
Outcome: The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world .
New directions in ELRA activities (L18-1)

Copied to clipboard

Challenge: ELRA is an indispensable middle-man in the field of Language Resources . new directions of work are being undertaken to answer the needs of this ever-moving community .
Approach: This paper addresses new directions of work undertaken by the ELRA in the field of Language Resources . it describes ELDA's regular activities updates and future projects .
Outcome: This paper addresses these new directions and describes ELRA regular activities updates.
A Framework for Multi-Language Service Design with the Language Grid (L18-1)

Copied to clipboard

Challenge: International NPO/NGOs are struggling with the design and development of tools and systems for multi-language communication in the real world.
Approach: They propose a framework for service design with the Language Grid by bridging the gap between language service infrastructures and multi-language systems.
Outcome: The proposed framework bridges the gap between language service infrastructures and multi-language systems by allowing users to design and develop multilingual communication services and tools in the real world.
Language Technology for Multilingual Europe: An Analysis of a Large-Scale Survey regarding Challenges, Demands, Gaps and Needs (L18-1)

Copied to clipboard

Challenge: a survey titled "Language Technology for Multilingual Europe" was conducted between May and June 2017 . 634 participants in 52 countries responded to the survey .
Approach: a large-scale survey was conducted to assess the best multilingual technologies in Europe. a total of 634 participants in 52 countries responded to the survey.
Outcome: The study aims to identify the biggest challenges, obstacles and gaps in European language technology . participants were encouraged to share concrete suggestions and recommendations on how present challenges can be turned into opportunities .
Annotating High-Level Structures of Short Stories and Personal Anecdotes (L18-1)

Copied to clipboard

Challenge: Existing theories for narrative structures have been challenging to operationalize . authors present an annotation scheme to help computer systems understand stories better .
Approach: They propose to consolidate and extend existing narratological theories and an annotation scheme . they will support an approach that enables systems to intelligently sustain complex communications with humans .
Outcome: The proposed method consolidates and extends existing narratological theories . it will support an approach that enables systems to intelligently sustain complex communications with humans .
Discovering the Language of Wine Reviews: A Text Mining Account (L18-1)

Copied to clipboard

Challenge: odors and flavors are often expressed in wine reviews, but they are often not.
Approach: They use a corpus of wine reviews to find out what wine is like in a review . they use lexical bag-of-words features, domain-specific terminology features and word embedding features to train machine learning.
Outcome: The proposed model predicts the wine's color, grape variety, and country of origin based on the review text alone.
Toward An Epic Epigraph Graph (L18-1)

Copied to clipboard

Challenge: a database of epigraphs is being developed to reveal literary influence as a set of connections between authors over time.
Approach: a database of epigraphs is created to map literary influence as a set of connections between authors . the database is being developed under an open license .
Outcome: a database of epigraphs is being developed to reveal literary influence over time . the database includes epigraph quotations from over 12,000 literary works . authors use epigraph to set theme and link work to existing body of literature .
Delta vs. N-Gram Tracing: Evaluating the Robustness of Authorship Attribution Methods (L18-1)

Copied to clipboard

Challenge: a novel authorship attribution method is developed for short texts . delta measures are well-established, but N-gram tracing is not robust enough .
Approach: They propose to use delta measures and N-gram tracing to compare short texts . they find they are highly sensitive to the choice of authors and texts in the corpus .
Outcome: The proposed methods are highly sensitive to the selection of authors and texts in the comparison corpus.
An Attribution Relations Corpus for Political News (L18-1)

Copied to clipboard

Challenge: Existing resources for recognizing attributions in context are limited in size and completeness.
Approach: They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction .
Outcome: The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date.
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)

Copied to clipboard

Challenge: a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images .
Approach: They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions .
Outcome: The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces.
Adapting Serious Game for Fallacious Argumentation to German: Pitfalls, Insights, and Best Practices (L18-1)

Copied to clipboard

Challenge: 'homeschooling' and 'death penalty' are non-existent in Germany, while being highly controversial topics of discussion in the United States.
Approach: They propose to port Argotario (serious game for learning argumentation fallacies) to another language and analyze users' behavior and in-game created data to assess dissemination strategies and qualitative aspects of the resulting corpus.
Outcome: The proposed game is based on a German-based game platform that can be used to learn argumentation fallacies.
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French.
Approach: They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French.
Outcome: The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French.
Improving Machine Translation of Educational Content via Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Using crowdsourcing to train neural machine translation models is expensive and expensive . professional outsourcing of bilingual data is expensive if the translations are of a lower quality .
Approach: They analyze the impact of crowdsourcing on the quality of in-domain training data . they use translations of MOOCs from English to eleven languages to fine-tune machine translation models .
Outcome: The proposed method improves on general-domain training data and with pre-existing in-domain corpora.
Grounding Gradable Adjectives through Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Often, texts describe interactions using vague, high-level language . crowdsourcing is expensive and requires extensive literature review and time .
Approach: They propose a method for estimating concrete groundings for a set of gradable adjectives by crowdsourcing human intuitions and fitting a mixed effects model to the text.
Outcome: The proposed model can generalize to unseen data and has a predictive R 2 of 0.632 in general and 0.677 on a subset of high-frequency adjectives.
Evaluation Phonemic Transcription of Low-Resource Tonal Languages for Language Documentation (L18-1)

Copied to clipboard

Challenge: Language documentation involves recording the speech of native speakers.
Approach: They propose to use a neural network architecture to model phonemes and tones versus modelling them separately.
Outcome: The proposed method improves efficiency, minimizes typographical errors and maintains transcription faithfulness to acoustic signal while highlighting phonetic and phonemic facts for linguistic consideration.
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)

Copied to clipboard

Challenge: a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels .
Approach: They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations .
Outcome: The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task .
Chahta Anumpa: A multimodal corpus of the Choctaw Language (L18-1)

Copied to clipboard

Challenge: a corpus of texts representing the Choctaw language is presented for use in linguistic studies.
Approach: They present a general use corpus for the Choctaw language in the southeastern u.s. the corpus contains audio, video, and text resources, with many texts also translated in english.
Outcome: The proposed corpus provides documentation support for the threatened language . the data set includes audio, video, and text resources .
BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools (L18-1)

Copied to clipboard

Challenge: Approximately 50 hours of Bàsàá speech were collected and then carefully re-spoken and orally translated into French .
Approach: They propose to provide an automatic phonetic transcription using a set of derived phone-like units.
Outcome: The proposed method provides an automatic phonetic transcription using a set of derived phone-like units.
Researching Less-Resourced Languages – the DigiSami Corpus (L18-1)

Copied to clipboard

Challenge: DigiSami project aims to support research on endangered languages . it uses spoken corpus and speech technology for the Fenno-Ugric language North Sami .
Approach: They describe the DigiSami project and its research results for the Fenno-Ugric language North Sami . they discuss ethical and privacy issues related to data collection for less-resourced languages and indigenous communities .
Outcome: The DigiSami project focuses on spoken corpus collection and speech technology for the Fenno-Ugric language North Sami.
The MADAR Arabic Dialect Corpus and Lexicon (L18-1)

Copied to clipboard

Challenge: Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation.
Approach: They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project.
Outcome: The proposed resources are the first of their kind in terms of their coverage and fine granularity.
Designing a Collaborative Process to Create Bilingual Dictionaries of Indonesian Ethnic Languages (L18-1)

Copied to clipboard

Challenge: a constraint-based approach has been proven useful for inducing bilingual dictionary for low-resource languages.
Approach: They propose a constraint-based approach for inducing bilingual dictionary for low-resource languages . they propose heuristic plan that only utilizes manual investment by native speaker .
Outcome: The proposed approach outperforms the heuristic plan with 63.3% cost reduction.
Constructing a Lexicon of Relational Nouns (L18-1)

Copied to clipboard

Challenge: Existing systems for extracting relations expressed using nouns do not exist for relational noun.
Approach: They contribute a lexicon of 6,224 labeled nouns which includes 1,446 relational noun.
Outcome: The proposed classifier achieves 70.4% F1 on held out nouns among the most common 2,500 word types in Gigaword.
Creating Large-Scale Multilingual Cognate Tables (L18-1)

Copied to clipboard

Challenge: Low-resource languages often suffer from a lack of high-coverage lexical resources.
Approach: They propose a method to generate cognate tables by clustering words from existing lexical resources.
Outcome: The proposed method outperforms baselines on the Romance and Turkic language families.
Lexical Profiling of Environmental Corpora (L18-1)

Copied to clipboard

Challenge: a method for distinguishing lexical layers in environmental corpora is described . the general environmental lexicon is well-distributed in a specialized corpus and specific to this type of corporata.
Approach: They propose a method for distinguishing lexical layers in environmental corpora . they aim to identify the general environmental lexicon (GEL) and assess its specificity .
Outcome: The proposed method is well-distributed in a specialized corpus and specific to this type of corpora.
Linking, Searching, and Visualizing Entities in Wikipedia (L18-1)

Copied to clipboard

Challenge: Existing systems to extract, index, search, and visualize entities in Wikipedia are not strings, but unique identifiers from Wikidata.
Approach: They propose a system to extract, index, search, and visualize entities in Wikipedia . they use a document model to store linguistic annotations and a string matching engine .
Outcome: The proposed system achieves CEAFm scores of 70.0 on English, 64.4 on Chinese, and 66.5 on Spanish.
Learning to Map Natural Language Statements into Knowledge Base Representations for Knowledge Base Construction (L18-1)

Copied to clipboard

Challenge: Currently, the construction and updating of knowledge bases rely on human labor.
Approach: They propose to map relational phrases in triples from natural language to knowledge base predicate format.
Outcome: The proposed mapping results show high quality and promising coverage on relational phrases compared to previous research.
Building a Knowledge Graph from Natural Language Definitions for Interpretable Text Entailment Recognition (L18-1)

Copied to clipboard

Challenge: a conceptual model for dictionary definitions is used to construct a knowledge graph from natural language definitions.
Approach: They propose a method for automatically building a graph world knowledge base from natural language definitions.
Outcome: The proposed method was used in an interpretable text entailment recognition approach.
Combining rule-based and embedding-based approaches to normalize textual entities with an ontology (L18-1)

Copied to clipboard

Challenge: a method to normalize multi-word terms with concepts from a domain-specific ontology is proposed . a large part of knowledge is expressed in textual form, such as in scientific articles .
Approach: They propose a method to normalize multi-word terms with concepts from a domain-specific ontology.
Outcome: The proposed method outperforms existing methods on a categorization task in bacterial habitats . the results are encouraging, and the proposed method is expected to be widely used in the biomedical/biological field .
T-REx: A Large Scale Alignment of Natural Language with Knowledge Base Triples (L18-1)

Copied to clipboard

Challenge: Existing datasets that provide alignments between natural language and knowledge bases (KB) triples are limited in size, lack coverage and are of unreported quality.
Approach: They propose to build a large scale dataset of alignments between Wikipedia abstracts and Wikidata triples that is two orders of magnitude larger than the largest available alignments dataset.
Outcome: The proposed dataset is two orders of magnitude larger than the largest available dataset and covers 2.5 times more predicates.
Multilingual Parallel Corpus for Global Communication Plan (L18-1)

Copied to clipboard

Challenge: In this paper, we introduce the Global Communication Plan (GCP) Corpus . the corpus is sentence-aligned and covers ten languages, including many Asian languages .
Approach: They introduce the Global Communication Plan (GCP) Corpus, a multilingual parallel corpus . it is sentence-aligned and covers ten languages, including many Asian languages .
Outcome: The proposed corpus is sentence-aligned and covers ten languages, including many Asian languages.
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)

Copied to clipboard

Challenge: Scielo database contains articles from several research domains.
Approach: They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish.
Outcome: The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish.
NegPar: A parallel corpus annotated for negation (L18-1)

Copied to clipboard

Challenge: NegPar is the first parallel corpus annotated for negation in the narrative domain.
Approach: They present NegPar, a parallel corpus annotated for negation in the narrative domain . they follow the annotation guidelines in the CONANDOYLE-NEG corpus .
Outcome: The proposed corpus is based on the CONANDOYLE-NEG corpus and is reannotated to ensure more consistent and interpretable representations.
The IIT Bombay English-Hindi Parallel Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of 1.49 million parallel segments is available in the public domain . the corpus is the largest publicly available English-Hindi parallel corpus .
Approach: They present the IIT Bombay English-Hindi Parallel Corpus . they present a compilation of public and private parallel corpora .
Outcome: The corpus contains 1.49 million parallel segments, of which 694k were not previously available in the public domain.
Extracting an English-Persian Parallel Corpus from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: Existing methods to extract parallel sentences from Wikipedia are limited for some language pairs such as Persian-English.
Approach: They propose a bidirectional method to extract parallel sentences from Wikipedia . they add extracted sentences to existing training data and use IR system to measure similarity .
Outcome: The proposed method outperforms the one-directional approach in analyzing translation data from two translation systems and IR systems.
Learning Word Vectors for 157 Languages (L18-1)

Copied to clipboard

Challenge: Distributed word representations, or word vectors, have been used in natural language processing for many tasks.
Approach: They propose to use the encyclopedia Wikipedia and the common crawl corpus to train distributed word representations on large corpora and use them in downstream tasks.
Outcome: The proposed model performs very well on 10 languages for which evaluation dataset exists.
SumeCzech: Large Czech News-Based Summarization Dataset (L18-1)

Copied to clipboard

Challenge: Summarization of documents is a well-studied NLP task, but only a few datasets are available for Czech.
Approach: They propose to use a Czech news-based summarization dataset to evaluate document summarizing . they propose a language-agnostic variant of the ROUGE metric to enable automatic evaluation .
Outcome: The proposed dataset contains more than a million Czech news articles . the proposed approach is strong abstractive and language-agnostic .
A Diachronic Corpus for Literary Style Analysis (L18-1)

Copied to clipboard

Challenge: Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data .
Approach: They propose a resource for diachronic style analysis in particular the analysis of literary authors over time.
Outcome: The proposed resource can be used to analyze literary authors over time.
Text Simplification from Professionally Produced Corpora (L18-1)

Copied to clipboard

Challenge: Existing approaches to Text Simplification rely on the Wikipedia-Simple Wikipedia parallel corpus, which is used for many tasks.
Approach: They propose to use the Newsela corpus to extract 550, 644 complex-simple sentence pairs from the corpus and introduce a lexical simplifier that uses the corpu to generate candidate simplifications.
Outcome: The proposed model outperforms state-of-the-art approaches and generates candidate simplifications from the newsela corpus.
Intertextual Correspondence for Integrating Corpora (L18-1)

Copied to clipboard

Challenge: Using intertextual correspondence, we can combine annotated text corpora to create new annotation connections.
Approach: They propose to use intertextual correspondence as an integrative technique for combining annotated text corpora.
Outcome: The proposed technique can be used to build argumentative arguments in two annotated text corpora.
A Gold Anaphora Annotation Layer on an Eye Movement Corpus (L18-1)

Copied to clipboard

Challenge: Anaphora resolution is a complex process in which multiple linguistic factors play a role.
Approach: They used annotated anaphorical pronouns from newspaper articles read by humans to model reading time of pronounes.
Outcome: The proposed resource allows to study human anaphora resolution on natural data.
Annotating Zero Anaphora for Question Answering (L18-1)

Copied to clipboard

Challenge: a large dataset of zero pronouns has been constructed to identify adjunct zero anaphoras . a lack of a dataset covering them has limited our ability to annotate them exhaustively .
Approach: They propose to annotate adjuncts marked by -de in Japanese and a second scheme to annnotate them in a more direct manner.
Outcome: The proposed annotation schemes are more accurate than the first one.
Building Named Entity Recognition Taggers via Parallel Corpora (L18-1)

Copied to clipboard

Challenge: Existing methods to generate semantic processors for languages lacking hand curated data are inefficiently slow and unaffordable in terms of human resources and economic costs.
Approach: They propose to use statistical word alignments to project annotations from multiple sources to a target language.
Outcome: The proposed method is effective to transport NER annotations across languages . it can generate a good statistical model for a new target language .
Cross-Document, Cross-Language Event Coreference Annotation Using Event Hoppers (L18-1)

Copied to clipboard

Challenge: Defined event hoppers for the DEFT Rich Entities, Relations and Events (Rich ERE) annotation task.
Approach: They propose an approach for cross-document, cross-lingual event coreference for the DEFT Rich Entities, Relations and Events (Rich ERE) annotation task.
Outcome: The proposed approach is based on the definition of event hoppers for the DEFT rich entities, relations, events and their attributes . it yields 389 cross-document event hoppings in 505 documents in three languages .
TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection (L18-1)

Copied to clipboard

Challenge: Detecting novelty of an entire document is an AI frontier problem . present state-of-the-art text matching techniques are unable to process such redundancy.
Approach: They propose a document-level novelty detection resource that can be used to benchmark techniques . they crawl news documents across several domains and use it to find out whether they contain new information .
Outcome: The proposed dataset is compared with a standard system for document novelty detection . the proposed system can detect elements that have not appeared before, or new or original .
A Corpus for Multilingual Document Classification in Eight Languages (L18-1)

Copied to clipboard

Challenge: a subset of the Reuters corpus volume 2 is used to evaluate cross-lingual document classification . current best practice is to evaluate document classification on resources in one language and transfer it to another without additional resources.
Approach: They propose to use a subset of the Reuters corpus to evaluate cross-lingual document classification . they propose to add Italian, Russian, Japanese and Chinese to the subset .
Outcome: The proposed subset of the Reuters corpus has balanced class priors for eight languages.
Analyzing Citation-Distance Networks for Evaluating Publication Impact (L18-1)

Copied to clipboard

Challenge: citation networks are used to study scholarly articles' semantic distances and their referencing patterns.
Approach: They propose to analyze the semantic distance of scholarly articles in a citation network to uncover patterns that reflect scientific impact.
Outcome: The proposed method combines semantic distance and content similarity to uncover scientific impact of articles in two different types of publications.
Annotating Educational Questions for Student Response Analysis (L18-1)

Copied to clipboard

Challenge: a new taxonomy and annotated educational corpus of questions is proposed for question answering systems.
Approach: They propose a taxonomy and annotated educational corpus of questions that can be used in automatic questions classification systems.
Outcome: The proposed approach achieves a weighted F1-score of 0.511, overtaking the baseline by 12%.
Incorporating Global Contexts into Sentence Embedding for Relational Extraction at the Paragraph Level with Distant Supervision (L18-1)

Copied to clipboard

Challenge: Existing approaches to relation extraction (RE) only extract relations from sentences that contain two target entities.
Approach: They propose to incorporate global contexts from paragraph-into-sentence embedding into RE . they propose to use a knowledge base to extract relations between pairs of entities .
Outcome: The proposed approach can learn an exact RE from sentences without syntactic parsing.
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge (L18-1)

Copied to clipboard

Challenge: Various approaches for script knowledge extraction and processing have been proposed in recent years.
Approach: They propose a dataset to evaluate natural language understanding approaches based on commonsense knowledge.
Outcome: The proposed dataset provides test cases for the broader natural language understanding community.
A Neural Network Based Model for Loanword Identification in Uyghur (L18-1)

Copied to clipboard

Challenge: Lexical borrowing happens in almost all languages, and we propose a new method to identify loanwords in Uyghur.
Approach: They propose a neural network based loanword identification model for Uyghur that captures past and future information and learns both word level and character level features automatically.
Outcome: The proposed model outperforms baseline models on Chinese, Arabic and Russian loanword detection in Uyghur.
Revisiting Distant Supervision for Relation Extraction (L18-1)

Copied to clipboard

Challenge: Existing approaches for relation extraction (RE) use supervised learning on relation-specific training data, which is expensive to acquire.
Approach: They propose to use a new testing dataset to re-examine distant supervision approaches . they aim to draw new conclusions based on the new testing data .
Outcome: The proposed method can generate training data without noise and bias issues . the proposed method is annotated by the researchers on Amzaon Mechanical Turk .
Incorporating Contextual Information for Language-Independent, Dynamic Disambiguation Tasks (L18-1)

Copied to clipboard

Challenge: a proposed multimodal system can resolve syntactic ambiguities by exploiting external evidence, says a researcher . a parser that processes linguistic information is expected to handle syntakically unambiguous sentences, but it cannot.
Approach: They propose to exploit external contextual information to resolve ambiguous sentences . they propose to use data-driven and grammar-based approaches to solve ambiguities .
Outcome: The proposed system confirms this hypothesis in experiments on syntactically ambiguous sentences.
Overcoming the Long Tail Problem: A Case Study on CO2-Footprint Estimation of Recipes using Information Retrieval (L18-1)

Copied to clipboard

Challenge: a particular challenge is the "long tail problem" that arises with the large diversity of possible ingredients.
Approach: They propose methods that use information retrieval methods for automatic calculation of CO2-footprints of cooking recipes.
Outcome: The proposed methods are generalizable to other use cases where a numerical value has to be calculated based on a list of textual elements.
Comparison of Pun Detection Methods Using Japanese Pun Corpus (L18-1)

Copied to clipboard

Challenge: A sampling survey of typology and component ratio analysis in Japanese puns revealed that the type of Japanese pun that had the largest proportion was a pun type with two sound sequences.
Approach: They propose a method to detect phonetically similar Japanese puns using phonological similarity and insertion / omission of prolonged sounds in addition to lexical feature.
Outcome: The proposed method is validated by adding the rule-based features to the baseline.
A vision-grounded dataset for predicting typical locations for verbs (L18-1)

Copied to clipboard

Challenge: Existing models for inferring location from text are often underestimating the probability of the most typical role fillers.
Approach: They propose a dataset which contains thematic fit judgments for 2,000 verb/location pairs.
Outcome: The proposed dataset can be used to evaluate text-based, vision-based or multimodal inference systems for the typicality of an event's location.
Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method (L18-1)

Copied to clipboard

Challenge: a mixed corpus composed of different dialects is sufficiently resourced to cluster them into dialects.
Approach: They propose a pipeline to derive clusters of dialects from a mixed corpus when their standard counterpart is sufficiently resourced.
Outcome: The proposed pipeline can identify dialectal content when its standard counterpart is sufficiently resourced and can then cluster it into four dialects.
A Fast and Flexible Webinterface for Dialect Research in the Low Countries (L18-1)

Copied to clipboard

Challenge: e-WBD and eWLD are webportals with search applications built to make the data from the Dictionary of the Brabantic and 39 volumes of the Limburgian accessible and retrievable.
Approach: This paper describes the development of webportals with search applications built to make the data accessible and retrievable for both the research community and general audience.
Outcome: The e-WBD and eWLD webportals are being defined in more detail.
Arabic Dialect Identification in the Context of Bivalency and Code-Switching (L18-1)

Copied to clipboard

Challenge: Existing methods for identifying Arabic dialects require significant amounts of annotated training data which is costly and time consuming to produce.
Approach: They propose a novel approach to Arabic dialect identification using language bivalency and written code-switching to identify Arabic dialects.
Outcome: The proposed method can reach more than 76% and score well (66%) when tested on unseen data.
Unified Guidelines and Resources for Arabic Dialect Orthography (L18-1)

Copied to clipboard

Challenge: Existing efforts to conventionalize the dialectal orthography of Arabic have focused on specific dialects and made ad hoc decisions.
Approach: They propose a set of guidelines and meta-guidelines for conventional orthography of Arabic dialects . they apply them to 28 Arab city dialects from Rabat to Muscat .
Outcome: The proposed guidelines and resources are being used by three large Arabic dialect processing projects in three universities.
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach (L18-1)

Copied to clipboard

Challenge: Automatic identification of Arabic dialects in texts is difficult, especially for Maghreb languages and when they are written in Arabic or Latin characters (Arabizi).
Approach: They propose a dictionary-based approach to detect Arabic dialects in texts . they focus on transliteration of Arabicizi into Latin script and code-switching .
Outcome: The proposed approach shows that it is possible to detect dialects in Arabic and Latin scripts.
Shami: A Corpus of Levantine Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world .
Approach: They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools.
Outcome: The proposed corpus is larger than existing corpora in terms of size, words and vocabularies.
You Tweet What You Speak: A City-Level Dataset of Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited.
Approach: They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects.
Outcome: The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects.
Visualizing the “Dictionary of Regionalisms of France” (DRF) (L18-1)

Copied to clipboard

Challenge: a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France.
Approach: They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France.
Outcome: The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas .
DART: A Large Dataset of Dialectal Arabic Tweets (L18-1)

Copied to clipboard

Challenge: The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic.
Approach: They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available.
Outcome: The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi.
Classification of Closely Related Sub-dialects of Arabic Using Support-Vector Machines (L18-1)

Copied to clipboard

Challenge: Existing studies on dialect identification have focused on binary classifications between colloquial Arabic and dialectal Egyptian .
Approach: They propose to use an n-gram based SVM to classify on a fine-grained sub-dialectal level and compare it to methods used in dialect classification such as vocabulary pruning.
Outcome: The proposed method is compared to methods used in dialect classification such as vocabulary pruning of shared items across dialects.
Page Stream Segmentation with Convolutional Neural Nets Combining Textual and Visual Features (L18-1)

Copied to clipboard

Challenge: (retro-)digitizing paper-based files is a major undertaking for private and public archives and an important task in electronic mailroom applications.
Approach: They propose to use convolutional neural networks to combine image and text features to achieve optimal document separation.
Outcome: The proposed approach achieves an accuracy of 93 % and is considered a state-of-the-art for this task.
Automating Document Discovery in the Systematic Review Process: How to Use Chaff to Extract Wheat (L18-1)

Copied to clipboard

Challenge: Systematic reviews address research questions by comprehensively examining the entire published literature.
Approach: They compare the impact of different schemes for choosing positive and negative examples from the different screening stages on the training of automated systems.
Outcome: The proposed ranking system achieves an AUC of 0.803 and 0.768 when relying on gold standard decisions based on title and abstracts of articles, and an AUT of 0.625 and 0.839 when based upon gold standard decision based in full text.
Two Multilingual Corpora Extracted from the Tenders Electronic Daily for Machine Learning and Machine Translation Applications. (L18-1)

Copied to clipboard

Challenge: European "Tenders Electronic Daily" is a valuable source of semi-structured and multilingual data . collecting and managing such kind of data is incredibly burdensome and takes time and resources .
Approach: They describe two documented and easy-to-use multilingual corpora extracted from the TED web site . they propose to make the extracted dataset available to the scientific community .
Outcome: The proposed dataset is based on the European tenders electronic daily (TED) web site . it is easy to use and can be used for text mining and natural language processing tasks.
Using Adversarial Examples in Natural Language Processing (L18-1)

Copied to clipboard

Challenge: Recent advances in machine learning have led to the use of adversarial examples in training of neural networks.
Approach: They investigate the effect of using adversarial examples during training of recurrent neural networks whose text input is in the form of a sequence of word/character embeddings.
Outcome: The proposed method provides regularization effect and enables training of models with greater number of parameters without overfitting.
Modeling Trolling in Social Media Conversations (L18-1)

Copied to clipboard

Challenge: a new classification of trolling allows for comment-based analysis from both the trolls' and the responders' perspectives . a trolled's intentions may cause a negative psychological impact on the participants .
Approach: They propose a trolling categorization that allows comment-based analysis from both trolls' and responders' perspectives . they annotate and release a dataset containing excerpts of Reddit conversations involving suspected trolled users .
Outcome: The proposed model allows comment-based analysis from both the trolls' and the responders' perspectives.
Automatic Annotation of Semantic Term Types in the Complete ACL Anthology Reference Corpus (L18-1)

Copied to clipboard

Challenge: a recent increase in quantitative studies of scientific text collections has led to a significant increase in the use of semantic labeling techniques.
Approach: They propose to use semantic class labels to enhance a well-known resource . they use semantic labels to assign semantic class labeling to technical terms .
Outcome: The proposed approach enhances the ACL Anthology Reference Corpus with semantic class labels for 20,000 technical terms . the goal is to use this information as one feature in the profiling of scientific papers, communities, and disciplines.
Annotated Corpus of Scientific Conference’s Homepages for Information Extraction (L18-1)

Copied to clipboard

Challenge: a corpus of scientific conferences contains homepages with annotations of important information . name of conference, abbreviation, place, submission, notification, camera ready dates are included .
Approach: They propose a corpus that contains 943 homepages of scientific conferences with annotations of interesting information.
Outcome: The proposed corpus contains 943 homepages of scientific conferences, 14794 including subpages . the results show that it can be used as a reference data set for this type of task.
Improving Unsupervised Keyphrase Extraction using Background Knowledge (L18-1)

Copied to clipboard

Challenge: Existing methods of keyphrase extraction are supervised and unsupervised . Topical PageRank uses topical information to extract the top topics of a document .
Approach: They propose an unsupervised method for keyphrase extraction based on Wikipedia . they construct a semantic graph and transform the extraction problem into an optimization problem .
Outcome: The proposed method improves over other state-of-the-art models by more than 2% in F1-score.
WikiDragon: A Java Framework For Diachronic Content And Network Analysis Of MediaWikis (L18-1)

Copied to clipboard

Challenge: WikiDragon is a Java Framework designed to give developers in computational linguistics an intuitive API to build, parse and analyze instances of MediaWikis.
Approach: They introduce WikiDragon, a Java Framework that allows developers to build, parse and analyze instances of MediaWikis on their computers.
Outcome: The framework is based on the Wikipedia, Wiktionary, WikiSource or WikiNews and evaluates link extraction, diachronic network analysis and the impact of different frameworks to text analysis.
Studying Muslim Stereotyping through Microportrait Extraction (L18-1)

Copied to clipboard

Challenge: Research shows that stereotypical ideas are often reflected in language use.
Approach: They propose to use microportraits to investigate stereotyping in the media to explore various dimensions of stereotypation.
Outcome: The proposed system allows social scientists to explore various dimensions of stereotyping compared to more basic models such as word clouds.
Analyzing the Quality of Counseling Conversations: the Tell-Tale Signs of High-quality Counseling (L18-1)

Copied to clipboard

Challenge: Behavioral and mental health disorders are the most costly and prevalent conditions worldwide.
Approach: They propose to use a dataset to analyze counseling interactions by using aspects such as mirroring, empathy, and reflective listening to build text-based classifiers.
Outcome: The proposed dataset can be used to build text-based classifiers able to predict the overall quality of a counseling conversation and provide insights into the linguistic differences between low-quality and high-quality counseling.
Interpersonal Relationship Labels for the CALLHOME Corpus (L18-1)

Copied to clipboard

Challenge: a lack of corpora makes exploration of this problem intractable, says nicolaus mills . mills: communication is one of the most invaluable tools humans have .
Approach: a new study uses a corpus of interpersonal relationship labels to help identify relationships . a set of labels is available for download on the website of the cnn.org team .
Outcome: a new set of interpersonal relationship labels is released for the CALLHOME English corpus . the labels are available for download on the cnn.com website .
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
Training and Adapting Multilingual NMT for Less-resourced and Morphologically Rich Languages (L18-1)

Copied to clipboard

Challenge: Using multilingual and multi-way neural machine translation approaches is a major advantage . training NMT systems for individual language pairs takes significantly more time than training of SMT systems .
Approach: They propose to employ multilingual and multi-way neural machine translation approaches for morphologically rich languages such as Estonian and Russian.
Outcome: The proposed approach improves translation quality by +3.27 BLEU points over baseline models.
Cross-lingual Terminology Extraction for Translation Quality Estimation (L18-1)

Copied to clipboard

Challenge: Using common statistical measures for termhood and unithood, we identify terms from monolingual texts and investigate the contribution of terminology to translation quality.
Approach: They propose to use common statistical measures for termhood and unithood as features to train classifiers for identifying terms in cross-domain and cross-language settings.
Outcome: The proposed method has shown some reliability in automatically identifying terms in human translations, but drawbacks in handling low frequency terms and term variations shall be dealt with in the future.
Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German (L18-1)

Copied to clipboard

Challenge: Using character-based neural MT, we normalize Swiss German input to address regional diversity.
Approach: They propose to use character-based neural MT to normalize Swiss German input and phrase-based statistical MT for a low-resource family of dialects.
Outcome: The proposed system achieves 36% BLEU score when translating from the Bernese dialect.
Improving domain-specific SMT for low-resourced languages using data from different domains (L18-1)

Copied to clipboard

Challenge: Evaluation of domain-specific statistical machine translation system for official government letters . use of pseudo in-domain data showed improvement for both test sets .
Approach: They develop a statistical machine translation system for official government letters . the system is based on a parallel in-domain dataset containing official letters based in Sinhala and Tamil .
Outcome: The proposed system improves on the in-domain data in the domain of official government letters . the evaluations show that the system requires quality data from diverse subject matters and sources to perform better.
Discovering Parallel Language Resources for Training MT Engines (L18-1)

Copied to clipboard

Challenge: Web crawling is an efficient way for compiling the monolingual, parallel and/or domain-specific corpora needed for machine translation and other HLT applications.
Approach: They propose a system for compiling monolingual, parallel and/or domain-specific corpora . ILSP-FC is a web crawling system that generates bilingual lexica and terminology lists .
Outcome: The ILSP Focused Crawler is a system developed by researchers at the IL SP/Athena RIC for the acquisition of such resources.
A fine-grained error analysis of NMT, SMT and RBMT output for English-to-Dutch (L18-1)

Copied to clipboard

Challenge: Since 2016, the landscape of automated translation has substantially changed with the arrival of neural machine translation (NMT).
Approach: They propose to use an annotated SCATE corpus of MT errors to enrich the SCATE error taxonomy to fit the neural MT output.
Outcome: The proposed system outperforms phrase-based and rule-based systems except for lexical issues.
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)

Copied to clipboard

Challenge: despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech.
Approach: They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective.
Outcome: The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective.
Multimodal Lexical Translation (L18-1)

Copied to clipboard

Challenge: Multimodal Lexical Translation (MLT) is a task that aims to translate ambiguous words given their context -an image and a sentence in the source language.
Approach: They introduce a task to translate an ambiguous word given its context -an image and a sentence in the source language.
Outcome: The proposed task is based on the Multi30K dataset and uses word-alignment followed by human inspection to select subsets of the dataset which are difficult to translate.
Literality and cognitive effort: Japanese and Spanish (L18-1)

Copied to clipboard

Challenge: pause-word ratios are indicators of cognitive effort during different translation modalities.
Approach: They propose a notion of pause-word ratio computed using ranges of a pause length rather than lower cutoffs for pauses . they compare translation and post-editing for language pairs that are different in terms of semantic and syntactic remoteness .
Outcome: The proposed pause-word ratio measures cognitive effort in translation and post-editing for language pairs that are different in terms of semantic and syntactic remoteness.
Evaluation of Machine Translation Performance Across Multiple Genres and Languages (L18-1)

Copied to clipboard

Challenge: In this paper, we evaluate the impact of genre differences on machine translation (MT) for a diverse set of language pairs . BLEU score differences between genres can be large for all genres and all language pairs.
Approach: They use multi-genre benchmarks to evaluate the impact of genre differences on machine translation (MT) they train and use genre classifiers to route test documents to the most appropriate genre systems .
Outcome: The proposed system can improve translation quality for all genres and language pairs .
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
Manual vs Automatic Bitext Extraction (L18-1)

Copied to clipboard

Challenge: targeted, site-specific crawling results in cleaner bitexts with a higher ratio of parallel sentences . general crawlers combined with boilerplate removal tools tend to retrieve shorter texts .
Approach: They compare manual and automatic approaches to extracting bitexts from the Web . they use targeted site-specific crawling to extract cleaner bitext sentences .
Outcome: The proposed methods extract more parallel sentences from the Web than manual methods.
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)

Copied to clipboard

Challenge: Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA).
Approach: They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use.
Outcome: The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety.
CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing (L18-1)

Copied to clipboard

Challenge: Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks.
Approach: They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages.
Outcome: The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks.
Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)

Copied to clipboard

Challenge: a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Approach: The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Outcome: The corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Evaluating Inflectional Complexity Crosslinguistically: a Processing Perspective (L18-1)

Copied to clipboard

Challenge: a cognitively motivated method for evaluating the inflectional complexity of a language is proposed . authors argue that some languages are inflectionally more complex than others .
Approach: They propose a cognitively motivated method for evaluating inflectional complexity of a language . they use a recurrent self-organising neural network to learn "raw" inflected word forms .
Outcome: The proposed method is independent of meta-linguistic issues and language-specific typological aspects.
Parser combinators for Tigrinya and Oromo morphology (L18-1)

Copied to clipboard

Challenge: morphological parsers for two Afroasiatic languages are developed using a parser-combinator paradigm . the paradigm allows rapid development and ease of integration with other systems, but at a cost of non-optimal theoretical efficiency.
Approach: They propose a rule-based morphological parser paradigm for Tigrinya and Oromo languages . they use a parsers-combinator paradigm instead of a finite-state paradigm .
Outcome: The proposed paradigm allows rapid development and ease of integration with other systems, but at cost of non-optimal theoretical efficiency.
Massively Translingual Compound Analysis and Translation Discovery (L18-1)

Copied to clipboard

Challenge: Morphological compounding is one of the most common and productive methods of word formation across the world's languages.
Approach: They propose a model for compounding using bilingual dictionaries and no annotated training data . they also release a massively multilingual dataset of compound words and their decompositions .
Outcome: The proposed model generates novel translations of English concepts on a multilingual dataset . the model can be applied to a wide range of languages and is highly reproducible.
Building a Morphological Treebank for German from a Linguistic Database (L18-1)

Copied to clipboard

Challenge: German is a language with complex morphological processes.
Approach: They propose a morphological treebank for German based on a German morphology database and a Perl script for the generation.
Outcome: The proposed treebank is based on the German lexical database CELEX and is able to generate 40,000 morphological trees with a grade of detail that can be chosen according to the requirements of the applications.
Baselines and Test Data for Cross-Lingual Inference (L18-1)

Copied to clipboard

Challenge: Recent research on textual entailment is limited to English, but it is expanding to other languages.
Approach: They propose to extend the research in SNLI-style natural language inference toward multilingual evaluation by using cross-lingual word embeddings and machine translation.
Outcome: The proposed system scores an average accuracy of just over 75%, but it is not perfect.
CATS: A Tool for Customized Alignment of Text Simplification Corpora (L18-1)

Copied to clipboard

Challenge: Existing corpora of original sentences and their manual simplifications are very scarce and small in size, hindering automated text simplification systems.
Approach: They propose a language-independent tool for sentence alignment from parallel/comparable TS resources.
Outcome: The proposed tool performs well on English and Spanish corpora and compares sentences based on their semantic overlap.
KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus (L18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are representations of words across languages in a shared continuous vector space.
Approach: They propose a multilingual word embedding corpus which is acquired by neural machine translation and is based on monolingual data.
Outcome: The proposed method is competitive with existing methods but on the cross-lingual document classification task, it obtains the best figures.
Multi-lingual Argumentative Corpora in English, Turkish, Greek, Albanian, Croatian, Serbian, Macedonian, Bulgarian, Romanian and Arabic (L18-1)

Copied to clipboard

Challenge: Argumentative corpora are costly to create and available only in few languages with English dominating the area.
Approach: They use 8 different argument mining classifiers trained for English to build a parallel corpora in which the source language is English and the target language is either a Balkan language or Arabic.
Outcome: The proposed method is based on 8 different argument mining classifiers trained for English and project the decision to the target language.
SemR-11: A Multi-Lingual Gold-Standard for Semantic Similarity and Relatedness for Eleven Languages (L18-1)

Copied to clipboard

Challenge: SemR-11 is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Approach: This paper describes a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Outcome: The dataset is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)

Copied to clipboard

Challenge: RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard.
Approach: They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard.
Outcome: The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages.
Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM (L18-1)

Copied to clipboard

Challenge: Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications.
Approach: They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy.
Outcome: The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler.
Web-based Annotation Tool for Inflectional Language Resources (L18-1)

Copied to clipboard

Challenge: Wasim is a web-based tool for semi-automatic morphosyntactic annotation of inflectional languages.
Approach: They present a web-based tool for semi-automatic morphosyntactic annotation of inflectional languages resources.
Outcome: The tool has high flexibility in segmenting tokens, editing, diacritizing, labelling tokens and segments.
HiNTS: A Tagset for Middle Low German (L18-1)

Copied to clipboard

Challenge: a non-standardized language such as Middle Low German has special requirements for annotating part of speech and morphology.
Approach: They describe a tagset for annotating parts-of-speech and morphology in Middle Low German texts . they describe two special features of the tagse, and prove their usefulness .
Outcome: The proposed tagset can be used to annotate parts-of-speech and morphology in Middle Low German texts.
Leveraging Lexical Resources and Constraint Grammar for Rule-Based Part-of-Speech Tagging in Welsh (L18-1)

Copied to clipboard

Challenge: POS tags are based on pre-annotated text, but there is not enough data to train a statistical POS tagger in lesser-resourced languages such as Welsh.
Approach: They propose a rule-based POS tagger for Welsh based on the VISL Constraint Grammar parser and extract a list of possible POS tags for each word token in a running text.
Outcome: The proposed approach is particularly useful in dealing with some of the specific intricacies of Welsh, such as morphological changes and word mutations.
Graph Based Semi-Supervised Learning Approach for Tamil POS tagging (L18-1)

Copied to clipboard

Challenge: Parts of Speech (POS) tagging is challenging for low resourced languages such as Tamil . low resource Tamil does not have large POS annotated corpus to build good quality POS taggers using supervised machine learning techniques.
Approach: They propose a graph-based semi-supervised learning approach to classify unlabelled data using a small POS labelled data set.
Outcome: The proposed method achieves 0.8743 over 0.7333 produced by a CRF tagger for the same limited size corpus.
What Causes the Differences in Communication Styles? A Multicultural Study on Directness and Elaborateness (L18-1)

Copied to clipboard

Challenge: Using a multi-cultural approach, we investigated the differences in the communication styles elaborateness and directness of human-computer interaction.
Approach: They propose to design a Spoken Dialogue System which adapts to the user's communication idiosyncrasies and to examine the influence of the user culture and gender on the system's elaborateness and directness.
Outcome: The proposed system could be used to communicate with computers in a human-computer interaction.
FARMI: A FrAmework for Recording Multi-Modal Interactions (L18-1)

Copied to clipboard

Challenge: a new framework for recording multi-modal data is needed to capture multi-party, richly recorded corpora and perform real-time processing of such data.
Approach: They propose an open-source processing architecture for corpora and real-time processing . they deploy the architecture in a multi-party deception game with six humans and one robot .
Outcome: The proposed architecture is agnostic to hardware and programming languages, although it's mostly written in Python.
Creating Large-Scale Argumentation Structures for Dialogue Systems (L18-1)

Copied to clipboard

Challenge: Argumentation is a process of reaching consensus through premises and rebuttals and is important for making decisions and exchanging views.
Approach: They propose to create argumentation structures in ten languages using argumentation databases . they also examine differences between the two languages to determine their effectiveness .
Outcome: The proposed arguments can be applied to argumentative dialogue systems and can be used as training data.
Exploring Conversational Language Generation for Rich Content about Hotels (L18-1)

Copied to clipboard

Challenge: a new method is needed to generate natural dialogues for hotel information . a recent study shows that hotel descriptions are not a good match for conversational interaction .
Approach: They propose to use stylistic features to generate and score hotel dialogues from hotel descriptions . they use hotel descriptions written by human writers within Google Content Studio .
Outcome: The proposed models can be used to generate natural dialogues for hotels . the authors show that the sentences in the original written hotel descriptions are not a good match for conversational interaction.
Identification of Personal Information Shared in Chat-Oriented Dialogue (L18-1)

Copied to clipboard

Challenge: An annotation scheme is developed to annotate the presence and type of personal information in chat-oriented dialogues.
Approach: They propose an annotation scheme that can be used to annotate the presence and type of personal information in chat-oriented dialogues.
Outcome: The proposed annotation scheme can be used to annotate the presence and type of personal information in chat-oriented dialogues.
A Vietnamese Dialog Act Corpus Based on ISO 24617-2 standard (L18-1)

Copied to clipboard

Challenge: standardized dialog act corpora are used for conversation mining research . different corporations often use different methods to understand interaction structure .
Approach: They propose to annotate dialog acts using ISO 24617-2 standard (2012) . they also annotated emotions using Ekman's six primitives and sentiment using tags "positive", "negative" and "neutral"
Outcome: The proposed corpus is constructed using the ISO 24617-2 standard (2012) . it is used for emotions, sentiment and positive, negative and neutral tags .
Annotating Reflections for Health Behavior Change Therapy (L18-1)

Copied to clipboard

Challenge: Existing studies show that depression can be treated by Motivational Interviewing (MI)
Approach: They annotated reflections, an essential counselor behavioral code in motivational interviewing for psychotherapy on conversations that are a combination of casual and therapeutic dialogue.
Outcome: The annotated transcripts are a vital resource for automated health behavior change therapy . the corpus is being constructed and annotating conversations by one annotator .
Annotating Attribution Relations in Arabic (L18-1)

Copied to clipboard

Challenge: Current studies focus on using lexical terms in long texts to verify author identity.
Approach: They propose to annotate attributed arguments to the source in Arabic news with required syntactical and semantic features with required features.
Outcome: The proposed method is applied to Arabic news and is compared with existing tools and methods.
The ADELE Corpus of Dyadic Social Text Conversations:Dialog Act Annotation with ISO 24617-2 (L18-1)

Copied to clipboard

Challenge: Recent studies have focused on task-based or instrumental dialogs, but there is increasing interest in social or interactional dialogs.
Approach: They describe a corpus of 193 dyadic text dialogs based on a novel 'getting to know you' social dialog elicitation paradigm and propose additional acts to better cover greeting and leavetaking.
Outcome: The proposed actions cover greeting and leavetaking, and the proposed acts improve the interaction between the dialogs and spoken language.
An Assessment of Explicit Inter- and Intra-sentential Discourse Connectives in Turkish Discourse Bank (L18-1)

Copied to clipboard

Challenge: Discourse parsing is a challenging task for NLP.
Approach: They propose to add a new set of explicit intra-sentential connectives to Turkish Discourse Bank 1.1 . they propose to evaluate the converb sense annotations and compare them to other Turkish corpus .
Outcome: The proposed annotations show that the subordinators tend to select certain senses not selected by explicit inter- and intra-sentential discourse connectives in the data.
Compilation of Corpora for the Study of the Information Structure–Prosody Interface (L18-1)

Copied to clipboard

Challenge: empirical studies on the Information Structure-prosody interface are scarce . thematicity defines how content is packaged in terms of "what is being talked about" a different view on thematicality is advocated by I. Mel'uk in the context of the MTT.
Approach: They propose a method for the compilation of annotated corpora to study the correspondence between Information Structure and prosody.
Outcome: The proposed method is applied to a corpus of read speech in English annotated with hierarchical thematicity and automatically extracted prosodic parameters.
Preliminary Analysis of Embodied Interactions between Science Communicators and Visitors Based on a Multimodal Corpus of Japanese Conversations in a Science Museum (L18-1)

Copied to clipboard

Challenge: a preliminary analysis of embodied interactions is based on a multimodal corpus of Japanese conversations . a corpus recording in a specific social setting is needed to analyze language use and non-verbal behaviors in situated activities.
Approach: They propose to draw on a multimodal corpus of Japanese conversations recorded at a museum in Tokyo . preliminary analyses show that science communicators are context-free and context-sensitive .
Outcome: The proposed method shows that science communicators are context-free and context-sensitive . it shows that the practices of science communiators are adapted to different situations .
Improving Crowdsourcing-Based Annotation of Japanese Discourse Relations (L18-1)

Copied to clipboard

Challenge: Discourse parsing is an important task in natural language processing, but few languages have corpora annotated with discourse relations . crowdsourcing-based annotations are of poor quality and require expensive and time-consuming . et al. (2009) evaluated the quality of annotations using expert annotations.
Approach: They construct a Japanese corpus with discourse annotations through crowdsourcing . they propose improvement techniques based on language tests .
Outcome: The proposed methods improve the quality of the annotations, and will make them publicly available.
Persian Discourse Treebank and coreference corpus (L18-1)

Copied to clipboard

Challenge: Currently, we are adding a new document-level discourse annotation to our new corpus.
Approach: They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution.
Outcome: The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens.
Automatic Labeling of Problem-Solving Dialogues for Computational Microgenetic Learning Analytics (L18-1)

Copied to clipboard

Challenge: This paper presents a recurrent neural network model to automate the analysis of students' computational thinking in problem-solving dialogue.
Approach: They propose a recurrent neural network model to automate the analysis of students' computational thinking in problem-solving dialogue.
Outcome: The proposed model outperforms the baseline model and outperformed the nave model by a large margin.
Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)

Copied to clipboard

Challenge: Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer.
Approach: They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines.
Outcome: The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures.
Semi-Supervised Clustering for Short Answer Scoring (L18-1)

Copied to clipboard

Challenge: Existing approaches to SAS use unsupervised clustering and have teachers label some items after clustering.
Approach: They propose to use semi-supervised clustering to provide structured groups of answers in addition to a score.
Outcome: The proposed method improves clustering performance from 0.504 kappa for unsupervised clustering to 0.566 kppa.
Analyzing Vocabulary Commonality Index Using Large-scaled Database of Child Language Development (L18-1)

Copied to clipboard

Challenge: a vocabulary commonality index is used to investigate to what extent each child acquires common words during the early stages of lexical development.
Approach: They propose a vocabulary commonality index to investigate to what extent each child acquires common words during the early stages of lexical development.
Outcome: The proposed index can be used to understand to what extent each child acquires common words during the early stages of lexical development.
The ICoN Corpus of Academic Written Italian (L1 and L2) (L18-1)

Copied to clipboard

Challenge: a corpus of academic written Italian is described in this paper . the corpus includes 2,115,000 tokens written by students having Italian as L2 .
Approach: They describe the ICoN corpus, a corpus of academic written Italian . it includes 2,115,000 tokens written by students having Italian as L2 and 1,769,000 tokens by students with Italian as a L1 .
Outcome: The ICoN corpus includes 2,115,000 tokens written by students having Italian as L2 and 1,769,000 tokens by students with Italian as a L1 . the corpus can be queried online while its complete contents are available on request for research purposes.
Revita: a Language-learning Platform at the Intersection of ITS and CALL (L18-1)

Copied to clipboard

Challenge: Existing language-learning tools do not address the fundamental requirements of language learners and teachers.
Approach: They propose a free-to-use platform for language learning beyond the beginner level . they outline the established desiderata of CALL and ITS .
Outcome: The proposed platform supports language learning beyond the beginner level.
The Distribution and Prosodic Realization of Verb Forms in German Infant-Directed Speech (L18-1)

Copied to clipboard

Challenge: Infant-directed speech is often seen as a predictor for infants' speech processing abilities, for instance speech segmentation or word learning.
Approach: They examine the syntactic distribution, accentuation and prosodic phrasing of German verb forms and show that many verb forms are prime candidates for early segmentation.
Outcome: The findings suggest that infants ought to be able to extract verbs as early as nouns, given appropriate stimulus materials.
Cross-linguistically Small World Networks are Ubiquitous in Child-directed Speech (L18-1)

Copied to clipboard

Challenge: In this paper we use network theory to model graphs of child-directed speech from caregivers of children from nine typologically and morphologically diverse languages.
Approach: They use network theory to model child-directed speech from caregivers of children from nine typologically and morphologically diverse languages.
Outcome: The proposed model adds to the repertoire of universal distributional patterns found in the input to children cross-linguistically.
L1-L2 Parallel Treebank of Learner Chinese: Overused and Underused Syntactic Structures (L18-1)

Copied to clipboard

Challenge: Currently, the treebank consists of 600 L2 sentences and 697 L1 sentences.
Approach: They propose to use "L1-L2 parallel treebanks" to facilitate analyses of learner language.
Outcome: The proposed treebank consists of 600 L2 sentences and 697 L1 sentences.
The Use of Text Alignment in Semi-Automatic Error Analysis: Use Case in the Development of the Corpus of the Latvian Language Learners (L18-1)

Copied to clipboard

Challenge: Using error annotation methods, the corpus of the Latvian language learners can be adapted for other languages with relatively free word order.
Approach: They propose a method for creating error annotated corpora using text correction, automated morphological analysis, automated text alignment and error annotation.
Outcome: The proposed method has been approbated in the development of the corpus of the Latvian language learners.
Error annotation in a Learner Corpus of Portuguese (L18-1)

Copied to clipboard

Challenge: Using the corpus architecture and the TEITOK platform, error tagging is a time-consuming task that has to be performed manually.
Approach: They propose a system that produces a final standoff, multilevel annotation with position-based tags that account for the main error types observed in the corpus.
Outcome: The proposed system annotates 47% of the corpus using the COPLE2 architecture and the TEITOK platform.
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)

Copied to clipboard

Challenge: a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures .
Approach: They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners .
Outcome: The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied .
Portable Spelling Corrector for a Less-Resourced Language: Amharic (L18-1)

Copied to clipboard

Challenge: a corpus-driven spelling corrector for Amharic is ported to other languages with little effort . a term list is used for spelling errors and can handle rare terms, proper nouns and neologisms.
Approach: They propose an automatic spelling corrector for Amharic, the working language of the Federal Government of Ethiopia.
Outcome: The proposed method outperforms baseline systems in Amharic and English . it has smoothed language model, generalized error model and ability to take into account context of misspellings.
A Speaking Atlas of the Regional Languages of France (L18-1)

Copied to clipboard

Challenge: a website is presented to show and promote the linguistic diversity of France through field recordings, a computer program and an orthographic transcription.
Approach: They propose to map linguistic diversity in France using field recordings and a computer program.
Outcome: The aim is to show and promote the linguistic diversity of France, through field recordings, a computer program and an orthographic transcription.
Towards Language Technology for Mi’kmaq (L18-1)

Copied to clipboard

Challenge: Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada .
Approach: They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages .
Outcome: The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue .
Pronunciation Dictionaries for the Alsatian Dialects to Analyze Spelling and Phonetic Variation (L18-1)

Copied to clipboard

Challenge: a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century .
Approach: They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French .
Outcome: The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing.
ChAnot: An Intelligent Annotation Tool for Indigenous and Highly Agglutinative Languages in Peru (L18-1)

Copied to clipboard

Challenge: Linguistic corpus annotation is one of the most important phases for addressing natural language processing (NLP) tasks.
Approach: They propose a web-based annotation tool for Peruvian indigenous and highly agglutinative languages that supports a variety of linguistic annotation tasks.
Outcome: The proposed tool supports a diverse set of linguistic annotation tasks, such as morphological segmentation markup, POS-tag markup and other.
The DLDP Survey on Digital Use and Usability of EU Regional and Minority Languages (L18-1)

Copied to clipboard

Challenge: the survey was launched by the Digital Language Diversity Project to investigate the real usage, needs and expectations of European minority language speakers regarding digital opportunities.
Approach: This paper reports on the results of an exploratory survey launched by the Digital Language Diversity Project about the digital use and usability of regional and minority languages on digital media and devices.
Outcome: The findings of the first exploratory survey on the use and usability of regional and minority languages on digital media and devices are presented in this paper.
ASR for Documenting Acutely Under-Resourced Indigenous Languages (L18-1)

Copied to clipboard

Challenge: Automatic speech recognition (ASR) has not been widely explored as a tool for documenting endangered languages.
Approach: They propose to use automatic speech recognition (ASR) to bootstrap new data to improve the acoustic model.
Outcome: The proposed system improves the model for a polysynthetic language with few audio and text resources.
Building a Sentiment Corpus of Tweets in Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: Sentiment analysis is a popular area of Natural Language Processing due to its subjective and semantic characteristics.
Approach: They propose to annotate Brazilian Portuguese sentences manually using a sentiment corpus . they run experiments on polarity classification using six machine learning classifiers .
Outcome: The proposed method is based on a Brazilian Portuguese sentiment corpus and achieved 80.38% on F-Measure and 64.87% when including the neutral class.
‘Aye’ or ‘No’? Speech-level Sentiment Analysis of Hansard UK Parliamentary Debate Transcripts (L18-1)

Copied to clipboard

Challenge: Transcripts of UK parliamentary debates are difficult for human readers to process due to the large quantity of textual data and the specialised language used.
Approach: They propose to use annotated sentiment labels and labels derived from speakers' votes to classify the sentiment polarity of speakers as being either positive or negative towards motions proposed in the debates.
Outcome: The proposed model outperforms existing models on a dataset of parliamentary debate transcripts using textual and contextual features.
Scalable Visualisation of Sentiment and Stance (L18-1)

Copied to clipboard

Challenge: a novel visualisation approach for sentiment and stance analysis is proposed for large datasets.
Approach: They propose a visualisation approach for scalable visualisation of sentiment and stance from large-scale data.
Outcome: The proposed visualisation approach can be used on a 9,278 user comments with stance explicitly declared by the author.
NoReC: The Norwegian Review Corpus (L18-1)

Copied to clipboard

Challenge: The Norwegian Review Corpus is a dataset of full-text reviews from major news sources.
Approach: This paper presents the Norwegian Review Corpus, created for document-level sentiment analysis.
Outcome: The corpus comprises more than 35,000 full-text reviews from a range of different domains.
SenSALDO: Creating a Sentiment Lexicon for Swedish (L18-1)

Copied to clipboard

Challenge: sentiment analysis has seen an explosive expansion over the last decade or so . many theoretical and methodological questions remain unanswered and resource gaps unfilled .
Approach: They develop a sentiment lexicon for written (standard) Swedish using an existing dataset . they assign a real value sentiment score in the range [-1,1] and produce a label for it .
Outcome: The proposed sentiment lexicon is an open source resource from the Swedish Language Bank . it is based on an existing gold standard dataset and is available from Sprkbanken .
Corpus Building and Evaluation of Aspect-based Opinion Summaries from Tweets in Spanish (L18-1)

Copied to clipboard

Challenge: a corpus of Spanish extractive and abstractive summaries of opinions is presented . the goal is to analyze the summary content and to show how different they are written .
Approach: They present a corpus of Spanish extractive and abstractive summaries of opinions . they analyze the summary agreement between them and their aspect coverage and sentiment orientation .
Outcome: The presented corpus of Spanish extractive and abstractive summaries is a reference for academic research.
Application and Analysis of a Multi-layered Scheme for Irony on the Italian Twitter Corpus TWITTIRÒ (L18-1)

Copied to clipboard

Challenge: Using a multi-layered scheme for the fine-grained annotation of irony on Italian Twitter is a challenging task to be performed by both human annotators and automatic NLP systems.
Approach: They propose to apply a multi-layered scheme for the fine-grained annotation of irony to an Italian Twitter corpus.
Outcome: The proposed scheme can be validated on Italian irony-laden social media contents and is available in the cross- and multi-lingual perspective.
Classifier-based Polarity Propagation in a WordNet (L18-1)

Copied to clipboard

Challenge: a wordnet-based sentiment lexicon can be built to express sentiment polarity in a way shared across domains.
Approach: They propose a method to build a sense-level sentiment lexicon on the basis of a wordnet . they use a rich set of wordnet-based features to recognize and assign sentiment polarity values .
Outcome: The proposed method allows for the construction of a more reliable sentiment lexicon . the proposed method is partially automated, but it's performance drops in cross-domain applications .
SMILE Swiss German Sign Language Dataset (L18-1)

Copied to clipboard

Challenge: The goal of an ongoing three-year project in Switzerland is to pioneer an assessment system for lexical signs of Swiss German Sign Language (Deutschschweizerische Gebärdensprache, DSGS) that relies on sign language recognition.
Approach: The goal of the project is to pioneer an assessment system for lexical signs of Swiss German Sign Language that relies on sign language recognition.
Outcome: The system will give adult L2 learners of DSGS feedback on the correctness of the manual parameters (handshape, hand position, location, and movement) of isolated signs they produce.
IPSL: A Database of Iconicity Patterns in Sign Languages. Creation and Use (L18-1)

Copied to clipboard

Challenge: a database of signs annotated according to iconicity parameters was created . the database contains 1542 signs in 19 sign languages .
Approach: a team of researchers has created a large-scale database of iconic signs . the database contains 1542 signs in 19 sign languages .
Outcome: the database contains 1542 sign annotated in 19 sign languages . the database can be used to further study iconicity in sign languages.
Sign Languages and the Online World Online Dictionaries & Lexicostatistics (L18-1)

Copied to clipboard

Challenge: Several online dictionaries documenting the lexicon of a variety of sign languages are available . methodological issues must be addressed regarding how these resources are used for research purposes.
Approach: They propose a web-based tool for annotating the articulatory features of signs . they compare handshapes for four Asian SLs and handshape for the entire sample .
Outcome: The proposed tool compares handshapes and handsights of Asian SLs with European, American, and Brazilian SL samples.
Elicitation protocol and material for a corpus of long prepared monologues in Sign Language (L18-1)

Copied to clipboard

Challenge: elicitation of long discourses is difficult in Sign Language, and is often a problem . e.g., elicitation of long texts is a technique that can be used to collect long discourse .
Approach: They propose a protocol and two tasks to collect long discourse in Sign Language . they propose to ensure both are collected and prepared in the language .
Outcome: The proposed protocol improves the produced data and the results of a test with LSF informants.
Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions (L18-1)

Copied to clipboard

Challenge: Existing technologies for CG-supported data display are not able to depict all relevant features of a natural signing sequence such as facial expression, spatial references or inter-sign movement.
Approach: They collected a corpus of Japanese Sign Language sentences for deep neural network learning.
Outcome: The proposed model could be used to train language features in Japanese Sign Language (JSL)
Modeling French Sign Language: a proposal for a semantically compositional system (L18-1)

Copied to clipboard

Challenge: Several studies have proposed linguistic models to describe sign languages, but none have succeeded to describe the specificities of SL.
Approach: They propose a linguistic approach to formalize the sign language (SL) they propose to take into account linguistic properties of the SL while respecting constraints of a modelisation process.
Outcome: The proposed model takes into account linguistic properties of the sign language while respecting constraints of a modelisation process.
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
Carcinologic Speech Severity Index Project: A Database of Speech Disorder Productions to Assess Quality of Life Related to Speech After Cancer (L18-1)

Copied to clipboard

Challenge: Increasing mortality in cancerology highlights the importance of reducing the impact on the Quality of Life after cancer treatment.
Approach: They collect a large database of french speech recordings aimed at validating Disorder Severity Indexes.
Outcome: The collected data will be available to the scientific community through the GIS Parolotheque.
Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)

Copied to clipboard

Challenge: BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages.
Approach: This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project.
Outcome: The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants.
A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks (L18-1)

Copied to clipboard

Challenge: Phonetic segmentation is the process of splitting speech into distinct phonetic units . methods for automatic segmentation are not always accurate enough .
Approach: They propose to model phonetic segmentation as close as possible to manual segmentation by recording experts performing a segmentation task.
Outcome: This corpus captures human segmentation behavior by recording experts performing a segmentation task.
Statistical Analysis of Missing Translation in Simultaneous Interpretation Using A Large-scale Bilingual Speech Corpus (L18-1)

Copied to clipboard

Challenge: Various types of omissions have been described in simultaneous interpretation to improve interpretation quality or train interpreters.
Approach: They analyze missing translations in simultaneous interpretations using a large-scale bilingual speech corpus.
Outcome: The authors found that a high proportion of adverbs were missed in the translations . the authors suggest that omissions can be improved to improve interpretation quality .
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)

Copied to clipboard

Challenge: a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers .
Approach: They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging.
Outcome: The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis .
Increasing the Accessibility of Time-Aligned Speech Corpora with Spokes Mix (L18-1)

Copied to clipboard

Challenge: Spokes Mix is an online service providing access to spoken corpora of Polish . high-quality corporata of conversational language are expensive to acquire .
Approach: a new online service provides access to spoken corpora of Polish . the service provides a centralized, easy-to-use corpus query engine with a responsive web interface .
Outcome: the proposed service provides access to spoken corpora of Polish, including three newly released time-aligned collections of manually transcribed spoken-conversational data.
The MonPaGe_HA Database for the Documentation of Spoken French Throughout Adulthood (L18-1)

Copied to clipboard

Challenge: Existing studies on life-span changes in the speech of adults are mainly based on English speakers and few studies have compared more than two extreme age groups.
Approach: They describe a MonPaGe_HealthyAdults database of spoken french with 405 speakers aged from 20 to 93 years old.
Outcome: The proposed database includes 405 speakers aged 20 to 93 years old and includes 4 regiolects.
Bringing Order to Chaos: A Non-Sequential Approach for Browsing Large Sets of Found Audio Data (L18-1)

Copied to clipboard

Challenge: a new approach to search for sound in large archives is being developed . speech and speech technology researchers struggle to access large amounts of data .
Approach: They propose a method for fast and efficient non-sequential browsing of sound in large archives that we know little about . they combine audio browsing through massively multi-object sound environments and an unsupervised dimensionality reduction algorithm to search for sound in public archives.
Outcome: The proposed method is shown to combine well, resulting in rapid and interpretable observations.
CoLoSS: Cognitive Load Corpus with Speech and Performance Data from a Symbol-Digit Dual-Task (L18-1)

Copied to clipboard

Challenge: Existing corpus of speech recordings under cognitive load is available for research purposes .
Approach: They propose to use a corpus of speech under cognitive load recorded in a learning task scenario to obtain a reference for cognitive load.
Outcome: The proposed corpus is available to the scientific community for use in cognitive load-based cognitive load recognition.
VAST: A Corpus of Video Annotation for Speech Technologies (L18-1)

Copied to clipboard

Challenge: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Approach: The Video Annotation for Speech Technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Outcome: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech detection, language identification, speaker identification, and speech recognition .
Edit me: A Corpus and a Framework for Understanding Natural Language Image Editing (L18-1)

Copied to clipboard

Challenge: a corpus of image edit requests is elicited for real world images, and an annotation framework is developed . evaluators evaluate crowd-sourced annotation as a means of efficiently creating a sizable corpus at a reasonable cost.
Approach: They propose a natural language interface for interacting with an image editing program . they propose an annotation framework for understanding natural language requests .
Outcome: The proposed tool interprets image edit requests and maps them to actionable commands.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
SimPA: A Sentence-Level Simplification Corpus for the Public Administration Domain (L18-1)

Copied to clipboard

Challenge: lexical simplification is the task of reducing lexically and/or structural complexity of texts.
Approach: They propose to collect manual simplifications for 1,100 original sentences using a sentence-level corpus from the Public Administration domain.
Outcome: The proposed corpus contains 1,100 original sentences with manual simplifications collected through a two-stage process.
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized .
Approach: They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content .
Outcome: The proposed corpus is based on a pipeline methodology and is available for querying and downloading.
Czech Text Document Corpus v 2.0 (L18-1)

Copied to clipboard

Challenge: a corpus of text documents for automatic document classification in Czech is presented . paper aims to facilitate a straightforward comparison of document classification approaches on Czech data .
Approach: This paper introduces a collection of text documents for automatic document classification in Czech language.
Outcome: The proposed corpus is based on the Czech news agency's real newspaper articles . it is used for evaluation of multi-label document classification approaches .
Corpora of Typical Sentences (L18-1)

Copied to clipboard

Challenge: Typical sentences of characteristic syntactic structures can be used for language understanding tasks like finding typical slotfiller for verbs.
Approach: They propose a method to select "typical sentences" from 5% of the original corpus . they use entropy measuring the distribution of words in a given position to identify larger sets of near-duplicate sentences .
Outcome: The proposed method works language independently.
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)

Copied to clipboard

Challenge: DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text.
Approach: They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports.
Outcome: The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year.
Risamálheild: A Very Large Icelandic Text Corpus (L18-1)

Copied to clipboard

Challenge: The corpus contains more than one billion running words from mostly contemporary texts.
Approach: They present the Icelandic Gigaword Corpus (IGC) with minimal work and resources.
Outcome: The Icelandic Gigaword Corpus (IGC) contains more than one billion running words from mostly contemporary texts.
TriMED: A Multilingual Terminological Database (L18-1)

Copied to clipboard

Challenge: a terminological tool is developed to solve communication problems in medical language . medical terminology is often semantically complex and difficult to understand .
Approach: They propose a terminological tool that solves problems related to the opacity of medical language . they use a multilingual terminological-phraseological resource called TriMED to analyze medical terminology .
Outcome: The proposed tool solves problems related to the opacity that characterizes communication in the medical field among its actors.
Preparation and Usage of Xhosa Lexicographical Data for a Multilingual, Federated Environment (L18-1)

Copied to clipboard

Challenge: lexicographical data is often difficult to find for less resourced languages . Xhosa is a popular language in south africa, but it is often suboptimal for many languages despite its multilingual nature .
Approach: They propose a new source of lexicographical data for Xhosa, a language spoken by 8 million speakers.
Outcome: The proposed model can be used in multilingual and federated environments and is extensible to other languages.
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)

Copied to clipboard

Challenge: lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels.
Approach: They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense.
Outcome: The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format.
One Language to rule them all: modelling Morphological Patterns in a Large Scale Italian Lexicon with SWRL (L18-1)

Copied to clipboard

Challenge: Linked data (LD) is a popular way of publishing lexical resources, but technical limitations and potentialities of LD are not understood as they should be.
Approach: They propose to use the Semantic Web Rule Language to encode morphological patterns for a lexicographic publication as linked open data.
Outcome: The proposed language allows the automatic derivation of inflectional variants of entries in the lexicon.
Metaphor Suggestions based on a Semantic Metaphor Repository (L18-1)

Copied to clipboard

Challenge: Existing algorithms for suggesting metaphors have been used to find related words . corpus studies have found that metaphors are very pervasive even in formal language .
Approach: They propose an algorithm that suggests metaphoric means of referring to concepts . they use MetaNet, a repository of conceptual metaphor, and lexical resources .
Outcome: The proposed model expands the potential of the original repository by enabling new connections to be drawn.
The Linguistic Category Model in Polish (LCM-PL) (L18-1)

Copied to clipboard

Challenge: a new version of the Linguistic Category Model (LCM) dictionary for the Polish language is available for use and integrates with the Polish WordNet.
Approach: They propose to use a dictionary that is annotated manually in its most important parts . they propose to add more manually annotating senses and increase quality of automated annotations .
Outcome: The proposed dictionary is the first widely usable version of the resource . it will have more manually annotated senses and more automated annotations .
WordNet-Shp: Towards the Building of a Lexical Database for a Peruvian Minority Language (L18-1)

Copied to clipboard

Challenge: WordNet-like resources are lexical databases with highly relevance information and data that could be exploited in more complex computational linguistics research and applications.
Approach: They propose to build a WordNet database for a low-resourced and indigenous language in Peru . they propose to use word2vec similarity to compare definition glosses in a dictionary with the content of a Spanish WordNet .
Outcome: The proposed database is based on a bilingual dictionary written in Spanish and an automatic evaluation process using a manually annotated Gold Standard in Shipibo-Koniba.
Retrieving Information from the French Lexical Network in RDF/OWL Format (L18-1)

Copied to clipboard

Challenge: a Java API to retrieve lexical information from the French Lexical Network is presented . RDF/OWL languages are not sufficient for a more detailed representation of linguistic information.
Approach: They propose a Java API to retrieve lexical information from the French Lexical Network . this API was used in the identification of collocations in a french corpus of 1.8 million sentences .
Outcome: The proposed API was used to identify collocations in a French corpus of 1.8 million sentences and in the semantic classification of these collocation.
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)

Copied to clipboard

Challenge: et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training.
Approach: They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks .
Outcome: The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score.
Error Analysis of Uyghur Name Tagging: Language-specific Techniques and Remaining Challenges (L18-1)

Copied to clipboard

Challenge: despite efforts at name tagging, there is limited understanding on the performance ceiling . despite the high-resource language, there are very few natural language processing tools available .
Approach: They propose to use a machine learning model to identify Uyghur name tagger errors . they conclude that such a model is unlikely to be effective for Uygur, or low-resource languages .
Outcome: The proposed model is unlikely to be effective for Uyghur, or low-resource languages in general, the authors argue . they show that the proposed model can be used for high-res languages with superficial features .
BiLSTM-CRF for Persian Named-Entity Recognition ArmanPersoNERCorpus: the First Entity-Annotated Persian Dataset (L18-1)

Copied to clipboard

Challenge: Named-entity recognition (NER) is a natural language processing component that aims to identify all the "named entities" (NEs) in an unstructured text.
Approach: They propose a deep learning approach for name-entity recognition in Persian . they publicize an entity-annotated Persian dataset and train word embeddings .
Outcome: The proposed approach achieves a 77.45% CoNLL F 1 score for Persian NER based on a deep learning architecture and pre-trained word embeddings.
Data Anonymization for Requirements Quality Analysis: a Reproducible Automatic Error Detection Task (L18-1)

Copied to clipboard

Challenge: a recent study focuses on identifying potential problems of ambiguity, completeness, conformity, singularity and readability in requirements specifications.
Approach: They propose to identify potential problems of ambiguity, completeness, conformity, singularity and readability in system and software requirements specifications.
Outcome: The proposed system achieves 79.47% for the F1 score on proposed evaluation data.
A German Corpus for Fine-Grained Named Entity Recognition and Relation Extraction of Traffic and Industry Events (L18-1)

Copied to clipboard

Challenge: Using text streams to extract events pertaining to specific companies, routes and routes remains a challenge.
Approach: They describe a corpus of German-language documents annotated with fine-grained geo-entities and standard named entity types.
Outcome: The proposed corpus consists of newswire texts, twitter messages, and traffic reports from radio stations, police and railway companies.
A Corpus Study and Annotation Schema for Named Entity Recognition and Relation Extraction of Business Products (L18-1)

Copied to clipboard

Challenge: Existing annotation guidelines for non-standard entity types and relations are lacking in news and forum texts.
Approach: They propose a corpus study and an annotation schema for the annotation of product entity and company-product relation mentions.
Outcome: The proposed annotation schema and guidelines are applied to the annotation of product entities and company-product relation mentions.
Portuguese Named Entity Recognition using Conditional Random Fields and Local Grammars (L18-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) involves automatically identifying and classifying entities such as persons, places, organizations and values.
Approach: They propose to use Conditional Random Fields for Named Entity Recognition in Portuguese texts using a Local Grammar as an additional informed feature.
Outcome: The proposed method outperforms competing systems in the literature.
M-CNER: A Corpus for Chinese Named Entity Recognition in Multi-Domains (L18-1)

Copied to clipboard

Challenge: NER is one of the most important natural language processing tasks.
Approach: They propose to annotate sentences from human-computer interaction, social media, and e-commerce using two rounds of annotation.
Outcome: The proposed system performs the best on all the data sets.
SlugNERDS: A Named Entity Recognition Tool for Open Domain Dialogue Systems (L18-1)

Copied to clipboard

Challenge: UCSC researchers have developed an open domain social bot aimed at casual conversation . NER and NEL are important preprocessing steps for understanding user intent in open domain dialogue systems.
Approach: They propose a tool for NER and NEL in open domain dialogue that addresses these challenges . they also propose two corpora based on 10,000 real user conversations .
Outcome: The proposed open domain social bot is aimed at casual conversation.
Transfer Learning for Named-Entity Recognition with Neural Networks (L18-1)

Copied to clipboard

Challenge: Existing approaches to named-entity recognition (NER) require additional lead time for developing and fine-tuning the rules.
Approach: They propose to transfer an ANN model trained on a large labeled dataset to another dataset with a limited number of labels to improve upon the state-of-the-art results for patient note de-identification.
Outcome: The proposed model can be transferred to a dataset with a limited number of labels, and improves on the state-of-the-art results on patient note de-identification.
ForFun 1.0: Prague Database of Forms and Functions – An Invaluable Resource for Linguistic Research (L18-1)

Copied to clipboard

Challenge: a new database, ForFun, is an invaluable resource for linguistic research . it allows researchers to elaborate various syntactic issues in depth .
Approach: They introduce the first version of ForFun, Prague Database of Forms and Functions . it uses syntactically annotated Prague Dependency Treebanks to organize annotations anew .
Outcome: The proposed database brings syntactic issues closer to researchers . it takes advantage of existing treebanks and offers a user-friendly access to real examples .
The LIA Treebank of Spoken Norwegian Dialects (L18-1)

Copied to clipboard

Challenge: a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material.
Approach: They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation.
Outcome: The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian.
Errator: a Tool to Help Detect Annotation Errors in the Universal Dependencies Project (L18-1)

Copied to clipboard

Challenge: UD project aims to develop cross-linguistically consistent treebank annotations for a wide array of languages.
Approach: They introduce tools that implement the annotation variation principle to help annotators find and correct errors in UD treebanks.
Outcome: The proposed tools can be used to correct errors in UD treebank annotations.
SandhiKosh: A Benchmark Corpus for Evaluating Sanskrit Sandhi Tools (L18-1)

Copied to clipboard

Challenge: Several important texts which are of interest to people all over the world were written in Sanskrit.
Approach: They develop a Sanskrit benchmark to evaluate the completeness and accuracy of tools . they use three most prominent tools to evaluate their completeness .
Outcome: The proposed tools have substantial scope for improvement and are available to researchers worldwide.
Czech Legal Text Treebank 2.0 (L18-1)

Copied to clipboard

Challenge: In this paper, we introduce the Czech Legal Text Treebank 2.0 with more elaborate syntactic annotations and enriched with two annotation layers.
Approach: They introduce a new version of the Czech Legal Text Treebank 2.0 with more elaborate syntactic annotations and two new annotation layers.
Outcome: The new version of the Czech Legal Text Treebank contains more elaborate syntactic annotations and two new annotation layers.
Creation of a Balanced State-of-the-Art Multilayer Corpus for NLU (L18-1)

Copied to clipboard

Challenge: Using full stack of language resources, we are creating a balanced text corpus for Latvian.
Approach: They propose to create a syntactically and semantically annotated multilayered corpus for Latvian . they use widely acknowledged and cross-lingual representations for the corpus .
Outcome: The proposed corpus adopts widely recognized and cross-lingual representations for natural language understanding and generation in Latvian.
Test Sets for Chinese Nonlocal Dependency Parsing (L18-1)

Copied to clipboard

Challenge: Chinese is a language rich in nonlocal dependencies.
Approach: They use trace annotations in the Penn Chinese Treebank to generate test sets of Chinese nonlocal dependencies which occur in different grammatical constructions.
Outcome: The proposed test sets can be used to evaluate nonlocal dependency recovery in Chinese.
Adding Syntactic Annotations to Flickr30k Entities Corpus for Multimodal Ambiguous Prepositional-Phrase Attachment Resolution (L18-1)

Copied to clipboard

Challenge: Using visual features extracted from an image, we propose to study the joint processing of image and language features for the Preposition-Phrase attachment disambiguation task.
Approach: They propose to add syntactic annotations to the captions of the Flickr30k Entities corpus to study the joint processing of image and language features for the Preposition-Phrase attachment disambiguation task.
Outcome: The proposed framework is based on the captions of the Flickr30k Entities corpus and is automatically projected on their French and German translations.
Analyzing Middle High German Syntax with RDF and SPARQL (L18-1)

Copied to clipboard

Challenge: Using CoNLL-RDF and SPARQL Update, we analyze the diachronic changes of Middle High German syntax.
Approach: They propose a rule-based shallow parser and an enrichment pipeline grounded in CoNLL-RDF and SPARQL Update for parsing.
Outcome: The proposed pipeline is based on CoNLL-RDF and SPARQL Update for syntactic annotation and semantic enrichment of Middle High German.
Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer (L18-1)

Copied to clipboard

Challenge: Using annotated corpus for linguistic purposes is no longer justified . hand-crafted syntactic resources such as grammars and lexicons can be used as sources of features to guide data driven systems.
Approach: They propose an efficient method for transferring annotations between two different treebanks of the same language.
Outcome: The proposed method is based on the Universal Dependency annotation scheme and was evaluated on the gold standard (94.75% of LAS, 99.40% UAS on the test set).
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
Undersampling Improves Hypernymy Prototypicality Learning (L18-1)

Copied to clipboard

Challenge: supervised hypernymy detection suffers from overfitting hypernies in training data.
Approach: They propose a method that can alleviate the problem of overfitting hypernyms in training data by using distributional representations for unknown word pairs.
Outcome: The proposed method alleviates the problem of overfitting hypernyms in training data and improves distributional prototypicality learning for unknown word pairs.
Interoperability of Language-related Information: Mapping the BLL Thesaurus to Lexvo and Glottolog (L18-1)

Copied to clipboard

Challenge: The Bibliography of Linguistic Literature (BLL Thesaurus) has been used since 2013 in the context of the Lin gu is tik portal, a hub for linguistically relevant information.
Approach: They propose to use Lexvo and Glottolog to facilitate interoperability between the BLL Thesaurus and terminological repositories in the Linguistic Linked Open Data cloud.
Outcome: The proposed model is based on Lexvo and Glottolog and is able to connect to the Linguistic Linked Open Data cloud.
Browsing and Supporting Pluricentric Global Wordnet, or just your Wordnet of Interest (L18-1)

Copied to clipboard

Challenge: a wordnet browser that allows to consult wordnet content is presented in this paper . the paper presents a browser that meets design requirements and complies with the most ample range of design features.
Approach: They propose a wordnet browser that meets design requirements for wordnets . they use existing browsers to analyze their functionalities and build a new browser .
Outcome: The proposed browser meets design requirements and complies with the most ample range of design features.
Cross-checking WordNet and SUMO Using Meronymy (L18-1)

Copied to clipboard

Challenge: Existing methods to validate knowledge encoded in WordNet, SUMO and their mapping have been mainly manual .
Approach: They propose to use WordNet and SUMO to validate knowledge by using automated theorem provers to evaluate the competency of SU MO-based ontologies.
Outcome: The proposed method enables validation of some pieces of information and also the detection of missing information or inconsistencies among knowledge resources.
Extended HowNet 2.0 – An Entity-Relation Common-Sense Representation Model (L18-1)

Copied to clipboard

Challenge: Extended HowNet 2.0 is a common-sense representation model for lexical senses .
Approach: They propose Extended HowNet 2.0 -an entity-relation common-sense representation model . a query system is being developed for flexibly clustering concepts .
Outcome: The proposed model can bring significant benefits to the community of lexical semantics and natural language understanding.
The Circumstantial Event Ontology (CEO) and ECB+/CEO: an Ontology and Corpus for Implicit Causal Relations between Events (L18-1)

Copied to clipboard

Challenge: a new ontology for calamity events models semantic circumstantial relations between event classes . a circumstancial relation makes clear "why" something happened, without necessarily predicting it.
Approach: They propose a circumstantial event ontology that models semantic circumstancial relations between event classes . they propose ECB+ annotated corpus for circumstantal relations and a meta model .
Outcome: The proposed model captures that the change yielded by one event explains to people the happening of the next event when observed.
Profiling Medical Journal Articles Using a Gene Ontology Semantic Tagger (L18-1)

Copied to clipboard

Challenge: a growing number of scientific publications are based on sub-divisions and sub-communities of expertise becoming disconnected from each other.
Approach: They propose to examine corpora derived from bodies of genetics literature and use it to make comparisons and improve retrieval methods.
Outcome: The proposed methods will help to make comparisons and improve retrieval methods using domain knowledge via an existing gene ontology.
Towards a Conversation-Analytic Taxonomy of Speech Overlap (L18-1)

Copied to clipboard

Challenge: a taxonomy for classifying speech overlap in natural language dialogue is presented . the scheme classifies overlap on the basis of several features, including onset point, local dialogue history, and management behavior.
Approach: They propose a taxonomy for classifying speech overlap in natural language dialogue . they describe the various dimensions of the scheme and show how it was applied to a corpus of collaborative dialogue based on onset point, dialogue history, and management behavior .
Outcome: The proposed taxonomy classifies overlap on the basis of onset point, dialogue history, management behavior.
Indian Language Wordnets and their Linkages with Princeton WordNet (L18-1)

Copied to clipboard

Challenge: Wordnets are rich lexico-semantic resources. Linked wordnets link similar concepts in wordnet of different languages.
Approach: They propose to map 18 Indian wordnets linked with Princeton WordNet . they use expansion approach with Hindi Wordnet as pivot .
Outcome: The proposed mappings of 18 Indian wordnets are based on Princeton WordNet . they show that availability of such resources will have a direct impact on NLP progress .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations