Challenge: a huge number of people use social media to express and exchange information in their own languages.
Approach: They propose to use a code-mixed environment to extract higher level features from text . they use 'gadget' algorithm that automatically discovers higher level feature from text.
Outcome: The proposed approach is generic and does not make use of handcrafted features or rules.

Similar Papers

Corpus Creation and Analysis for Named Entity Recognition in Telugu-English Code-Mixed Social Media Data (P19-2)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of Information Extraction in NLP.
Approach: They present a Telugu-English code-mixed corpus with the corresponding named entity tags.
Outcome: The proposed model scored 0.96, 0.94 and 0.95 on a Telugu-English code-mixed corpus.
Language Identification and Named Entity Recognition in Hinglish Code Mixed Tweets (P18-3)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is an important text analysis task . code-mixing occurs when lexical items and grammatical features from two languages appear in one sentence .
Approach: They propose to use language identifiers, parts-of-speech tags and chunkers to analyze code-mixed data.
Outcome: The proposed method outperforms the best baseline by 33.18%.
Corpus Creation and Emotion Prediction for Hindi-English Code-Mixed Social Media Text (N18-4)

Copied to clipboard

Challenge: Emotion Prediction is a natural language processing task dealing with detection and classification of emotions in monolingual and bilingual texts.
Approach: They propose a machine learning system which uses various machine learning techniques to detect emotion associated with tweets.
Outcome: The proposed system uses various machine learning techniques to detect emotion associated with the text.
De-Mixing Sentiment from Code-Mixed Text (P19-2)

Copied to clipboard

Challenge: Code-mixing is the phenomenon of mixing the vocabulary and syntax of multiple languages in the same sentence.
Approach: They propose a hybrid architecture for the task of Sentiment Analysis of English-Hindi code-mixed data using CNNs to generate subword representations for the sentences.
Outcome: The proposed architecture achieves 83.54% accuracy and 0.827 F1 score on a benchmark dataset.
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)

Copied to clipboard

Challenge: a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people.
Approach: They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook.
Outcome: The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.
A Neural Network Model for Part-Of-Speech Tagging of Social Media Texts (L18-1)

Copied to clipboard

Challenge: Recent approaches based on end-to-end Deep Neural Networks (DNNs) have shown promising results for Natural Language Processing (NLP).
Approach: They propose a neural network model for part-of-speech (POS) tagging of User-Generated Content (UGC) such as Twitter, Facebook and Web forums that uses character and word representations.
Outcome: The proposed model is end-to-end and uses character and word representations . it is compared with existing models on social media in English, german, french, italian and spanish .
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment.
Approach: They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions.
Outcome: The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages.
A Platform for Event Extraction in Hindi (2020.lrec-1)

Copied to clipboard

Challenge: Event Extraction is an important task in the widespread field of NLP, but there is no benchmark setup in Hindi.
Approach: They propose an Event Extraction framework for Hindi language and develop deep learning based models to set as the baselines.
Outcome: The proposed framework crawls more than seventeen hundred disaster related Hindi news articles from various news sources.
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)

Copied to clipboard

Challenge: a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text .
Approach: They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags.
Outcome: The proposed method detects humor in code-mixed tweets in English-Hindi.
Neural Adaptation Layers for Cross-domain Named Entity Recognition (D18-1)

Copied to clipboard

Challenge: Named entity recognition is a type of information extraction task whereby features can be designed based on domain-specific knowledge.
Approach: They propose to use existing neural architectures to adapt to new domains without retraining . they propose to add adaptation layers to existing neural models to minimize re-training based on source data.
Outcome: The proposed approach significantly outperforms state-of-the-art methods on social media domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations