Papers by Behnam Sabeti

5 papers
Optimizing Annotation Effort Using Active Learning Strategies: A Sentiment Analysis Case Study in Persian (2020.lrec-1)

Copied to clipboard

Challenge: Existing deep learning approaches require huge amounts of data to be trained properly.
Approach: They propose to use Persian as a model to choose the samples for annotation instead of labeling the whole dataset.
Outcome: The proposed models achieve the baseline performance with a significantly lower amount of labeled data.
MirasText: An Automatically Generated Text Corpus for Persian (L18-1)

Copied to clipboard

Challenge: Natural language processing is one of the most important fields of artificial intelligence.
Approach: They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens .
Outcome: The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles .
Irony Detection in Persian Language: A Transfer Learning Approach Using Emoji Prediction (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for emotion extraction and sentiment analysis produce invalid results due to the use of irony.
Approach: They propose to use emoji prediction to fine tune a model using hand labeled tweets with irony tags.
Outcome: The proposed method outperforms the state-of-the-art method on Persian dataset with an accuracy of 83.1% and offers strong baseline for further research in Persian language.
MirasVoice: A bilingual (English-Persian) speech corpus (L18-1)

Copied to clipboard

Challenge: Existing research and development areas in speech recognition are focused on the language of speakers.
Approach: They propose to use a bilingual (English-Farsi) speech corpus to validate and explore speaker verification systems.
Outcome: The proposed corpus can be used in a variety of language dependent and independent applications.
Twitter Trend Extraction: A Graph-based Approach for Tweet and Hashtag Ranking, Utilizing No-Hashtag Tweets (2020.lrec-1)

Copied to clipboard

Challenge: Twitter has become a major platform for users to express their opinions on any topic and engage in debates.
Approach: They propose to use tweets as graph nodes to extract trends from tweets graph . they propose to employ RankClus algorithm to rank tweets, words and hashtags in each trend .
Outcome: The proposed algorithm can extract trends from tweets and rank tweets, words and hashtags based on their importance and relevance to the topic.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations