Analyzing Linguistic Differences between Owner and Staff Attributed Tweets (P19-1)

Copied to clipboard

Challenge: Existing studies on social media assume that all tweets are authored by the same person . this paper examines the linguistic differences between posts signed by the account owner and staff attributed to the account's owner .
Approach: They analyze linguistic differences between tweets signed by the account owner and staff attributed to their staff.
Outcome: The proposed model predicts owner and staff attributed tweets with good accuracy even without training data from the account.

Similar Papers

The Trumpiest Trump? Identifying a Subject’s Most Characteristic Tweets (D19-1)

Copied to clipboard

Challenge: characterization scores are associated with popularity of a given short text, but are not always representative of the source.
Approach: They use a dataset of tweets from 15 celebrities to quantify the extent to which a given short text is characteristic of a specific person.
Outcome: The proposed model shows a statistically significant correlation between characterization scores and popularity of the associated texts for 13 of the 15 celebrities in the study.
Categorizing and Inferring the Relationship between the Text and Image of Twitter Posts (P19-1)

Copied to clipboard

Challenge: Social media posts often contain images to provide content, provide context, or express feelings.
Approach: They build and release a dataset of image tweets annotated with four different classes which express whether the text or the image provides additional information to the other modality.
Outcome: The proposed method can be used in several downstream applications including pre-training image tagging models and collecting distantly supervised data for image captioning.
Exploring Author Context for Detecting Intended vs Perceived Sarcasm (P19-1)

Copied to clipboard

Challenge: Existing studies on textual sarcasm detection use manual labelling and tag-based distant supervision to detect sarcasm.
Approach: They define author context as the embedded representation of their historical tweets and suggest neural models that extract these representations.
Outcome: The proposed models achieve state-of-the-art on two datasets labelled manually and via tag-based distant supervision indicating a difference between intended and perceived sarcasm .
Biographically Relevant Tweets – a New Dataset, Linguistic Analysis and Classification Experiments (2022.coling-1)

Copied to clipboard

Challenge: Unlike previous work, we do not restrict biographical relevance to a small fixed set of pre-defined relations.
Approach: They propose a dataset comprising tweets for the novel task of detecting biographically relevant utterances.
Outcome: The proposed dataset focuses on biographical information on ordinary users of Twitter.
TSix: A Human-involved-creation Dataset for Tweet Summarization (L18-1)

Copied to clipboard

Challenge: a new dataset for tweet summarization is available for free.
Approach: They propose a dataset for tweet summarization that uses human annotations to evaluate extractive summarizing methods.
Outcome: The proposed dataset includes six events collected from Twitter . human-annotated gold-standard references facilitate evaluation, the study shows .
Towards Automated Semantic Role Labelling of Hindi-English Code-Mixed Tweets (D19-55)

Copied to clipboard

Challenge: a new system for semantic role labelling of Hindi-English code-mixed tweets is proposed . code-mixing is a largely observed phenomenon in colloquial usage and on social media .
Approach: They propose a system for automating Semantic Role Labelling of Hindi-English code-mixed tweets.
Outcome: The proposed system gives an overall accuracy of 84% for Argument Classification, a 10% increase over the existing rule-based model.
Extracting Possessions from Social Media: Images Complement Language (D19-1)

Copied to clipboard

Challenge: Existing studies show that authors of tweets possess objects they tweet about.
Approach: They propose a dataset and experiments to determine whether tweet authors possess objects they tweet about.
Outcome: The proposed strategy incorporates visual information into any neural network beyond weights from pretrained networks.
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation.
Approach: They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic .
Outcome: The proposed models do not generalize, indicating heterogeneous political users.
What does the language of foods say about us? (D19-62)

Copied to clipboard

Challenge: Using a dataset of 24 million food-related tweets, we can predict if states in the United States are above the median rates for type 2 diabetes mellitus (T2DM) income, poverty, and education are important factors in predicting T2DM rates, but socioeconomic factors do not capture this information.
Approach: They use a dataset of 24 million food-related tweets to investigate the signal contained in the language of food on social media.
Outcome: The language of food can predict health risks, political orientation, and geographic location, and outperform previous work by 4–18%.
Disambiguating False-Alarm Hashtag Usages in Tweets for Irony Detection (P18-2)

Copied to clipboard

Challenge: Existing methods to collect self-labeled data for irony detection are based on false-alarm hashtags.
Approach: They propose a neural network-based model which disambiguates hashtag usages and prunes the self-labeled training data.
Outcome: The proposed model outperforms the models trained on the less but cleaner training instances.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations