Analyzing Linguistic Differences between Owner and Staff Attributed Tweets (P19-1)
Copied to clipboard
| Challenge: | Existing studies on social media assume that all tweets are authored by the same person . this paper examines the linguistic differences between posts signed by the account owner and staff attributed to the account's owner . |
| Approach: | They analyze linguistic differences between tweets signed by the account owner and staff attributed to their staff. |
| Outcome: | The proposed model predicts owner and staff attributed tweets with good accuracy even without training data from the account. |
Similar Papers
The Trumpiest Trump? Identifying a Subject’s Most Characteristic Tweets (D19-1)
Copied to clipboard
| Challenge: | characterization scores are associated with popularity of a given short text, but are not always representative of the source. |
| Approach: | They use a dataset of tweets from 15 celebrities to quantify the extent to which a given short text is characteristic of a specific person. |
| Outcome: | The proposed model shows a statistically significant correlation between characterization scores and popularity of the associated texts for 13 of the 15 celebrities in the study. |
Categorizing and Inferring the Relationship between the Text and Image of Twitter Posts (P19-1)
Copied to clipboard
| Challenge: | Social media posts often contain images to provide content, provide context, or express feelings. |
| Approach: | They build and release a dataset of image tweets annotated with four different classes which express whether the text or the image provides additional information to the other modality. |
| Outcome: | The proposed method can be used in several downstream applications including pre-training image tagging models and collecting distantly supervised data for image captioning. |
Exploring Author Context for Detecting Intended vs Perceived Sarcasm (P19-1)
Copied to clipboard
| Challenge: | Existing studies on textual sarcasm detection use manual labelling and tag-based distant supervision to detect sarcasm. |
| Approach: | They define author context as the embedded representation of their historical tweets and suggest neural models that extract these representations. |
| Outcome: | The proposed models achieve state-of-the-art on two datasets labelled manually and via tag-based distant supervision indicating a difference between intended and perceived sarcasm . |
Biographically Relevant Tweets – a New Dataset, Linguistic Analysis and Classification Experiments (2022.coling-1)
Copied to clipboard
| Challenge: | Unlike previous work, we do not restrict biographical relevance to a small fixed set of pre-defined relations. |
| Approach: | They propose a dataset comprising tweets for the novel task of detecting biographically relevant utterances. |
| Outcome: | The proposed dataset focuses on biographical information on ordinary users of Twitter. |
TSix: A Human-involved-creation Dataset for Tweet Summarization (L18-1)
Copied to clipboard
| Challenge: | a new dataset for tweet summarization is available for free. |
| Approach: | They propose a dataset for tweet summarization that uses human annotations to evaluate extractive summarizing methods. |
| Outcome: | The proposed dataset includes six events collected from Twitter . human-annotated gold-standard references facilitate evaluation, the study shows . |
Towards Automated Semantic Role Labelling of Hindi-English Code-Mixed Tweets (D19-55)
Copied to clipboard
| Challenge: | a new system for semantic role labelling of Hindi-English code-mixed tweets is proposed . code-mixing is a largely observed phenomenon in colloquial usage and on social media . |
| Approach: | They propose a system for automating Semantic Role Labelling of Hindi-English code-mixed tweets. |
| Outcome: | The proposed system gives an overall accuracy of 84% for Argument Classification, a 10% increase over the existing rule-based model. |
Extracting Possessions from Social Media: Images Complement Language (D19-1)
Copied to clipboard
| Challenge: | Existing studies show that authors of tweets possess objects they tweet about. |
| Approach: | They propose a dataset and experiments to determine whether tweet authors possess objects they tweet about. |
| Outcome: | The proposed strategy incorporates visual information into any neural network beyond weights from pretrained networks. |
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation. |
| Approach: | They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic . |
| Outcome: | The proposed models do not generalize, indicating heterogeneous political users. |
What does the language of foods say about us? (D19-62)
Copied to clipboard
| Challenge: | Using a dataset of 24 million food-related tweets, we can predict if states in the United States are above the median rates for type 2 diabetes mellitus (T2DM) income, poverty, and education are important factors in predicting T2DM rates, but socioeconomic factors do not capture this information. |
| Approach: | They use a dataset of 24 million food-related tweets to investigate the signal contained in the language of food on social media. |
| Outcome: | The language of food can predict health risks, political orientation, and geographic location, and outperform previous work by 4–18%. |
Disambiguating False-Alarm Hashtag Usages in Tweets for Irony Detection (P18-2)
Copied to clipboard
| Challenge: | Existing methods to collect self-labeled data for irony detection are based on false-alarm hashtags. |
| Approach: | They propose a neural network-based model which disambiguates hashtag usages and prunes the self-labeled training data. |
| Outcome: | The proposed model outperforms the models trained on the less but cleaner training instances. |