Papers by Janosch Haber
Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore (2023.acl-long)
Copied to clipboard
Janosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Röttger
| Challenge: | Toxic content is a global problem, but most resources for detecting toxic content are in English . new datasets and models for non-English languages focus exclusively on one language or dialect . |
| Approach: | They propose to use a multilingual dataset of online attacks to identify code-mixed toxic content in Singapore . they collect reddit comments in Indonesian, Malay, Singlish, and other languages and provide fine-grained hierarchical labels for attacks . |
| Outcome: | The proposed dataset provides fine-grained hierarchical labels for online attacks in Singapore . it shows that the metadata can be used for granular error analysis . |
The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue (P19-1)
Copied to clipboard
| Challenge: | Using the PhotoBook dataset, we investigate shared dialogue history accumulating during conversation . human interlocutors are known to collaboratively establish a shared repository of mutual information during a conversation - this common ground is then used to optimise understanding and communication efficiency. |
| Approach: | They propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context and previously established referring expressions. |
| Outcome: | The proposed model takes into account shared information accumulated in a reference chain and is important to resolve later descriptions. |
Assessing Polyseme Sense Similarity through Co-predication Acceptability and Contextualised Embedding Distance (2020.starsem-1)
Copied to clipboard
| Challenge: | Co-predication is a commonly used linguistic test to tell apart shifts in polysemic sense from changes in homonymic meaning. |
| Approach: | They examine how co-predication acceptability relates to explicit ratings of polyseme word sense similarity and how well they can be predicted through the distance between target words’ contextualised word embeddings. |
| Outcome: | The proposed measures can be predicted through the distance between target words’ contextualised word embeddings. |
Patterns of Polysemy and Homonymy in Contextualised Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study has focused on homonymy, a variety of multiplicity of meanings exemplified by word forms with unrelated meanings. |
| Approach: | They investigate the extent to which contextualised embeddings reflect traditional distinctions of polysemy and homonymy. |
| Outcome: | The proposed model shows that it can distinguish between polysemy and homonymy . it shows that the model fails to replicate the results of the human-annotated dataset . |