Papers by Janosch Haber

4 papers
Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore (2023.acl-long)

Copied to clipboard

Challenge: Toxic content is a global problem, but most resources for detecting toxic content are in English . new datasets and models for non-English languages focus exclusively on one language or dialect .
Approach: They propose to use a multilingual dataset of online attacks to identify code-mixed toxic content in Singapore . they collect reddit comments in Indonesian, Malay, Singlish, and other languages and provide fine-grained hierarchical labels for attacks .
Outcome: The proposed dataset provides fine-grained hierarchical labels for online attacks in Singapore . it shows that the metadata can be used for granular error analysis .
The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue (P19-1)

Copied to clipboard

Challenge: Using the PhotoBook dataset, we investigate shared dialogue history accumulating during conversation . human interlocutors are known to collaboratively establish a shared repository of mutual information during a conversation - this common ground is then used to optimise understanding and communication efficiency.
Approach: They propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context and previously established referring expressions.
Outcome: The proposed model takes into account shared information accumulated in a reference chain and is important to resolve later descriptions.
Assessing Polyseme Sense Similarity through Co-predication Acceptability and Contextualised Embedding Distance (2020.starsem-1)

Copied to clipboard

Challenge: Co-predication is a commonly used linguistic test to tell apart shifts in polysemic sense from changes in homonymic meaning.
Approach: They examine how co-predication acceptability relates to explicit ratings of polyseme word sense similarity and how well they can be predicted through the distance between target words’ contextualised word embeddings.
Outcome: The proposed measures can be predicted through the distance between target words’ contextualised word embeddings.
Patterns of Polysemy and Homonymy in Contextualised Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: a recent study has focused on homonymy, a variety of multiplicity of meanings exemplified by word forms with unrelated meanings.
Approach: They investigate the extent to which contextualised embeddings reflect traditional distinctions of polysemy and homonymy.
Outcome: The proposed model shows that it can distinguish between polysemy and homonymy . it shows that the model fails to replicate the results of the human-annotated dataset .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations