Papers with bigrams
Big BiRD: A Large, Fine-Grained, Bigram Relatedness Dataset for Examining Semantic Composition (N19-1)
Copied to clipboard
| Challenge: | Existing datasets of semantic relatedness only include pairs of unigrams (single words) Existing data suffer from inconsistent annotations and scale region bias due to rating scales. |
| Approach: | They propose to use a large, fine-grained, bigram relatedness dataset to compare the relatedness of 3,345 English term pairs using a comparative annotation technique called Best–Worst Scaling. |
| Outcome: | The proposed datasets are highly reliable and have a split-half reliability of 0.937. |
Finding Dataset Shortcuts with Grammar Induction (2022.emnlp-main)
Copied to clipboard
| Challenge: | Prior work on shortcut detection focused on enumerating features like unigrams or bigrams . prior work relied on post-hoc models that reveal qualitative patterns without a clear statistical interpretation . |
| Approach: | They propose to use probabilistic grammars to characterize and discover shortcuts in NLP datasets using context-free grammars and synchronous context- free grammars. |
| Outcome: | The proposed grammars reveal interesting shortcut features in a number of datasets, including simple and high-level features, and automatically identify groups of test examples on which conventional classifiers fail. |
Mutual Gaze and Linguistic Repetition in a Multimodal Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | a study of linguistic repetitions and mutual understanding is conducted . we find no compelling correlation between mutual gaze and duration of the event . |
| Approach: | They investigate the correlation between mutual gaze and linguistic repetition, a form of alignment, which they take as evidence of mutual understanding. |
| Outcome: | The proposed method is based on the Multisimo corpus, a multimodal corpus which provides authentic task-based interactions among three participants. |
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains. |
| Approach: | They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic. |
| Outcome: | The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes. |