On the Idiosyncrasies of the Mandarin Chinese Classifier System (N19-1)

Copied to clipboard

Challenge: idiosyncrasies of the Chinese classifier system have been studied, but little work has been done to quantify them with statistical methods.
Approach: They propose an information-theoretic approach to measuring idiosyncrasies in Mandarin Chinese by calculating the mutual information between the distribution over classifiers and distributions over other linguistic quantities.
Outcome: The proposed method reduces uncertainty in Mandarin Chinese classifiers by knowing semantic information about nouns that they modify.

Similar Papers

Mandarin classifier systems optimize to accommodate communicative pressures (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies suggest that gendered languages are inherently optimized to accommodate communicative pressures on language learning and processing.
Approach: They propose to use grammatical or probabilistic modifiers to smooth the entropy of nouns in context to find the same frequency, similarity, and co-occurrence interactions that structure gender systems.
Outcome: The proposed noun classification device is sensitive to frequency, similarity, and co-occurrence interactions that structure gender systems.
Understanding the Use of Quantifiers in Mandarin (2022.findings-aacl)

Copied to clipboard

Challenge: a corpus of short texts in Mandarin is analyzed to examine the "coolness" hypothesis . quantified expressions are used to describe short texts, but are not as informative as English .
Approach: They propose a corpus of Mandarin in which quantified expressions figure prominently.
Outcome: The proposed corpus of short texts in Mandarin is compared with an English corpus.
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to detect cross-linguistic associations are not effective, but their effects are minor.
Approach: They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon.
Outcome: The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average).
The UIR Uncertainty Corpus for Chinese: Annotating Chinese Microblog Corpus for Uncertainty Identification from Social Media (L18-1)

Copied to clipboard

Challenge: Uncertainty identification is an important semantic processing task, critical to the quality of information in terms of factuality in many NLP techniques and applications.
Approach: They propose to annotate Chinese microblogs with an open uncertainty corpus . they propose to use contextual uncertain semantics rather than traditional cue-phrases to identify uncertainty .
Outcome: The proposed corpus can be used to identify uncertainty in social media texts.
Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
Comparing Static and Contextual Distributional Semantic Models on Intrinsic Tasks: An Evaluation on Mandarin Chinese Datasets (2024.lrec-main)

Copied to clipboard

Challenge: Distributional Semantics has undergone significant changes with the introduction of contextualized distributional models.
Approach: They compare static and contextual distributional models for Mandarin Chinese . they find that static models are stronger for some of the classical tasks .
Outcome: The proposed models perform better on some of the classical tasks that consider word meaning independent of context, while contextualized models excel in identifying semantic relations between word pairs and categorization of words into abstract semantic classes.
Ambiguity Meets Uncertainty: Investigating Uncertainty Estimation for Word Sense Disambiguation (2023.findings-acl)

Copied to clipboard

Challenge: Existing supervised methods treat word sense disambiguation as a classification task but ignore uncertainty estimation (UE) in the real-world setting, the data is always noisy and out of distribution.
Approach: They propose to use word sense disambiguation to determine an appropriate sense for a word given its context to determine the most appropriate sense.
Outcome: The proposed model reflects data uncertainty satisfactorily but underestimates model uncertainty.
Mandarinograd: A Chinese Collection of Winograd Schemas (2020.lrec-1)

Copied to clipboard

Challenge: Mandarinograd is a corpus of Winograd Schemas in Mandarin Chinese . WS are hard to collect and few datasets are publicly available .
Approach: They introduce a corpus of Winograd Schemas in Mandarin Chinese . they describe the difficulties faced when building the corpus and explain how they overcome the anomalies.
Outcome: The proposed corpus of Winograd Schemas in Mandarin Chinese is hard to build and resistant to statistical methods.
Why is penguin more similar to polar bear than to sea gull? Analyzing conceptual knowledge in distributional models (2020.acl-srw)

Copied to clipboard

Challenge: Several analysis methods have been shown to be limited and are not well understood . thesis aims to understand distributional semantic representations based on linguistic data .
Approach: They propose a framework for investigating the information encoded in distributional semantic models . they combine observations made on corpora with insights obtained from data manipulation experiments .
Outcome: The proposed framework pairs observations made on corpora with insights obtained from data manipulation experiments.
WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)

Copied to clipboard

Challenge: Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count.
Approach: They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP.
Outcome: The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations