Challenge: Polysemy between adpositions and case markers is high in casemarked languages . a singular adeposition can cover a wide range of semantic fields while occupying the same syntactic context.
Approach: They propose to manually annotate adposition and case marker tokens in Finnish and Latin translations of Le Petit Prince.
Outcome: The proposed method can be applied to Finnish and Latin translations of Le Petit Prince . it uses k-means clustering to group raw, contextualized BERT embeddings .

Similar Papers

J-SNACS: Adposition and Case Supersenses for Japanese Joshi (2024.lrec-main)

Copied to clipboard

Challenge: adpositions are used to mark a variety of semantic relations in languages such as English and Korean.
Approach: They propose a Japanese extension of the SNACS framework for annotating adpositions in corpora from several languages.
Outcome: The proposed framework captures similarities not seen in multilingual embedding space.
MASALA: Modelling and Analysing the Semantics of Adpositions in Linguistic Annotation of Hindi (2022.lrec-1)

Copied to clipboard

Challenge: Existing work on SNACS annotation for a variety of typologically diverse languages focuses on semantic role labelling and upstream applications in related languages.
Approach: They propose to use the multilingual SNACS annotation scheme to attempt automatic labelling of SNAC supersenses in Hindi.
Outcome: The proposed method is competitive with previous work on English and Gujarati.
A Corpus of Adpositional Supersenses for Mandarin Chinese (2020.lrec-1)

Copied to clipboard

Challenge: Adpositions are frequent markers of semantic relations, but they are highly ambiguous and vary significantly from language to language.
Approach: They propose to annotate Chinese adpositions in a corpus with all aforementioned supersenses . they adapt a framework that defined a set of supersens according to ostensibly language-independent criteria .
Outcome: The proposed corpus is the first to be broadly annotated with adposition semantics in Chinese . it shows that the supersense categories are well-suited to Chinese adepositions despite syntactic differences from English .
Multilingual Supervision Improves Semantic Disambiguation of Adpositions (2025.coling-main)

Copied to clipboard

Challenge: a corpus-based cross-linguistic investigation into the lexical semantics of adpositions is conducted . a significant amount of ambiguity and flexibility in their meanings are present in a variety of languages .
Approach: They conduct a corpus-based corpus analysis of adpositions using SNACS . they find distributional differences in a language's adequacy and disambiguation performance .
Outcome: The proposed framework is suited for analyzing adpositions across languages . it provides a framework for a wide-coverage corpus annotation of high-level senses .
Xposition: An Online Multilingual Database of Adpositional Semantics (2022.lrec-1)

Copied to clipboard

Challenge: Xposition is an online platform for documenting adpositional semantics across languages . SNACS provides a unified metalanguage for characterizing the major classes of meanings expressed with appositions .
Approach: They propose to use Xposition to document adpositional semantics across languages . Xpos houses annotation guidelines, structured lexicographic documentation, annotated corpora .
Outcome: The proposed platform houses annotation guidelines, structured lexicographic documentation, and annotated corpora.
Semantic Supersenses for English Possessives (L18-1)

Copied to clipboard

Challenge: Existing semantic categories for possessive constructions are limited to nominals and s-genitives.
Approach: They propose to use a supersense inventory to annotate English possessives . they show existing supersensor categories are readily applicable to possessives.
Outcome: The proposed annotations are applied to English possessives in a corpus of web reviews.
Comprehensive Supersense Disambiguation of English Prepositions and Possessives (P18-1)

Copied to clipboard

Challenge: Frequent prepositions like for are maddeningly polysemous, their interpretation depends especially on the object of the preposition.
Approach: They propose a new annotation scheme, corpus, and task for the disambiguation of prepositions and possessives in English.
Outcome: The proposed annotations are comprehensive with respect to types and tokens of these markers and use broadly applicable supersense classes rather than fine-grained dictionary definitions.
CaMEL: Case Marker Extraction without Labels (2022.acl-long)

Copied to clipboard

Challenge: Existing models for morphological case marking and semantic content are not isomorphic.
Approach: They propose a model that extracts case markers from a multilingual corpus using a noun phrase chunker and an alignment system.
Outcome: The proposed model can extract case markers in 83 languages and visualise similarities and differences between case systems and annotate fine-grained deep cases in languages where they are not overtly marked.
Dr. Livingstone, I presume? Polishing of foreign character identification in literary texts (2022.naacl-srw)

Copied to clipboard

Challenge: Current state-of-the-art models that use neural networks can help with character identification in agglutinative languages.
Approach: They propose to use a search for the shortest version of the name to identify the baseform of the character's lemma to align different appearances of the same character in the narrative.
Outcome: The proposed method is the easiest, best performing and resource-independent method.
More Embeddings, Better Sequence Labelers? (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing work suggests contextual embeddings improve sequence labeling accuracy . but, there is no definite conclusion on whether concatenating different kinds of embeddables is effective .
Approach: They propose a family of contextual embeddings that improves sequence labeling accuracy . they conduct extensive experiments on 3 tasks over 18 datasets and 8 languages .
Outcome: The proposed family of contextual embeddings improves the accuracy of sequence labelers over non-contextual embedders.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations