Using English Baits to Catch Serbian Multi-Word Terminology (L18-1)

Copied to clipboard

Challenge: a new method for bilingual terminology extraction is proposed for a source language and a target language.
Approach: They propose to use a bilingual terminology extraction approach for a source language and a target language to extract the terminology for sri lanka.
Outcome: The proposed method extracts terminology for a source language and a target language from it.

Similar Papers

Towards a unified framework for bilingual terminology extraction of single-word and multi-word terms (C18-1)

Copied to clipboard

Challenge: Existing methods for extracting bilingual terminology from comparable corpora are limited to a set of syntactic patterns.
Approach: They propose a framework for aligning bilingual terms independently of term lengths . they introduce some enhancements to the context-based and neural network based approaches .
Outcome: The proposed framework improves the performance of the context-based and neural network based approaches and can be adapted in specialized domains.
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.
Multi-lingual Entity Discovery and Linking (P18-5)

Copied to clipboard

Challenge: This tutorial reviews the framework of cross-lingual EL and motivates it as a broad paradigm for the Information Extraction task.
Approach: This tutorial will review the framework of cross-lingual EL and motivate it as a broad paradigm for the Information Extraction task.
Outcome: The aim of this tutorial is to review the framework of cross-lingual EL and motivate it as a broad paradigm for the Information Extraction task.
Cross-lingual and Cross-domain Transfer Learning for Automatic Term Extraction from Low Resource Data (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Term Extraction (ATE) is a key component for domain knowledge understanding and can be used for further NLP applications.
Approach: They propose to fine-tune pre-trained BERT models for automatic Term Extraction (ATE) using cross-lingual and cross-domain transfer learning to extract single and multi-word terms.
Outcome: The proposed models can capture cross-domain and cross-lingual terminologically-marked contexts shared by terms, opening a new design-pattern for ATE.
Crossing Domains without Labels: Distant Supervision for Term Extraction (2025.emnlp-industry)

Copied to clipboard

Challenge: Current state-of-the-art methods require expensive human annotation and struggle with domain transfer, limiting their practical deployment.
Approach: They propose a benchmark spanning seven diverse domains to evaluate ATE performance . they propose psuedo-labels and post-hoc heuristics to ensure generalizability .
Outcome: The proposed model outperforms supervised cross-domain encoder models and few-shot learning baselines on the document- and corpus-levels and its GPT-4o teacher on the benchmark.
Multilingualization of Medical Terminology: Semantic and Structural Embedding Approaches (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for multilingual terminology curation are limited as they do not fit the term within existing terminology.
Approach: They propose a method to encode the structural property of a term by aligning embeddings using graph convolutional networks trained from separate languages.
Outcome: The proposed method can encode the structural property of a term by aligning embeddings using graph convolutional networks trained from separate languages.
Transforming Term Extraction: Transformer-Based Approaches to Multilingual Term Extraction Across Domains (2021.findings-acl)

Copied to clipboard

Challenge: Automated Term Extraction (ATE) is a challenging task, with few exceptions.
Approach: They propose to use a transformer-based term extraction model to extract terms from sentences . they also propose to employ a language model for token classification and a sequence model to reduce sentences to terms .
Outcome: The proposed models outperform baselines on the ATE challenge TermEval 2020 dataset in English, French, and Dutch.
Leveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora (C18-1)

Copied to clipboard

Challenge: Recent studies on bilingual lexicon extraction from specialized comparable corpora show differences in performance . lack of large specialized corporan to build efficient representations can be partially explained .
Approach: They propose to use character-based embedding models to combine different embeddable models . they emphasize how character-driven embeddance models outperform other models on quality .
Outcome: The proposed model outperforms other models on quality of extracted bilingual lexicons . comparable corpora are an interesting and practical alternative to parallel corporation .
word2word: A Collection of Bilingual Lexicons for 3,564 Language Pairs (2020.lrec-1)

Copied to clipboard

Challenge: Our dataset provides top-k word translations in 3,564 (directed) language pairs across 62 languages in OpenSubtitles2018.
Approach: They propose a dataset and an open-source Python package for cross-lingual word translations extracted from sentence-level parallel corpora.
Outcome: The proposed bilingual lexicons have high coverage and achieve competitive translation quality for several language pairs.
MINERS: Multilingual Language Models as Semantic Retrievers (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks have evaluated language models to evaluate their performance across a range of embedding tasks.
Approach: They propose a benchmark to evaluate the robustness of multilingual language models in semantic retrieval tasks including bitext mining and classification via retrieval-augmented contexts.
Outcome: The proposed framework evaluates the robustness of multilingual LMs in retrieval tasks across over 200 languages, including extremely low-resource languages in challenging cross-lingual and code-switching settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations