Towards a unified framework for bilingual terminology extraction of single-word and multi-word terms (C18-1)
Copied to clipboard
| Challenge: | Existing methods for extracting bilingual terminology from comparable corpora are limited to a set of syntactic patterns. |
| Approach: | They propose a framework for aligning bilingual terms independently of term lengths . they introduce some enhancements to the context-based and neural network based approaches . |
| Outcome: | The proposed framework improves the performance of the context-based and neural network based approaches and can be adapted in specialized domains. |
Similar Papers
Building Comparable Corpora for Assessing Multi-Word Term Alignment (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to extract bilingual terminologies from corpora are limited . MWTs pose serious challenges for alignment and machine translation systems . |
| Approach: | They propose an approach to build comparable corpora and bilingual term dictionaries that evaluate bilingual term alignment in comparable corpus. |
| Outcome: | The proposed method is validated on an existing dataset and manually annotated data. |
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)
Copied to clipboard
| Challenge: | Terms are notoriously difficult to identify, both automatically and manually. |
| Approach: | They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information . |
| Outcome: | The proposed method provides a tool for evaluation and rich source of information about terms. |
Leveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora (C18-1)
Copied to clipboard
| Challenge: | Recent studies on bilingual lexicon extraction from specialized comparable corpora show differences in performance . lack of large specialized corporan to build efficient representations can be partially explained . |
| Approach: | They propose to use character-based embedding models to combine different embeddable models . they emphasize how character-driven embeddance models outperform other models on quality . |
| Outcome: | The proposed model outperforms other models on quality of extracted bilingual lexicons . comparable corpora are an interesting and practical alternative to parallel corporation . |
Multilingualization of Medical Terminology: Semantic and Structural Embedding Approaches (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for multilingual terminology curation are limited as they do not fit the term within existing terminology. |
| Approach: | They propose a method to encode the structural property of a term by aligning embeddings using graph convolutional networks trained from separate languages. |
| Outcome: | The proposed method can encode the structural property of a term by aligning embeddings using graph convolutional networks trained from separate languages. |
Using English Baits to Catch Serbian Multi-Word Terminology (L18-1)
Copied to clipboard
| Challenge: | a new method for bilingual terminology extraction is proposed for a source language and a target language. |
| Approach: | They propose to use a bilingual terminology extraction approach for a source language and a target language to extract the terminology for sri lanka. |
| Outcome: | The proposed method extracts terminology for a source language and a target language from it. |
Cross-lingual Terminology Extraction for Translation Quality Estimation (L18-1)
Copied to clipboard
| Challenge: | Using common statistical measures for termhood and unithood, we identify terms from monolingual texts and investigate the contribution of terminology to translation quality. |
| Approach: | They propose to use common statistical measures for termhood and unithood as features to train classifiers for identifying terms in cross-domain and cross-language settings. |
| Outcome: | The proposed method has shown some reliability in automatically identifying terms in human translations, but drawbacks in handling low frequency terms and term variations shall be dealt with in the future. |
A Hybrid Approach for Automatic Extraction of Bilingual Multiword Expressions from Parallel Corpora (L18-1)
Copied to clipboard
| Challenge: | Specific-domain bilingual lexicons are composed of MultiWord Expressions (MWEs) the manual construction of MWEs bilingual dictionaries is costly and time-consuming. |
| Approach: | They propose to use word alignment approaches to automatically construct bilingual lexicons of MWEs from parallel corpora by formalizing the alignment process as an integer linear programming problem. |
| Outcome: | The proposed approach extracts and aligns multiword expressions from parallel corpora and then filters them using linguistic patterns to build bilingual lexicons. |
Transforming Term Extraction: Transformer-Based Approaches to Multilingual Term Extraction Across Domains (2021.findings-acl)
Copied to clipboard
| Challenge: | Automated Term Extraction (ATE) is a challenging task, with few exceptions. |
| Approach: | They propose to use a transformer-based term extraction model to extract terms from sentences . they also propose to employ a language model for token classification and a sequence model to reduce sentences to terms . |
| Outcome: | The proposed models outperform baselines on the ATE challenge TermEval 2020 dataset in English, French, and Dutch. |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
A Closer Look at Clustering Bilingual Comparable Corpora (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for clustering comparable corpora are not suitable for bilingual corpors. |
| Approach: | They propose new clustering models fully adapted to comparable corpora based on a deep variant of Kmeans . they illustrate their behavior on bilingual collections created from Wikipedia . |
| Outcome: | The proposed models show that they can cluster comparable corpora on bilingual collections . the proposed models are based on a state-of-the-art deep variant of Kmeans . |