Papers with MWTs
Word Embedding Approach for Synonym Extraction of Multi-Word Terms (L18-1)
Copied to clipboard
| Challenge: | MWTs are motivated combinations that clearly convey the concept they designate. |
| Approach: | They propose a word-embedding-based approach for automatic acquisition of MWT synonyms that manage length variability. |
| Outcome: | The proposed approach improves on two specialized domain corpora and shows that it is more efficient than baseline approaches. |
Multi-word Tokenization for Sequence Compression (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Large Language Models have proven successful at modelling tasks, but they are expensive and slow to scale. |
| Approach: | They propose a Multi-Word Tokenizer that represents frequent multi-word expressions as single tokens. |
| Outcome: | The proposed tokenizer is more robust across shorter sequence lengths, allowing for major speedups via early sequence truncation. |
Representing Multiword Term Variation in a Terminological Knowledge Base: a Corpus-Based Study (2020.lrec-1)
Copied to clipboard
| Challenge: | Multiword terms are the most frequent type of lexical units in scientific and technical communication. rendering them in another language is not easy due to their cognitive complexity, proliferation of different forms, and their unsystematic representation in terminographic resources. |
| Approach: | They evaluated Spanish translation variants of multiword terms in three parallel corpora, two comparable corporales and two terminological resources. |
| Outcome: | The results show that multiword terms exhibit a significant degree of term variation . the proposed model is based on a set of criteria for determining which variants should be selected . |
Building Comparable Corpora for Assessing Multi-Word Term Alignment (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to extract bilingual terminologies from corpora are limited . MWTs pose serious challenges for alignment and machine translation systems . |
| Approach: | They propose an approach to build comparable corpora and bilingual term dictionaries that evaluate bilingual term alignment in comparable corpus. |
| Outcome: | The proposed method is validated on an existing dataset and manually annotated data. |