Challenge: a recent study has focused on term technicality, but there are still few studies on it.
Approach: They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces .
Outcome: The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons.

Similar Papers

A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.
A Laypeople Study on Terminology Identification across Domains and Task Definitions (N18-2)

Copied to clipboard

Challenge: Existing studies on term annotation show that even experts differ in their understanding of termhood .
Approach: They propose a new dataset of term annotation that examines the common understanding of what constitutes a term.
Outcome: The proposed datasets show that even experts differ in their understanding of termhood . the findings suggest that there is a common understanding of what constitutes a term .
Automatic Annotation of Semantic Term Types in the Complete ACL Anthology Reference Corpus (L18-1)

Copied to clipboard

Challenge: a recent increase in quantitative studies of scientific text collections has led to a significant increase in the use of semantic labeling techniques.
Approach: They propose to use semantic class labels to enhance a well-known resource . they use semantic labels to assign semantic class labeling to technical terms .
Outcome: The proposed approach enhances the ACL Anthology Reference Corpus with semantic class labels for 20,000 technical terms . the goal is to use this information as one feature in the profiling of scientific papers, communities, and disciplines.
Unsupervised Term Extraction for Highly Technical Domains (2022.emnlp-industry)

Copied to clipboard

Challenge: Term extraction is an important task for knowledge discovery platforms because domain specific terms are the linguistic representation of domainspecific concepts.
Approach: They propose a term extraction subsystem that uses an unsupervised annotator to generate training data to fine-tune transformer models.
Outcome: The proposed system can generalize across domains while reducing latency and inference time while preserving the high performance of the existing system.
Varying Vector Representations and Integrating Meaning Shifts into a PageRank Model for Automatic Term Extraction (2020.lrec-1)

Copied to clipboard

Challenge: a comparative study for automatic term extraction from domain-specific language using a PageRank graph algorithm with different edge-weighting methods.
Approach: They propose to use a PageRank algorithm to extract automatic terms from domain-specific language using different edge-weighting methods.
Outcome: The proposed model is compared with a PageRank model with different edge-weighting methods.
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)

Copied to clipboard

Challenge: lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems.
Approach: They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation.
Outcome: The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers.
Extracting Text Representations for Terms and Phrases in Technical Domains (2023.acl-industry)

Copied to clipboard

Challenge: Large pre-trained language models are extensively used in modern NLP systems.
Approach: They propose an unsupervised approach to encoding using character-based models and pre-trained sentence encoders to reconstruct large pre-trained embedding matrices.
Outcome: The proposed approach matches the quality of sentence encoders in technical domains and is 5 times smaller and up to 10 times faster on high-end GPUs.
Quantifying Compositionality of Classic and State-of-the-Art Embeddings (2025.findings-emnlp)

Copied to clipboard

Challenge: Static word embeddings make strong claims about compositionality, but the SOTA generative models go too far in the other direction.
Approach: a new study evaluates the compositionality of word embeddings by canonical correlation analysis . strong compositional signals are observed in later training stages across data modalities .
Outcome: a new evaluation of compositional models shows that they exploit access meanings when justified . strong compositional signals are observed in later training stages and in deeper layers of the transformer-based model before a decline at the top layer.
Leveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora (C18-1)

Copied to clipboard

Challenge: Recent studies on bilingual lexicon extraction from specialized comparable corpora show differences in performance . lack of large specialized corporan to build efficient representations can be partially explained .
Approach: They propose to use character-based embedding models to combine different embeddable models . they emphasize how character-driven embeddance models outperform other models on quality .
Outcome: The proposed model outperforms other models on quality of extracted bilingual lexicons . comparable corpora are an interesting and practical alternative to parallel corporation .
Cross-lingual Terminology Extraction for Translation Quality Estimation (L18-1)

Copied to clipboard

Challenge: Using common statistical measures for termhood and unithood, we identify terms from monolingual texts and investigate the contribution of terminology to translation quality.
Approach: They propose to use common statistical measures for termhood and unithood as features to train classifiers for identifying terms in cross-domain and cross-language settings.
Outcome: The proposed method has shown some reliability in automatically identifying terms in human translations, but drawbacks in handling low frequency terms and term variations shall be dealt with in the future.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations