Challenge: Schema translation is not well studied in the community because of morphological difference and context difference between plain text and tabular data.
Approach: They propose a schema translation model augmented with schema context . they model a target header and its context as a directed graph to represent their entities .
Outcome: The proposed model outperforms state-of-the-art models on schema translation . it uses a graph to represent entity types and relations, and a relational-aware transformer .

Similar Papers

Transformers for Tabular Data Representation: A Survey of Models and Applications (2023.tacl-1)

Copied to clipboard

Challenge: Recent research efforts extend LMs by developing neural representations for structured data.
Approach: They propose to extend transformer-based language models to tabular data by analyzing inputs, model training, and supported downstream tasks.
Outcome: The proposed models are compared against existing models and are based on a traditional pipeline.
Tagged Back-translation Revisited: Why Does It Really Work? (2020.acl-main)

Copied to clipboard

Challenge: In this paper, we show that neural machine translation systems trained on large back-translated data overfit some of the characteristics of machine-transcribed texts.
Approach: They propose to add a tag to back-translations to help distinguish back-translated data from original parallel training data.
Outcome: The proposed tag helps the system distinguish back-translated data from original parallel training data and is as effective as a tag in high-resource training.
Towards Modeling the Style of Translators in Neural Machine Translation (2021.naacl-main)

Copied to clipboard

Challenge: a key ingredient of neural machine translation is the use of large datasets with different but consistent translation styles . however, the models do not capture the variety of translators' styles from the data . a recent study shows that style-augmented models can capture the style variations of translator .
Approach: They propose to augment a neural machine translation model with translator information . they use TED talk datasets to model and control translator-related stylistic variations .
Outcome: The proposed models capture the style variations of translators and generate translations with different styles on new data.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Context-aware Neural Machine Translation with Coreference Information (D19-65)

Copied to clipboard

Challenge: Existing models for translating a sentence in a text do not consider coreference relations provided within the text.
Approach: They propose a graph-based encoder which can consider coreference relations provided within the text explicitly.
Outcome: The proposed model improves on the previous approach by 0.9 points on the BLEU score . the graph-based encoder can handle a longer text well, compared with the previous model .
Storytelling from Structured Data and Knowledge Graphs : An NLG Perspective (P19-4)

Copied to clipboard

Challenge: tutorial aims to explain the basic concepts of translating structured data into natural language . Various solutions for structured data translation will be discussed .
Approach: tutorial aims to cover foundational, methodological, and system development aspects of translating structured data into natural language . Various solutions starting from traditional rule based/heuristic driven and modern data-driven will be discussed .
Outcome: The tutorial aims to convey challenges and nuances in structured data translation, data representation techniques, and domain adaptable solutions for translation of the data into natural language form.
Challenges in Context-Aware Neural Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: despite well-reasoned intuitions, most context-aware neural machine translation models show only modest improvements over sentence-level systems.
Approach: They propose a more realistic setting for document-level translation called paragraph-to-paragraph (PARA2PARA) they collect a dataset of Chinese-English novels to promote future research .
Outcome: The proposed model improves translation quality across document-level metrics and discourse phenomena.
ASTRA: Automatic Schema Matching using Machine Translation (2024.emnlp-industry)

Copied to clipboard

Challenge: Currently, eCommerce platforms use schema matching to structure product information from disparate sources.
Approach: They propose to model the schema matching problem as a neural machine translation task . they propose to use open-source seq2seq models fine-tuned on product attribute mappings to build a framework .
Outcome: The proposed model achieves a significant performance boost (15% precision and 7% recall uplift) it can support new attributes with precision 95% using only five labeled samples per attribute.
CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing code translation datasets focus on a single pair of programming languages . early software systems are developed using programming languages such as Fortran and COBOL .
Approach: They propose a large-scale comprehensive benchmark that supports the largest variety of programming languages for code translation.
Outcome: The proposed framework supports translations between multiple programming languages and a cross-framework dataset for deep learning code across different frameworks.
Rethinking Document-level Neural Machine Translation (2022.findings-acl)

Copied to clipboard

Challenge: Neural machine translation models are weak enough for document-level translation . current models only translate sentences individually, resulting in poor document coherence .
Approach: They propose to use the original Transformer model to test document-level neural machine translation . they find that the original transformer models can achieve strong results for document translation if trained properly .
Outcome: The proposed model outperforms sentence-level models on nine datasets and two sentence- level datasets across six languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations