A Deep Factorization of Style and Structure in Fonts (D19-1)

Copied to clipboard

Challenge: Using a variational inference procedure, we factor each training glyph into a combination of a character-specific content embedding and a latent font-specific style variable.
Approach: They propose a deep factorization model that disentangles content from style by factorizing each training glyph into a latent content embedding and a learned embeddable character.
Outcome: The proposed model outperforms a strong nearest neighbors baseline and state-of-the-art discriminative model on reconstructing missing glyphs from an unknown font given only a small number of observations.

Similar Papers

Scalable Font Reconstruction with Dual Latent Manifolds (2021.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that fonts with a large number of missing glyphs are difficult to model due to the relative sparsity of most fonts.
Approach: They propose a deep generative model that performs typography analysis and font reconstruction by learning disentangled manifolds of both font style and character shape.
Outcome: The proposed model scales up the number of character types we can model compared to previous methods . it can generalize to characters that were not observed during training time, and it compares favorably to other models .
A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing (2020.acl-main)

Copied to clipboard

Challenge: Scholars often need to go beyond textual analysis for establishing provenance of historical documents.
Approach: They propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents by generating a latent vector responsible for inking variations, jitter, noise and other unforeseen phenomena.
Outcome: The proposed model outperforms interpretable clustering baselines and overly-flexible deep generative models on the task of completely unsupervised discovery of typefaces in mixed-fonts documents.
Learning Interpretable Style Embeddings via Prompting LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Prior work has treated the style of a text as separable from the content.
Approach: They use prompting to perform stylometry on a large number of texts to generate a synthetic stylometric dataset.
Outcome: The proposed model trains human-interpretable representations on a large stylometric dataset and a linguistic model for style representation learning.
Disentangled Representation Learning for Non-Parallel Text Style Transfer (P19-1)

Copied to clipboard

Challenge: a paper aims to disentangle latent representations of style and content in language models . auxiliary multi-task and adversarial objectives are used to disentangle the latent space .
Approach: They propose a simple yet effective approach to disentangling latent representations . they propose auxiliary multi-task and adversarial objectives to disentangle style and content .
Outcome: The proposed approach achieves high performance in terms of transfer accuracy, content preservation, and language fluency compared to previous approaches .
Style Vectors for Steering Generative Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) can be trained on vast corpora and can generate text in a nuanced and parameterisable way.
Approach: They propose to add style vectors to the activations of hidden layers during text generation to steer output towards specific styles.
Outcome: The proposed approach differs from prompt engineering in that it can be nuanced and parameterisable.
TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted Restyling (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for text style transfer require style-labeled training data, but use only labeled data at inference time.
Approach: They propose a method that uses readily-available unlabeled text to train style transfer . they use a style vector to condition a decoder to perform style transfer using unlabelled text .
Outcome: The proposed method is competitive on sentiment transfer, even compared to models trained fully on labeled data.
Deep Latent Variable Models of Natural Language (D18-3)

Copied to clipboard

Challenge: In this tutorial, we will discuss the challenges of applying neural variational inference to NLP problems.
Approach: The tutorial will cover deep latent variable models in the case where exact inference over the latent variables is tractable.
Outcome: The proposed tutorial will cover deep latent variable models in the case where inference cannot be performed tractably and when it is not .
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for embedding text are limited by the imperfect nature of data acquired under such assumptions.
Approach: They propose a new approach to training stronger content-independent style embeddings using a synthetic dataset of near-exact paraphrases with controlled style variations.
Outcome: The proposed model outperforms existing methods in real-world benchmarks and outperformed leading style representations in downstream applications.
Compositional Generalization by Factorizing Alignment and Translation (2020.acl-srw)

Copied to clipboard

Challenge: a crucial property underlying the expressive power of human language is its systematicity.
Approach: They propose to make an analogous separation between alignment and translation in neural machine translation to capture compositional structure.
Outcome: The proposed architecture outperforms existing neural networks on a compositional generalization task without supervision.
Deep Generative Model for Joint Alignment and Word Representation (N18-1)

Copied to clipboard

Challenge: EmbedAlign model embeds words in their complete observed context and learns by marginalisation of latent lexical alignments.
Approach: They exploit translation as a distributional context and embed words as posterior probability densities, rather than point estimates, which allows them to compare words in context using a measure of overlap between distributions.
Outcome: The proposed model performs on a range of lexical semantics tasks and achieves competitive results on benchmarks including natural language inference, paraphrasing, and text similarity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations