Papers with PE

25 papers
Systematic Analysis for Pretrained Language Model Priming for Parameter-Efficient Fine-tuning (2024.naacl-srw)

Copied to clipboard

Challenge: Parameter-efficient (PE) methods for adapting pre-trained language models to downstream tasks are still lacking in many cases.
Approach: They propose a general PE priming framework to enhance few-shot adaptation and generalization ability of PE methods.
Outcome: The proposed framework reveals that the best priming strategy facilitates adaptation to target tasks.
MMPE: A Multi-Modal Interface using Handwriting, Touch Reordering, and Speech Commands for Post-Editing Machine Translation (2020.acl-demos)

Copied to clipboard

Challenge: a shift from traditional translation to post-editing (PE) of machine-translated text can save time and reduce errors, but it also affects the design of translation interfaces.
Approach: They propose a prototype that combines traditional input modes with pen, touch, and speech modalities for post-editing of machine-translated (MT) they propose to use these modalités to cross out or hand-write new text, drag and drop words for reordering, or use spoken commands to update the text in place.
Outcome: The proposed interfaces can be used to cross out or hand-write new text, drag and drop words for reordering, or use spoken commands to update the text in place.
Sample Design Engineering: An Empirical Study on Designing Better Fine-Tuning Samples for Information Extraction with LLMs (2024.emnlp-industry)

Copied to clipboard

Challenge: Prompt Engineering (PE) is renowned for improving IE performance through prompt modifications, but the realm of sample design for downstream fine-tuning remains unexplored.
Approach: They propose a methodical approach to enhancing LLMs’ post-tuning performance by refining input, output, and reasoning designs.
Outcome: The proposed approach outperforms heuristic design strategies on three complex IE tasks with four additional LLMs.
Exploring the Potential of ChatGPT on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations (2024.findings-eacl)

Copied to clipboard

Challenge: Recent studies have demonstrated ChatGPT's remarkable few-shot, even zero-shot learning abilities when compared to other models.
Approach: They quantitatively evaluate the performance of ChatGPT on inter-sentential relations such as temporal relations, causal relations, and discourse relations.
Outcome: The proposed model performs well on temporal relations, causal relations, and discourse relations.
PE-QAT: Parameter-Efficient Quantization-Aware Training for Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Quantization Aware Training (QAT) is expensive to train and unscalable to large models.
Approach: They propose a parameter-efficient framework targeting per-channel 4-bit weight-activation quantization of large language models.
Outcome: The proposed framework preserves accuracy within 0.11 percentage points of the full-precision baseline on Llama-2-7B zero-shot tasks while training only 1.26% of total parameters.
DecBERT: Enhancing the Language Understanding of BERT with Causal Attention Masks (2022.findings-naacl)

Copied to clipboard

Challenge: Experimental results show that Transformer Encoder model can't automatically capture word order, so explicit position embeddings are required to be fed into the target model.
Approach: They propose a Transformer-based language model DecBERT that uses a causal attention mask to capture word order.
Outcome: The proposed model improves on the GLUE language understanding benchmark and accelerates the pre-training process.
Self-Attention with Cross-Lingual Position Representation (2020.acl-main)

Copied to clipboard

Challenge: Position encoding (PE) is used to preserve word order information for natural language processing tasks, generating fixed position indices for input sequences.
Approach: They propose to augment SANs with cross-lingual position representations to model bilingually aware latent structure for the input sentence.
Outcome: The proposed model significantly improves translation quality over baselines on EnglishGerman, JapaneseEnglish, and ChineseEnglish translation tasks.
MMPE: A Multi-Modal Interface for Post-Editing Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Current advances in machine translation (MT) increase the need for translators to switch from traditional translation to post-editing (PE) of machine-translated text.
Approach: They propose to combine traditional input modes with pen, touch, and speech modalities for post-editing of machine-translated text.
Outcome: The proposed interfaces are designed to reduce errors and save time.
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers (2026.findings-eacl)

Copied to clipboard

Challenge: Length generalization is the ability of language models to maintain performance on inputs longer than those seen during pretraining.
Approach: They propose a position encoding strategy that uses random float sampling to generalize to unseen lengths.
Outcome: The proposed strategy can generalize to lengths unseen during training and in benchmarks.
Incorporating Noisy Length Constraints into Transformer with Length-aware Positional Encodings (2020.coling-main)

Copied to clipboard

Challenge: Neural Machine Translation suffers from an under-translation problem due to limited modeling of output sequence lengths.
Approach: They propose a method to train a Transformer model using length constraints based on positional encoding.
Outcome: The proposed method outperforms a vanilla Transformer in an English-to-Japanese translation by 3.22 points . the noise injection improved robustness for length prediction errors, especially within the window size.
Automatic Post-Editing of Machine Translation: A Neural Programmer-Interpreter Approach (D18-1)

Copied to clipboard

Challenge: Existing approaches to inducing APE have suffered from over-correction, where the APE system tends to keep the machine translated text without any modification.
Approach: They propose a neural programmer-interpreter approach to automated post-editing (APE) that mimics human perform post- editing using discrete edit operations . their model outperforms previous neural models for inducing PE programs on the WMT17 APE task for German-English up to +1 BLEU score and -0.7 TER scores.
Outcome: The proposed model outperforms previous neural models for inducing PE programs on the WMT17 APE task for German-English up to +1 BLEU score and -0.7 TER scores.
MedicalSum: A Guided Clinical Abstractive Summarization Model for Generating Medical Reports from Patient-Doctor Conversations (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing models for summarizing medical conversations do not take clinical knowledge into account and are difficult to control.
Approach: They propose a transformer-based sequence-to-sequence architecture for summarizing medical conversations by integrating medical domain knowledge from the Unified Medical Language System (UMLS).
Outcome: The proposed model achieves state-of-the-art ROUGE score improvements of 0.8-2.1 points (including 6.2% error reduction in the PE section) it incorporates medical domain knowledge from the Unified Medical Language System (UMLS).
WeTS: A Benchmark for Translation Suggestion (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on overall performance of machine translation but ignore TS performance, authors say . if TS is applied into post-editing, it will reduce the time and cost of post-production.
Approach: They propose to use a golden corpus annotated by experts to generate a translation suggestion model.
Outcome: The proposed model improves on the golden corpus annotated by translators on four translation directions.
An Anchor-based Relative Position Embedding Method for Cross-Modal Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: Position Embedding (PE) is essential for transformer to capture the sequence ordering of input tokens.
Approach: They propose a unified position embedding method that bridges the semantic gap between modalities and embeds the anchor-based distance to guide computation of cross-attention.
Outcome: The proposed method obtains new SOTA results on a wide range of benchmarks.
Probing Simile Knowledge from Pre-trained Language Models (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to learn generic knowledge from a large corpus are time-consuming and labor-intensive.
Approach: They propose a framework to probe simile knowledge from pre-trained language models to solve SI and SG tasks.
Outcome: The proposed framework solves the SI and SG tasks in a simile triple completion task.
Argument Mining with Fine-Tuned Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Argument Mining (AM) pipelines use fine-tuned large language models (LLMs) . initial approaches employ supervised machine learning algorithms, such as Maximum Entropy classifiers and Logistic Regressions.
Approach: They propose to model the three main AM sub-tasks as text generation tasks and fine-tune eight popular quantized and non-quantized large language models (LLMs) on the benchmark PE, AbstRCT, and CDCP datasets.
Outcome: The proposed pipeline achieves state-of-the-art across all AM sub-tasks and datasets, showing significant improvements over previous benchmarks.
PE: A Poincare Explanation Method for Fast Text Hierarchy Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work on feature interactions neglects underlying linguistic information in feature representations.
Approach: They propose a method for modeling feature interactions with hyperbolic spaces using Poincare Explanation.
Outcome: The proposed method is able to model feature interactions with hyperbolic spaces in a time efficient manner.
Mid-Air Hand Gestures for Post-Editing of Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: In a well-connected world, translation is of everincreasing importance.
Approach: They propose to use mid-air hand gestures in combination with the keyboard for editing in machine translation and post-editing workflows to improve quality.
Outcome: The proposed prototype supports mid-air hand gestures for cursor placement, text selection, deletion, and reordering.
Evaluating Automatic Subtitling: Correlating Post-editing Effort and Automatic Metrics (2024.lrec-main)

Copied to clipboard

Challenge: Existing metrics for automatic subtitling are not yet fully explored.
Approach: They propose to use machine translation metrics to measure post-editing effort in automatic subtitling to collect data on product-, process- and participant-based data.
Outcome: The proposed metrics correlate with measures of post-editing effort in automatic subtitling.
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to enhance length extrapolation of large language models have been developed, but a systematic survey is lacking.
Approach: They propose to examine the effects of positional encoding on length extrapolation.
Outcome: The proposed methods improve the extrapolation of large language models, but they are still lacking a systematic survey.
Disentangle to Decay: Linear Attention with Trainable Decay Factor (2025.coling-main)

Copied to clipboard

Challenge: Existing linear attention models use a decay factor based positional encoding (PE), but the decay factor is manually designed and non-trainable, limiting further optimization.
Approach: They propose a PE-based positional encoding that disentangles decay factor into two parts to achieve further optimization and stable training.
Outcome: The proposed model achieves stable training of decay factor and improves inference efficiency in normal context and extrapolation scenarios.
Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation Output (2021.emnlp-main)

Copied to clipboard

Challenge: Post-editing (PE) machine translation (MT) output can save time and reduce errors.
Approach: They propose to use automatic word-level quality estimation to predict correctness of MT output to flag problematic output.
Outcome: The proposed model is not good enough to support human translations, but is based on a visualization reflecting uncertainty of the model.
How Real Are Synthetic Therapy Conversations? Evaluating Fidelity in Prolonged Exposure Dialogues (2025.findings-emnlp)

Copied to clipboard

Challenge: Synthetic data adoption in healthcare is driven by privacy concerns, data access limitations, and high annotation costs.
Approach: They compare real and synthetic PTSD therapy conversations using linguistic, structural, and protocol-specific metrics like turn-taking and treatment fidelity.
Outcome: The proposed framework assesses clinical fidelity beyond surface fluency.
Understanding How Positional Encodings Work in Transformer Model (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have reported superiority of relative PEs in translation tasks.
Approach: They analyze in which part of a transformer model PEs work and compare them using experiments . they find that relative PEs should be added only to query and key of attention mechanism .
Outcome: The results show that relative and absolute PEs work in a transformer model, and should be added to the query and key of an attention mechanism, not to the value.
GSM-Noise: Exploring and Enhancing Large Language Models’ Reasoning under Noisy Inputs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models struggle when dealing with complex, ill-formed, or noisy inputs . open-source models are less robust, while closed-source ones are more robust .
Approach: They propose to use GSM-Noise to refine inputs before engaging in in-depth analysis to improve LLM robustness under noisy conditions.
Outcome: The proposed model can achieve consistent performance gains under noisy conditions with prompt engineering, supervised finetuning, and reinforcement learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations