Papers with LCS

7 papers
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval (2026.eacl-srw)

Copied to clipboard

Challenge: Existing large language models are not designed for semantic retrieval and PDF-based legislative sources introduce substantial noise due to imperfect text extraction.
Approach: They propose a large-scale multilingual corpus of EU environmental legislation constructed from 24,953 official EUR-Lex PDF documents covering 25 languages.
Outcome: The proposed model improves Top-k retrieval accuracy in monolingual and bilingual settings . it also improves accuracy in low- and high-resource languages .
Rethinking Negative Instances for Generative Named Entity Recognition (2024.findings-acl)

Copied to clipboard

Challenge: Named Entity Recognition (NER) models are constrained by a pre-defined label set and require extensive human annotations, which limits their flexibility and adaptability to unseen tasks.
Approach: They propose a Generative NER system that shows improved zero-shot performance across unseen entity domains by introducing contextual information and delineating label boundaries.
Outcome: The proposed model outperforms state-of-the-art methods in zero-shot evaluation.
Active Learning for Rumor Identification on Social Media (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for rumor tracking depend on a significant amount of labeled data.
Approach: They propose an Active-Transfer Learning strategy to identify rumors with limited amount of annotated data.
Outcome: The proposed approach achieves faster convergence in terms of the F-score while requiring fewer annotated samples (42% of the whole dataset for the best model).
CaLcs: Continuously Approximating Longest Common Subsequence for Sequence Level Optimization (D18-1)

Copied to clipboard

Challenge: Maximum-likelihood estimation (MLE) is widely used for text-generation based natural language processing applications.
Approach: They propose a method to train models with maximum-likelihood estimation using a differentiable surrogate of longest common subsequence measure that captures sequence-level structure similarity.
Outcome: Experimental results show that the proposed approach improves on the current MLE approach for downstream tasks like text summarization and machine translation.
LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation (2024.findings-acl)

Copied to clipboard

Challenge: Existing LT strategies cannot indicate the desired target language on zero-shot translation, i.e., the off-target issue.
Approach: They propose a language converter strategy that embeds the target language into the top encoder layers to mitigate confusion in the encoder and ensures stable language indication for the decoder.
Outcome: The proposed language converter strategy significantly mitigates off-target issue on multiUN, TED, and OPUS-100 datasets.
Contextualized Semantic Distance between Highly Overlapped Texts (2023.findings-acl)

Copied to clipboard

Challenge: Conventional semantic metrics are based on word representations and are vulnerable to disturbance of overlapped components with similar representations.
Approach: They propose a mask-and-predict strategy to evaluate the semantic distance between the overlapped sentences using words in the longest common sequence as neighboring words and use masked language modeling to predict their positions.
Outcome: The proposed method outperforms the state-of-the-art in domain adaption by a huge margin.
GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for procedural planning over-rely on visual inputs and lack structured semantic information.
Approach: They propose a vision–language framework for multimodal procedural planning that exploits implicit spatial relations and deep semantics encoded in object attributes.
Outcome: The proposed framework outperforms existing methods in terms of execution success rate, LCS, and planning correctness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations