Papers with processing

64 papers
ReasonGraph: Visualization of Reasoning Methods and Extended Inference Paths (2025.acl-demo)

Copied to clipboard

Challenge: Large Language Models (LLMs) reasoning processes are complex and lack of organized visualization tools creates barriers to understanding, evaluation, and improvement.
Approach: They propose a web-based platform for visualizing and analyzing LLM reasoning processes.
Outcome: The proposed platform shows high parsing reliability, efficient processing, and excellent usability across various downstream applications.
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can explain quantum mechanics and write sophisticated code, yet fail to parse sentences like "The cat that the mouse feared chased meowed"
Approach: They propose a framework to distinguish structural understanding from semantic pattern matching . they use a set of 9,720 comprehension questions on center-embedded sentences .
Outcome: a new framework shows that models lose performance when they abandon structural analysis for semantic associations.
AutoRE: Document-Level Relation Extraction with Large Language Models (2024.acl-demos)

Copied to clipboard

Challenge: Existing methods for relation extraction are limited to Sentence-level Relation Extraction (SentRE) tasks.
Approach: They propose an end-to-end DocRE model that adopts a novel RE extraction paradigm named RHF (Relation-Head-Facts) Unlike existing approaches, AutoRE does not rely on the assumption of known relation options, making it more reflective of real-world scenarios.
Outcome: The proposed model surpasses TAG by 10.03% and 9.03% on the dev and test set.
Data Management Plan (DMP) for Language Data under the New General Da-ta Protection Regulation (GDPR) (L18-1)

Copied to clipboard

Challenge: ELRA proposes its own template for the Data Management Plan, which is being updated to take the new law into account.
Approach: They propose a framework for the data management plan to be updated to take the new law into account and propose how it can be integrated into the DMP to increase transparency and spread good practices .
Outcome: The proposed framework will strengthen certain principles related to the processing of personal data, which will also affect many projects in the field of natural language processing.
Gaining Insights into Unrecognized User Utterances in Task-Oriented Dialog Systems (2022.emnlp-industry)

Copied to clipboard

Challenge: Goal-oriented dialog systems fail to recognize the intent of natural language requests due to system errors, incomplete service coverage, or insufficient training.
Approach: They propose an end-to-end pipeline for processing unrecognized user utterances, deployed in a commercial task-oriented dialog system, including a specifically-tailored clustering algorithm, a novel approach to cluster representative extraction, and cluster naming.
Outcome: The proposed components show that they improve the performance of the proposed system in the analysis of unrecognized user requests.
Text Extraction and Script Completion in Images of Arabic Script-Based Calligraphy: A Thesis Proposal (2025.naacl-srw)

Copied to clipboard

Challenge: despite its artistic elements, Arabic calligraphy is difficult to read, even for those fluent in Arabic.
Approach: They analyze the variability in calligraphic styles and the influence of artistic distortions to improve text extraction and script completion.
Outcome: The proposed methods improve text extraction and script completion in Arabic calligraphy . the authors show that the proposed techniques are more efficient than traditional methods .
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters .
Approach: They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters .
Outcome: The proposed pipeline is optimized for high performance cluster environments and runs in high performance.
Incremental processing of noisy user utterances in the spoken language understanding task (D19-55)

Copied to clipboard

Challenge: triggered actions with high executions times can cause dialog systems to react slowly due to high latency and high latex.
Approach: They propose a model-agnostic method to achieve high quality in processing incrementally produced partial utterances.
Outcome: The proposed method improves the metric F1-score by 47.91 percentage points . the proposed method can be used to create low-latency natural language understanding components on ATIS datasets.
Conversing with databases: Practical Natural Language Querying (2023.emnlp-industry)

Copied to clipboard

Challenge: Large amount of companies' data is stored in relational databases . quick hypotheses validation is rarely, if ever, possible for majority of nontechnical business stakeholders.
Approach: They propose a hybrid NLQ system for conversational DB querying that allows non-technical users to formulate data requests as natural language questions.
Outcome: The proposed system is based on a hybrid NLQ (Natural Language Querying) system for conversational DB querying.
Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches (P19-2)

Copied to clipboard

Challenge: a study using non-canonical text normalization shows that it can surpass the current best performing system by a large margin.
Approach: They propose a fully automated, context-aware machine translation approach with fewer stages of processing.
Outcome: The proposed approach surpasses the current best-performing system by a large margin . the proposed method is more data-hungry and more data sensitive than other methods .
TabGenie: A Toolkit for Table-to-Text Generation (2023.acl-demo)

Copied to clipboard

Challenge: TabGenie enables researchers to explore, preprocess, and analyze data-to-text generation datasets.
Approach: They present TabGenie, a toolkit which enables researchers to explore, preprocess, and analyze a variety of data-to-text generation datasets.
Outcome: The toolkit provides an interactive mode for debugging table-to-text generation, side-by-side comparison of generated system outputs, and easy exports for manual analysis.
CLCL: Non-compositional Expression Detection with Contrastive Learning and Curriculum Learning (2023.acl-long)

Copied to clipboard

Challenge: Non-compositional expressions are a substantial challenge for natural language processing systems, necessitating more intricate processing compared to general language tasks.
Approach: They propose a dynamic curriculum learning framework specifically designed to take advantage of scarce available training data for modeling non-compositionality.
Outcome: The proposed framework improves on idiom usage recognition and metaphor detection tasks.
Gender Bias in Machine Translation (2021.tacl-1)

Copied to clipboard

Challenge: Interest in understanding, assessing, and mitigating gender bias in machine translation (MT) still lacks cohesion.
Approach: They propose to review current conceptualizations of gender bias in machine translation (MT) they summarize previous studies and propose ways to mitigate bias.
Outcome: This paper summarizes the current conceptualizations and proposes strategies to mitigate biases in machine translation (MT) .
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness (2025.findings-emnlp)

Copied to clipboard

Challenge: Multi-modal large language models have been used for processing and understanding information from diverse modalities.
Approach: They propose to evaluate the audio-visual capabilities of multi-modal large language models . they focus on effectiveness, efficiency, generalizability, and robustness .
Outcome: The proposed models exhibit strong zero-shot and few-shot generalization abilities . their success relies heavily on the vision modality, which impairs performance when visual input is corrupted or missing.
Bratly: A Python Extension for BRAT Functionalities (2025.emnlp-demos)

Copied to clipboard

Challenge: BRAT is a widely used web-based text annotation tool, but lacks robust Python support for effective annotation management and processing.
Approach: They propose an open-source extension of BRAT that introduces a solid Python backend and enables advanced annotation functions such as annotation typings, collection typings with statistical insights, corpus and annotation handling, object modifications, and entity-level evaluation.
Outcome: The proposed extension streamlines annotation workflows, improves usability, and facilitates high-quality NLP research.
Multimodal fusion via cortical network inspired losses (2022.acl-long)

Copied to clipboard

Challenge: Recent work in deep fusion models has led to substantial improvements over unimodal approaches in areas like speech recognition, emotion recognition and analysis.
Approach: They propose to introduce neural dependencies into the loss functions to allow for fusion of different modalities while keeping the model complexity manageable.
Outcome: Experiments on multimodal sentiment analysis tasks show that the proposed approach provides a consistent performance boost.
GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification (P19-1)

Copied to clipboard

Challenge: Existing methods to extract information from evidence are unable to grasp relational and logical information among the evidence.
Approach: They propose a graph-based evidence aggregating and reasoning framework to integrate evidence from multiple pieces of evidence.
Outcome: The proposed framework achieves significant performance improvements on a large-scale benchmark dataset.
Systematic Evaluation of Long-Context LLMs on Financial Concepts (2024.emnlp-industry)

Copied to clipboard

Challenge: Long-context large language models (LC LLMs) are promising for tasks with long context windows . however, their ability to reliably utilize their growing context windows remains under investigation .
Approach: They evaluate the performance of long-context large language models using a real-world financial news dataset.
Outcome: The proposed models exhibit brittleness at longer context lengths even for simple tasks, the authors show . they advocate for more rigorous evaluation of LC LLMs by employing holistic metrics such as F1 (rather than recall)
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)

Copied to clipboard

Challenge: Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed.
Approach: They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data.
Outcome: The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration.
Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Cantonese is considered a low-resource language due to the dominance of Mandarin . rich colloquial vocabulary of Cantone, English loanwords, and code-switching characteristics add to the complexity of corpus collection and processing.
Approach: We collect Cantonese texts from open source corpora, Hong Kong-specific forums, Wikipedia . we refine the model through supervised fine-tuning on curated Cantonesian tasks .
Outcome: The model achieves state-of-the-art (SOTA) performance on four Cantonese benchmarks.
Metacognitive Prompting Improves Understanding in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in prompting have enhanced reasoning in logic-intensive tasks for LLMs, yet the nuanced understanding abilities of these models remain underexplored.
Approach: They propose a strategy inspired by human introspective reasoning processes to enhance LLMs' understanding abilities.
Outcome: The proposed method outperforms chain-of-thought prompting and its advanced versions on ten natural language understanding (NLU) datasets.
Sorting through the noise: Testing robustness of information processing in pre-trained language models (2021.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive performance on downstream NLP tasks, but we have yet to establish a clear understanding of their sophistication when it comes to processing, retaining, and applying information presented in their input.
Approach: They examine how robustly pre-trained LMs retain and apply relevant context information in the face of distracting content.
Outcome: The proposed models retain and use critical context information in the face of distracting content, while models are susceptible to factors of semantic similarity and word position.
DoLFIn: Distributions over Latent Features for Interpretability (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to interpret neural networks face a trade-off between a model's usefulness and its complexity.
Approach: They propose a novel approach to achieve interpretability that avoids this trade-off by using probability as the central quantity instead of a fixed quantity.
Outcome: The proposed approach outperforms the classical CNN and BiLSTM classifiers on the SST2 and AG-news datasets.
TED-Q: TED Talks and the Questions they Evoke (2020.lrec-1)

Copied to clipboard

Challenge: Evoked questions represent a hitherto unexplored type of linguistic data, promising to open up important new lines of research.
Approach: They propose a method to annotate TED-talks with the questions they evoke and, where available, the answers to these questions.
Outcome: The proposed method is designed to scale up, relying on crowdsourcing by non-expert annotators, with its utility for Natural Language Processing in mind.
INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models perform well on unseen tasks in English, but their abilities in non-English languages are less explored due to limited benchmarks and training data.
Approach: They propose to release a large dataset for context-grounded question answering in 11 major Indian languages.
Outcome: The Indic-QA Benchmark compared large datasets of large LLMs on extractive and abstractive tasks in 11 major Indian languages.
Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation (2026.findings-acl)

Copied to clipboard

Challenge: Talk-to-Your-Slides is a high-efficiency slide editing agent that uses language-driven structured data manipulation instead of the image modality.
Approach: They propose a language-driven slide editing agent that uses language-based structured data manipulation instead of image modality.
Outcome: The proposed system achieves faster processing and better instruction fidelity than GUI-based agents.
Prompting ChatGPT in MNER: Enhanced Multimodal Named Entity Recognition with Auxiliary Refined Knowledge (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to enhance textual entity prediction neglect the need for external knowledge or encounter high redundancy in the retrieved knowledge.
Approach: They propose a framework that leverages ChatGPT as an implicit knowledge base and heuristically generates auxiliary knowledge for more efficient entity prediction.
Outcome: The proposed framework outperforms state-of-the-art methods on two classic datasets and exhibits a stronger robustness and generalization capability.
DocQueryNet: Value Retrieval with Arbitrary Queries for Form-like Documents (2022.coling-1)

Copied to clipboard

Challenge: Existing methods that only address a fixed set of fields are difficult to use for different form types.
Approach: They propose a value retrieval method with arbitrary queries for form-like documents . they propose 'docQueryNet' to predict target value based on understanding of layout and semantics of a form .
Outcome: The proposed method outperforms existing methods on value retrieval . it improves document understanding on large-scale model pre-training by 17% .
Automated Refugee Case Analysis: A NLP Pipeline for Supporting Legal Practitioners (2023.findings-acl)

Copied to clipboard

Challenge: In Canada, retrieving similar cases and their analysis is a key part of legal work . long processing times are due to a significant backlog and to the amount of work required from counsels .
Approach: They propose to extend existing neural named-entity recognition models to retrieve 19 categories of items from refugee cases.
Outcome: The proposed pipeline achieves a superior F1- score on five of the targeted categories and superior to 80% on an additional 4 categories.
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are fast but require expensive pre-training . a new approach to scale large language models into MoEs reduces inference costs .
Approach: They propose an analytical post-training framework that rapidly restructures FFNs into sparse MoE architectures using only a small calibration dataset.
Outcome: The proposed framework outperforms existing methods on a small calibration dataset.
A Bayesian Framework for Information-Theoretic Probing (2021.emnlp-main)

Copied to clipboard

Challenge: a recent paper suggests that probing should be seen as approximating a mutual information.
Approach: They propose a Bayesian mutual information framework that probes probing representations from the perspective of Bayes' agents.
Outcome: The proposed framework allows for more intuitive results in scenarios with finite data.
Benchmarking Data Science Agents (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have emerged as promising data science aids, assisting humans in data analysis and processing.
Approach: They propose an evaluation paradigm and benchmarks that assess the performance of data science agents throughout the entire data science lifecycle.
Outcome: The proposed evaluation paradigm streamlines dataset preparation, improves coverage, and expands benchmarking comprehensiveness.
A Finite-State Morphological Analyser for Evenki (2020.lrec-1)

Copied to clipboard

Challenge: Evenki is a language with rich morphology, therefore a morphological analyser is highly desirable for processing Evenki texts.
Approach: They propose to use a morphological analyser for Evenki to analyze half of the corpus . they evaluate the morphology of available corpora and estimate accuracy, recall and F-score .
Outcome: The proposed morphological analyser can analyse less than a half of the available corpora on Evenki . it is based on the Helsinki Finite-State Transducer toolkit (HFST).
KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question Answering (2022.acl-long)

Copied to clipboard

Challenge: Open-Domain Question Answering (ODQA) models typically include a retrieving module and a reading module.
Approach: They propose a new open-domain question-answering framework that uses a knowledge-enhanced version of FiD to improve the approach.
Outcome: The proposed model improves on ODQA benchmark datasets with less than 40% computation cost.
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data.
Approach: They propose a framework for table structure parsing based on the image-to-text model and a vision guider to refine the model’s capability to understand textual semantics in table images.
Outcome: The proposed framework improves on a dataset of PubTabNet, PubTables1M, WTW, and iFLYTAB and will be made publicly available.
A Myanmar (Burmese)-English Named Entity Transliteration Dictionary (2020.lrec-1)

Copied to clipboard

Challenge: Currently, there are no data available for the transcription of borrowed English words in Myanmar . lack of resources is a problem for many understudied languages .
Approach: They construct a dictionary of Myanmar-English transliteration instances using a CC BY-NC-SA license.
Outcome: The proposed model outperforms the statistical model significantly on the character level.
Too Big to Fail: Larger Language Models are Disproportionately Resilient to Induction of Dementia-Related Linguistic Anomalies (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that the attention mechanism in transformer-based NLMs may present an analogue to the notions of cognitive and brain reserve.
Approach: They propose a bidirectional ablation method that masks attention heads to display degradation of similar magnitude to masking in smaller models.
Outcome: The proposed method exhibits properties attributed to the concepts of cognitive and brain reserve in human brain studies.
MuLD: The Multitask Long Document Benchmark (2022.lrec-1)

Copied to clipboard

Challenge: Existing benchmarks for NLP focus on tasks for one or two sentences, but efficient techniques are needed for processing much longer sequences.
Approach: They propose to modify existing NLP tasks to create a long document benchmark which requires models to successfully model long-term dependencies in the text.
Outcome: The proposed benchmark is much more challenging than its ‘short document’ equivalents.
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas (2025.coling-main)

Copied to clipboard

Challenge: Existing studies on LMs have focused on linguistic generalizations and representations from developmentally plausible data.
Approach: They propose to use phoneme- and grapheme-based language models to learn linguistic units at and below the word level.
Outcome: The proposed models can achieve strong performance on syntactic and novel benchmarks and match grapheme-based models in standard tasks and novel evaluations.
VLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training (2025.findings-emnlp)

Copied to clipboard

Challenge: a significant drawback of Vision-language Models is their reliance on static training data, leading to outdated information and limited contextual awareness.
Approach: They propose a framework with knowledge-enhanced reranking and noise-injected training to improve the VLM's ranking ability.
Outcome: The proposed framework is based on a simple yet effective instruction template and is able to induce its ranking ability and serve it as a reranker to precisely filter the top-k retrieved images.
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated significant capabilities in processing and understanding text data.
Approach: They propose a structure-based instruction-based method to enhance LLM performance on complex graph tasks.
Outcome: The proposed framework outperforms open-source models on graph problem-solving, but the gap is narrowing.
Corpus for Automatic Structuring of Legal Documents (2022.lrec-1)

Copied to clipboard

Challenge: In populous countries, pending legal cases are growing exponentially.
Approach: They propose a corpus of legal judgment documents in English that is annotated with a label coming from a list of pre-defined rhetorical roles.
Outcome: The proposed corpus of legal judgment documents is annotated with a label coming from a list of pre-defined rhetorical roles.
Joint Speech Transcription and Translation: Pseudo-Labeling with Out-of-Distribution Data (2023.findings-acl)

Copied to clipboard

Challenge: a recent study shows that self-training can improve upon fully supervised baselines in low-resource settings for several sequence-to-sequence tasks.
Approach: They propose to use pseudo-labeling to label unsupervised data and add it to the training pool.
Outcome: The proposed setup improves on the unsupervised data by using pseudo-labeling . the proposed setup provides 0.4% absolute WER and 2.1 BLEU points for En–De .
Up-cycling Data for Natural Language Generation (L18-1)

Copied to clipboard

Challenge: Existing systems for creating adaptive texts from cultural heritage data require expert input . a number of research projects have focused on using NLG systems to create multilingual adaptive texts .
Approach: They propose automatic processes which aim to reduce the need for expert input . they normalize the dates and names which occur in the data and link to the Semantic Web .
Outcome: The proposed processes reduce the need for expert input during conversion and up-cycling process.
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have expanded their capabilities to multimodal contexts, including comprehensive video understanding.
Approach: They propose to store and retrieve relevant video frames for specific queries and a Divide-and-Conquer loop capable of autonomous reasoning.
Outcome: The proposed model efficiently stores and retrieves relevant video frames for specific queries, preserving the detailed content of videos.
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for enhancing performance through increased use of expert knowledge often result in diminishing sparsity during expert selection.
Approach: They propose a framework that integrates the computational processes of MoE with the concept of knowledge transferring in multi-task learning.
Outcome: The proposed framework outperforms existing methods under identical conditions concerning the number of experts.
LLM-as-Scheduler: Agentic Workflow Dynamic Scheduling (2026.acl-long)

Copied to clipboard

Challenge: Experiments show that LAS cuts token usage by 43% and reduces end-to-end latency by more than 36%, while causing at most a 1.4 percentage-point drop in accuracy compared with a strong fixed workflow.
Approach: They propose a system that dynamically chooses the right workflow for each query.
Outcome: Experiments show that LAS cuts token usage by 43% and reduces end-to-end latency by more than 36% while causing at most a 1.4 percentage-point drop in accuracy compared with a strong fixed workflow.
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly permeating daily lives and require real-time interactions that mirror human conversations.
Approach: They propose to use time-division-multiplexing to process queries and responses pseudo-simultaneously.
Outcome: The proposed model can listen to users while generating output and adjust to provide instant feedback.
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities in handling long context inputs, but this comes at the cost of increased computational resources and latency.
Approach: They propose an algorithm that uses early LLM layers as filters to select and compress input tokens, reducing the context length for subsequent processing.
Outcome: The proposed method outperforms existing techniques on the Needle in a Haystack task while demonstrating comparable performance on the LongBench challenge.
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limitations when it comes to comprehending and expressing world knowledge that extends beyond the boundaries of natural language.
Approach: They propose a model that integrates symbolic data into LLM training without loss of generality ability.
Outcome: The proposed model performs better on symbol- and NL-centric tasks.
SandhiKosh: A Benchmark Corpus for Evaluating Sanskrit Sandhi Tools (L18-1)

Copied to clipboard

Challenge: Several important texts which are of interest to people all over the world were written in Sanskrit.
Approach: They develop a Sanskrit benchmark to evaluate the completeness and accuracy of tools . they use three most prominent tools to evaluate their completeness .
Outcome: The proposed tools have substantial scope for improvement and are available to researchers worldwide.
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags (2024.findings-acl)

Copied to clipboard

Challenge: XL-HeadTags is a dataset that includes 20 languages across 6 diverse language families.
Approach: They propose to leverage auxiliary information such as images and captions embedded in news articles to retrieve relevant sentences and utilize instruction tuning with variations to generate both headlines and tags for news articles in a multilingual context.
Outcome: The proposed approach generates headlines and tags in a multilingual context using images and captions embedded in the articles and instruction tuning with variations.
PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for Indian languages are limited in terms of coverage and size.
Approach: They propose a multilingual and massively parallel summarization corpus focused on languages in India that provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
Outcome: The proposed dataset provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to meeting summarization are limited due to noise, lengthy transcripts, and scattered salient information.
Approach: They propose a two-step framework for meeting summarization that leverages a self-supervised paradigm to reconstruct transcripts and a relative positional bucketing algorithm to equip models to generate the summary.
Outcome: The proposed method significantly reduces memory consumption and processing time on two meeting summarization datasets.
Optimizing Multi-Hop Document Retrieval Through Intermediate Representations (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to addressing multi-hop queries are computationally expensive . despite their success, large language models often generate factually incorrect answers .
Approach: They propose a layer-by-layer reasoning approach that leverages intermediate representations from the middle layers to retrieve external knowledge.
Outcome: The proposed method outperforms existing RAG methods on open-domain multi-hop question-answering datasets while maintaining inference overhead similar to that of standard RAG.
Is a cute puyfred cute? Context-dependent form-meaning systematicity in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: valence is encoded in meaningful ways in large language models and in some LLMs, pseudowords affect the representation of whole sentences similarly to words.
Approach: They investigate how LLMs represent valence, a key semantic attribute, and how they deal with contextualisation of pseudowords in sentences.
Outcome: The results show that the models represent valence, a key semantic attribute, in sentences and in context, and that they handle the contextualisation of pseudowords differently.
Exploring Intra and Inter-language Consistency in Embeddings with ICA (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that ICA can reveal universal semantic axes across languages but lack verification of consistency of independent components within and across languages.
Approach: They propose to use independent component analysis to identify independent components that are more interpretable than PCA to find universal semantic axes.
Outcome: The proposed framework ensures the reliability and universality of semantic axes.
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have shown promising results in arithmetic and symbolic reasoning by expressing intermediate reasoning in text as a chain of thought, yet struggle to extend this capability to answer text queries that are easily solved by visual reasoning.
Approach: They propose a method to unlock the visual reasoning capabilities of multimodal large language models by using a metaphorical ‘whiteboard’ to draw out reasoning steps as images and return these images back to the model for further processing.
Outcome: The proposed method shows that it can be used on four difficult tasks that involve visual and spatial reasoning with no demonstrations or specialized modules.
CAPA: Contribution-Aware Pruning and FFN Approximation for Efficient Large Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Efficient inference in Large Vision Language Models is constrained by the high cost of processing thousands of visual tokens.
Approach: They propose a framework that prunes visual tokens using attention contribution at critical functional transitions and reduces computations using efficient linear approximations.
Outcome: The proposed framework achieves competent efficiency–performance trade-offs with improved robustness.
Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for historical document restoration focus on single modality or limited-size restoration, failing to meet practical needs.
Approach: They propose a full-page HDR dataset and an automated HDR solution to replace manual restoration methods.
Outcome: The proposed solution improves OCR accuracy from 46.83% to 84.05% when processing severely damaged documents, with enhancement to 94.25% through human-machine collaboration.
LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder (2025.emnlp-main)

Copied to clipboard

Challenge: Prior research on linguistic mechanisms of large language models is limited by coarse granularity, limited analysis scale, and narrow focus.
Approach: They propose a framework for analyzing the linguistic mechanisms of large language models based on Sparse Auto-Encoders.
Outcome: The proposed framework extracts Chinese and English linguistic features across four dimensions . it uncovers intrinsic representations of linguistic knowledge in LLMs and can control outputs .
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks (2025.acl-long)

Copied to clipboard

Challenge: a growing number of recorded human speech is recorded for automated processing, resulting in errors in the transcripts . a configurable framework is proposed to analyze transcript noise impact across noise levels and transcript-cleaning techniques.
Approach: They propose a configurable framework for assessing task models in diverse noisy settings . framework facilitates investigation of task model behavior, which can support effective SLU solutions.
Outcome: The proposed framework can analyze model behavior in various noise levels and transcript-cleaning techniques.
Glyph: Scaling Context Windows via Visual-Text Compression (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) traditionally represent text as sequences of discrete tokens . a long-context scaling problem requires processing more tokens more efficiently .
Approach: They propose a framework that renders long texts into compact visual pages and processes them with a vision-language model.
Outcome: The proposed framework renders long texts into compact visual pages and processes them with a vision-language model.
Dual Alignment Between Language Model Layers and Human Sentence Processing (2026.acl-long)

Copied to clipboard

Challenge: Existing studies have demonstrated both the successes and limitations of accurate predictability estimation by modern LMs in cognitive modeling.
Approach: They propose to use internal layers to better estimate human cognitive effort observed in syntactic ambiguity processing in English.
Outcome: The proposed models can be modeled using surprisal from early layers of large language models (LLMs) this raises the question whether such advantages extend to more syntactically challenging constructions, where surprised estimates underestimate human cognitive effort.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations