Challenge: Existing solutions for document QA fail to provide personalized and up-to-date information efficiently.
Approach: They propose to deploy a self-evolving, efficient LLM system that can offer personalized research services, maintaining a real-time updated database.
Outcome: The proposed system saves 69.92% of time after efficient deployment.

Similar Papers

arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation (2026.acl-long)

Copied to clipboard

Challenge: Literature review tables are essential for summarizing and comparing collections of scientific papers.
Approach: They propose to generate a database of literature review tables from a pool of papers and to model retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators.
Outcome: The proposed method improves over strong baselines while the absolute scores remain modest, underscoring the task’s difficulty.
LLM-Independent Adaptive RAG: Let the Question Speak for Itself (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to retrieve Large Language Models (LLMs) are inefficient and impractical.
Approach: They propose a lightweight adaptive retrieval method that leverages external information to achieve comparable quality while achieving significant efficiency gains.
Outcome: The proposed methods achieve comparable quality while achieving significant efficiency gains on 6 QA datasets.
On Improving Repository-Level Code QA for Large Language Models (2024.acl-srw)

Copied to clipboard

Challenge: Commercial AI-assisted programming Chatbots may generate incorrect information when requests go beyond the model training data or require additional knowledge.
Approach: They propose to implement different self-alignment processes and retrieval-augmented generation pipelines to improve the copilot performance.
Outcome: The proposed model improves the copilot performance on repository-level semantics, dependency between files, and meta-information about the repository.
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs (2025.acl-long)

Copied to clipboard

Challenge: Recent surveys of literature highlight the overwhelming growth of Large Language Models (LLMs).
Approach: They propose a semi-automated literature analysis approach that automates literature analysis using LLMs.
Outcome: The proposed approach reduces paper surveying and data extraction by 93% compared to manual methods.
LLM-Evolve: Evaluation for LLM’s Evolving Capability on Benchmarks (2024.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models evaluate LLMs on i.i.d. tasks, overlooking their ability to learn iteratively from past experiences.
Approach: They propose a framework which extends established benchmarks to sequential problem-solving settings and provides feedback after each round to build a demonstration memory that the models can query in future tasks.
Outcome: The proposed framework can improve performance of LLMs by learning from past interactions and improve models' performance over time.
FalconCopilot: Empowering LLMs Towards Integrated Human-Machine Systems for Aviation Autonomy (2026.findings-acl)

Copied to clipboard

Challenge: Complex flight tasks require both intricate, long-horizon decision-making and precise operations.
Approach: They propose a LLM-based copilot system that addresses deficiencies in adaptability and fine-grained decision support while integrating with a high-fidelity environment.
Outcome: The proposed system shortens task completion time while attaining a level of performance approaching that of a human instructor.
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing memory systems can support long-horizon human-LLM interactions by persisting historical interactions beyond limited context windows.
Approach: They propose a framework that augments memory systems with a self-evolving meta-memory . meta-meso is iteratively distilling transferable knowledge utilization experiences . results show MetaMem outperforms strong baselines by over 3.6% .
Outcome: The proposed framework outperforms baselines by over 3.6% in the long-horizon human-LLM interaction.
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to improve data quality face limitations in static dataset curation that fail to adapt to evolving model capabilities.
Approach: They propose a self-evolving framework that uses model-aware data selection and context-preserving data refinement to improve LLM performance.
Outcome: The proposed framework improves the quality of seed data and boosts LLM’s performance with improving accuracy by 7.15% on average while maintaining the original dataset scale.
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)

Copied to clipboard

Challenge: Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results.
Approach: They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories.
Outcome: The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories.
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have demonstrated remarkable performance across tasks.
Approach: They propose a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models.
Outcome: The proposed framework extends existing benchmarks to extend models across tasks and tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations