Papers by Jing Bi
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing (2026.findings-acl)
Copied to clipboard
Hongzhi Zhang, Yuanze Hu, Tinghai Zhang, Jia Fu, Tao Wang, Junwei Jing, Zhaoxin Fan, Wei Bi, Ruiming Tang, Han Li, Guorui Zhou, Kun Gai
| Challenge: | Large Language Models (LLMs) are evolving towards autonomous agents . retrieval capabilities are well-benchmarked, but post-retrieval synthesis is under-evaluated due to open-ended writing. |
| Approach: | They propose a benchmark to evaluate information consolidation capabilities using survey papers as gold standards. |
| Outcome: | The proposed benchmark analyzes the post-retrieval synthesis stage of large language models . it leverages high-quality survey papers as gold standards and reverse-engineers research requests . the proposed benchmark outperforms single-turn generation and reduces hallucinations . |
Enhancing the Reasoning Capabilities of Small Language Models via Solution Guidance Fine-Tuning (2025.coling-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks. |
| Approach: | They propose a new reasoning strategy Solution Guidance (SG) and a plug-and-play training paradigm Solution-Guidance Fine-Tuning (SGFT) which focuses on problem understanding and decomposition at the semantic and logical levels, rather than specific computations. |
| Outcome: | The proposed reasoning strategy Solution Guidance (SG) and plug-and-play training paradigm Solution-Guidance Fine-Tuning (SGFT) improves the reasoning capabilities of small language models on various reasoning tasks. |
OSCaR: Object State Captioning and State Change Representation (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods to extrapolate and comprehend changes in object states are limited . relying on a small set of symbolic words to represent changes has restricted expressiveness of language. |
| Approach: | They propose a dataset and benchmark to evaluate multimodal large language models . they investigate causal relations between a concrete action and the change . |
| Outcome: | The proposed method achieves near parity with GPT-4V ratings across helpfulness, accuracy, reasoning, and other key metrics. |
How to Make Large Language Models Generate 100% Valid Molecules? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) can learn to perform a wide range of tasks, but generating valid molecules using representations like SMILES is challenging in few-shot settings. |
| Approach: | They propose a language framework that converts invalid SMILES to SELFIES and LLMs as post-hoc correctors to ensure that the molecules generated by LLM are 100% valid. |
| Outcome: | The proposed model performs worse with SELFIES than with SMILES and improves on other metrics. |