Papers by Kuzman Ganchev

4 papers
Dolomites: Domain-Specific Long-Form Methodical Tasks (2025.tacl-1)

Copied to clipboard

Challenge: Experts in various fields perform methodical writing tasks to plan, organize, and report their work.
Approach: They propose a benchmark with specifications for 519 methodical writing tasks . they use expert revisions of up to 10 model-generated examples to evaluate contemporary language models.
Outcome: The proposed benchmark includes specifications for 519 methodical writing tasks . it includes examples with input and output examples, and is available at https://dolomites-benchmark.github.io/ .
State-of-the-art Chinese Word Segmentation with Bi-LSTMs (D18-1)

Copied to clipboard

Challenge: A wide variety of neural-network architectures have been proposed for the task of Chinese word segmentation.
Approach: They propose a bidirectional LSTM model with standard deep learning techniques and best practices for the task of Chinese word segmentation.
Outcome: The proposed model outperforms models based on standard deep learning techniques and best practices on Chinese word segmentation datasets.
Conditional Generation with a Question-Answering Blueprint (2023.tacl-1)

Copied to clipboard

Challenge: Neural generation models often struggle to identify which content units are salient.
Approach: They propose a new conceptualization of text plans as a sequence of question-answer pairs . they propose QA blueprints as QA proxy for content selection and planning .
Outcome: The proposed model improves existing datasets with QA blueprints as proxy for content selection and planning.
Text-Blueprint: An Interactive Platform for Plan-based Conditional Generation (2023.eacl-demo)

Copied to clipboard

Challenge: Recent work shows that conditional generation models can be useful to control the text generation process, leading to irrelevant, repetitive, and hallucinated content.
Approach: They propose a web browser-based demonstration for query-focused summarization that uses a sequence of question-answer pairs as a blueprint plan for guiding text generation.
Outcome: The proposed model can be used to generate query-focused summarization text using question-answer pairs as a blueprint plan.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations