Challenge: Recent advances in generative AI technologies like large language models have boosted the incorporation of AI assistance in writing workflows.
Approach: They conduct an experimental study to determine whether disclosure of AI assistance in the writing process would affect people's evaluation on the quality of the writing and ranking of different writings.
Outcome: The disclosure of AI assistance decreases the average quality ratings for argumentative essays and creative stories, and increases the quality of the writings.

Similar Papers

Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated (2025.findings-acl)

Copied to clipboard

Challenge: Prior research on AI mistrust focused primarily on AI's bias towards different human pop-ups.
Approach: They examine how bias shapes the perception of AI versus human generated content . they found that raters favored content labeled "Human Generated" even when labels were deliberately swapped .
Outcome: The findings highlight the limitations of human judgment in interacting with AI and offer a foundation for improving human-AI collaboration.
An Exploration of Post-Editing Effectiveness in Text Summarization (2022.naacl-main)

Copied to clipboard

Challenge: Automated summarization methods are efficient but can suffer from low quality.
Approach: They conducted an experiment with 72 participants to compare post-editing provided summaries with manual summarization for summary quality, human efficiency, and user experience.
Outcome: The results show that post-editing improves summary quality, human efficiency, and user experience on formal (XSum news) and informal (Reddit posts) text.
Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have shattered the ceiling of human-like text generation.
Approach: They compared human-AI interaction types in LLM-assisted news headline generation to determine whether humans can best leverage them for writing.
Outcome: The guiding and selecting model outputs added the most benefit with the lowest cost (in time and effort) Furthermore, AI assistance did not harm participants’ perception of control compared to freeform editing.
Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing Assistance (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly utilized in educational tasks such as providing writing suggestions to students.
Approach: They conduct a large-scale user study with 231 students writing business case peer reviews in german.
Outcome: The proposed model does not carry bias in the feedback loops of the students .
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies comparing AI-generated and human-authored literary texts have produced conflicting results.
Approach: They hypothesize that differences in reading quality can be explained by genuine differences in how readers interpret and value literature .
Outcome: The authors show that the differences in reading quality are largely explained by differences in how readers interpret and value literature, rather than by an intrinsic quality of the texts evaluated.
Automatic Authorship Analysis in Human-AI Collaborative Writing (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for authorship analysis and text detection are limited . authors: human-AI collaborative writing poses a potential challenge for existing methods .
Approach: They investigate the extent to which existing AI detection and authorship analysis models can perform classification on data generated in human-AI collaborative writing sessions.
Outcome: The proposed models outperform existing models on human-AI collaborative writing data . authors say human- AI co-written text will require adapting models in the near future .
Help Me Write a Story: Evaluating LLMs’ Ability to Generate Writing Feedback (2025.acl-long)

Copied to clipboard

Challenge: Current models provide specific and mostly accurate writing feedback, but they fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback.
Approach: They propose a task that corrupts 1,300 stories to intentionally introduce writing issues to study model performance.
Outcome: The proposed model performs well in a controlled task with human and automatic evaluation metrics.
AI-Assisted Human Evaluation of Machine Translation (2025.naacl-long)

Copied to clipboard

Challenge: Annotation metrics are misaligned with the ideal measure of text quality and human evaluation remains the most accurate, reliable, and ultimate standard.
Approach: They propose an annotation protocol that helps annotators mark erroneous parts of the translation and assign a final score.
Outcome: The proposed protocol reduces the time per span annotation by half . the method reduces annotation budget by 25% with filtering of examples that the AI deems to be likely to be correct.
Position Paper: How Should We Responsibly Adopt LLMs in the Peer Review Process? (2026.findings-eacl)

Copied to clipboard

Challenge: a recent paper criticizes the current use of Large Language Models (LLMs) for simple review text generation.
Approach: They propose to use Large Language Models to support key aspects of the review process . they argue that this approach overlooks more meaningful applications of LLMs . authors argue that the increased reviewing burden per reviewer is a factor .
Outcome: The proposed approach would support reproducibility, correctness and relevance of citations and ethics review flagging.
Measuring Human Contribution in AI-Assisted Content Generation (2026.acl-long)

Copied to clipboard

Challenge: generative AI has created a new way to generate content with humans . varying degrees of human contribution in content generation poses significant challenges for the delineation of originality .
Approach: They propose a framework to measure human contribution in AI-assisted content generation by calculating mutual information between human input and AI-aided output relative to self-information of AI-assist output.
Outcome: The proposed measure discriminates between varying degrees of human contribution across multiple creative domains and is validated in real-world applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations