Papers by Cristiano Ciaccio

3 papers
Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation.
Approach: They evaluate the resilience of state-of-the-art MGT detectors to linguistically informed adversarial attacks by using Direct Preference Optimization to shift the MGT style toward human-written text.
Outcome: The proposed pipeline fine-tunes language models to shift the MGT style toward human-written text (HWT) it obtains generations more challenging to detect by current models, and shows that detectors can be easily fooled with relatively few examples, resulting in a significant drop in detecting performances.
Beyond the Spelling Miracle: Investigating Substring Awareness in Character-Blind Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Current Pre-trained Language Models are character-blind and struggle in spelling tasks . ability to identify characters and substrings within words is trivial but fundamental to robust language understanding.
Approach: They propose to evaluate pre-trained language models with a binary substring identification task . they propose to examine where, when, and how a PLMs develop awareness of characters and substrings .
Outcome: The proposed model identifies characters and substrings in a binary substring identification task.
Evaluating Lexical Proficiency in Neural Language Models (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Natural Language Processing have been significantly shaped by the Deep Learning tsunami and the introduction of Transformer-based Language Models.
Approach: They validated a framework to assess the lexical proficiency and linguistic creativity of Transformer-based Language Models (LMs) by analyzing performance of LMs of different sizes across tasks involving the generation, definition, and contextual usage of lexicals, neologisms, and nonce words.
Outcome: The framework evaluates LMs in mono- and multilingual configuration across tasks involving the generation, definition, and contextual usage of lexicalized words, neologisms, and nonce words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations