Papers by Thomas Arnold

5 papers
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection (2024.eacl-long)

Copied to clipboard

Challenge: Large language models generate fluent responses to user queries, but they are also susceptible to misuse in journalism, education, and academia.
Approach: They propose a large-scale benchmark for machine-generated text detection that is a multi-generator, multi-domain, and multi-lingual corpus.
Outcome: The proposed system can detect machine-generated text and pinpoint misuse . the proposed system is based on a large-scale benchmark dataset .
M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have brought an unprecedented surge in machine-generated text (MGT) societal implications are posed by their potential misuse and lack of training data.
Approach: They propose a benchmark to detect machine-generated text in multiple languages . they use multi-domain and multi-generator corpus to identify which model generated the text .
Outcome: The proposed benchmark compares a multilingual, multi-domain and multi-generator corpus of MGTs with human-generated content.
Beyond Generic Summarization: A Multi-faceted Hierarchical Summarization Corpus of Large Heterogeneous Data (L18-1)

Copied to clipboard

Challenge: Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader.
Approach: They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically.
Outcome: The proposed method can be used to develop and evaluate hierarchical summarization systems.
Logging Keystrokes in Writing by English Learners (2024.lrec-main)

Copied to clipboard

Challenge: Essay writing is a skill commonly taught and practised in schools.
Approach: They collect and analyse data representing the essay writing process from start to finish by recording every keystroke from multiple writers participating in the study.
Outcome: The data collected from 1,006 writers is compared against a standard dataset of texts, keystroke logs and metadata for public release.
DP-Rewrite: Towards Reproducibility and Transparency in Differentially Private Text Rewriting (2022.coling-1)

Copied to clipboard

Challenge: Existing systems for differentially private text rewriting lack the means to validate privacy-preserving claims.
Approach: They propose an open-source framework for differentially private text rewriting which is modular, extensible and highly customizable.
Outcome: The proposed framework provides a way to lead and validate private text rewriting research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations