Papers by Michiharu Yamashita

4 papers
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: MULTITuDE benchmarks lack authentic and machine-generated text in languages other than English . defining characteristic of new generation of LLMs is increased quality of text .
Approach: They propose a benchmarking dataset for multilingual machine-generated text detection that compares detectors with authentic and machine-generated texts in 11 languages.
Outcome: The proposed dataset compares detectors with zero-shot and fine-tuned detectors in 11 languages.
Unmasking Fake Careers: Detecting Machine-Generated Career Trajectories via Multi-layer Heterogeneous Graphs (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate convincing career trajectories in fake resumes . a novel heterogeneous, hierarchical multi-layer graph framework is proposed to model career entities and their relations in a unified global graph built from genuine resumes.
Approach: They propose a novel heterogeneous, hierarchical multi-layer graph framework that models career entities and their relations in a unified global graph built from genuine resumes.
Outcome: The proposed framework outperforms state-of-the-art models by 5.8-85.0% relative to baselines.
Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation (2023.emnlp-main)

Copied to clipboard

Challenge: Recent ubiquity and disruptive impacts of large language models have raised concerns about their potential to be misused.
Approach: They propose a strategy that leverages LLMs' generative and emergent reasoning capabilities to counter human-written and LLM-generated disinformation.
Outcome: The proposed strategy synthesizes authentic and deceptive LLM-generated content through paraphrase-based and perturbation-based prefix-style prompts, respectively.
Authorship Obfuscation in Multilingual Machine-Generated Text Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Language Modeling have birthed Large Language Models (LLMs), which exhibit significant improvements, including the ability to generate texts easily misconstrued as humanwritten.
Approach: They compare authorship obfuscation methods against machine-generated text (MGT) in 11 languages and analyze their performance against 37 well-known AO methods.
Outcome: The proposed methods can cause evasion of detection in all languages, with homoglyph attacks particularly successful.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations