Papers by Philippe Martin
Paraphrase Generation Evaluation Powered by an LLM: A Semantic Metric, Not a Lexical One (2025.coling-main)
Copied to clipboard
| Challenge: | Existing measures for automatic paraphrase generation are based on lexical distances or semantic embedding alignments. |
| Approach: | They propose a measure based on a log likelihood ratio from an LLM to assess the quality of a potential paraphrase. |
| Outcome: | The proposed measure is better for sorting pairs of sentences by semantic proximity and provides an interpretable classification threshold between paraphrases and non-paraphrases. |
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)
Copied to clipboard
Delphine Bernhard, Anne-Laure Ligozat, Fanny Martin, Myriam Bras, Pierre Magistry, Marianne Vergez-Couret, Lucie Steiblé, Pascale Erhart, Nabil Hathout, Dominique Huck, Christophe Rey, Philippe Reynés, Sophie Rosset, Jean Sibille, Thomas Lavergne
| Challenge: | RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard. |
| Approach: | They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard. |
| Outcome: | The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages. |
A German Corpus for Fine-Grained Named Entity Recognition and Relation Extraction of Traffic and Industry Events (L18-1)
Copied to clipboard
Martin Schiersch, Veselina Mironova, Maximilian Schmitt, Philippe Thomas, Aleksandra Gabryszak, Leonhard Hennig
| Challenge: | Using text streams to extract events pertaining to specific companies, routes and routes remains a challenge. |
| Approach: | They describe a corpus of German-language documents annotated with fine-grained geo-entities and standard named entity types. |
| Outcome: | The proposed corpus consists of newswire texts, twitter messages, and traffic reports from radio stations, police and railway companies. |