Papers by Laurent Romary

10 papers
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies have demonstrated that Large Language Models (LLMs) are capable of performing grounding acts such as requesting clarification or producing acknowledgments, yet relatively little work has investigated how common ground can be explicitly represented and stored for later use.
Approach: They propose to use relational references to represent common ground in situated dialogues and propose to improve both the establishment of common ground and its subsequent use in the conversation.
Outcome: The proposed models can establish and exploit common ground in situated dialogues and improve its subsequent use.
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names.
Approach: They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations .
Outcome: The French TreeBank is the main source of morphosyntactic and syntactical annotations for French.
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction.
Approach: They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it.
Outcome: The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities.
Modelling Etymology in LMF/TEI: The Grande Dicionário Houaiss da Língua Portuguesa Dictionary as a Use Case (2020.lrec-1)

Copied to clipboard

Challenge: In this article, we will introduce two of the new parts of the Lexical Markup Framework (LMF) ISO standard . part 3 deals with etymological and diachronic data and part 4 consists of a TEI serialisation of all of the prior parts of TEIS model.
Approach: They introduce two parts of the Lexical Markup Framework (LMF) ISO standard, part 3 dealing with etymological and diachronic data and part 4 containing TEI serialisation of all prior parts of a model.
Outcome: The proposed models are based on examples taken from a Portuguese dictionary conversion and are then compared with TEI-XML models.
BERTrade: Using Contextual Embeddings to Parse Old French (2022.lrec-1)

Copied to clipboard

Challenge: a growing interest in digital humanities for automatic processing and annotation of historical texts is generating new models for historical languages.
Approach: They use POS-tagging and dependency parsing to evaluate contextual word embedding models . Old French is one of the historical languages for which they have the largest amount of syntactically annotated data .
Outcome: The proposed model can be used to improve performance in Old French, the authors show . they use POS-tagging and dependency parsing to evaluate the model's quality .
CamemBERT: a Tasty French Language Model (2020.acl-main)

Copied to clipboard

Challenge: Pretrained language models are now ubiquitous in Natural Language Processing, but their use in other languages is limited.
Approach: They propose to train monolingual Transformer-based model for other languages using web crawled data instead of Wikipedia data and a relatively small web crawl dataset leads to better results.
Outcome: The proposed model performs as well as those obtained using larger datasets.
Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding (2024.emnlp-main)

Copied to clipboard

Challenge: despite its importance, there has been limited research on conversational grounding in recent years . pre-trained language models have been costly and time-consuming to evaluate .
Approach: They evaluate the performance of large language models in various aspects of conversational grounding . they propose ways to enhance the capabilities of the models that lag in this aspect .
Outcome: The proposed model performance is based on pre-trained language models and a large pre-training dataset.
On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)

Copied to clipboard

Challenge: TEI and OntoLex deal with corpus citations in lexicons.
Approach: They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding .
Outcome: The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary .
Conversational Grounding: Annotation and Analysis of Grounding Acts and Grounding Units (2024.lrec-main)

Copied to clipboard

Challenge: Successful conversations often rest on common understanding, says a researcher . despite recent advances in dialog systems, there is a noticeable deficit in their grounding capabilities .
Approach: They propose to use a framework to build conversational grounding in dialogs . they propose to analyze two dialog corpora using grounding acts and grounding units .
Outcome: The proposed model shows that language models are not enough to ground dialogs with machines . the proposed model can be used to test the performance of existing Language Models .
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages (2020.acl-main)

Copied to clipboard

Challenge: a recent trend in neural NLP has been the introduction of feature-based and fine-tuning methods . we train monolingual contextualized word embeddings for five mid-resource languages .
Approach: They use common Crawl corpus to train monolingual contextualized word embeddings . they compare performance of OSCAR-based and Wikipedia-based embeddables on part-of-speech tasks .
Outcome: The results show that OSCAR-based and Wikipedia-based embeddings perform better than Wikipedia-style embedders on part-of-speech tagging and parsing tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations