Papers by Camille Pradel

4 papers
Mining Discourse Markers for Unsupervised Sentence Representation Learning (N19-1)

Copied to clipboard

Challenge: Current state of the art systems in NLP heavily rely on manually annotated datasets, which are expensive to obtain and are ineffective to extract.
Approach: They propose to automatically discover sentence pairs with relevant discourse markers and apply it to massive amounts of data.
Outcome: The proposed method can learn transferable sentence embeddings from 174 discourse markers even for rare markers such as “coincidentally” or “amazingly”.
DiscSense: Automated Semantic Analysis of Discourse Markers (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for predicting discourse markers have been used to study link between markers and semantic relations .
Approach: They use a model trained to predict discourse markers between sentence pairs to predict plausible markers between sentences with a known semantic relation.
Outcome: The proposed method predicts markers between sentence pairs with a known semantic relation . the resulting dataset, named DiscSense, is publicly available .
A Pragmatics-Centered Evaluation Framework for Natural Language Understanding (2022.lrec-1)

Copied to clipboard

Challenge: a number of studies have suggested that models induce universal text representations . current benchmarks focus on semantic phenomena, so pragmatics needs to be the focus .
Approach: They propose a benchmark that unites 11 pragmatics-focused evaluation datasets for English.
Outcome: The proposed benchmark shows that natural language inference does not result in genuinely universal representations.
How’s Business Going Worldwide ? A Multilingual Annotated Corpus for Business Relation Extraction (2022.lrec-1)

Copied to clipboard

Challenge: The 21st century economy has shaped the economic landscape and changed the way market stakeholders interact with each other in the global market where national borders have melted and trades became more open and free.
Approach: They propose a multilingual dataset for automatic extraction of binary business relations involving organizations from the web.
Outcome: The proposed dataset is the first multilingual dataset for automatic extraction of binary business relations involving organizations from the web.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations