Papers by Pierre Zweigenbaum

16 papers
GNEG: Graph-Based Negative Sampling for word2vec (P18-2)

Copied to clipboard

Challenge: Generally speaking, negative sampling is the best choice for distributed word representation learning.
Approach: They hypothesize that taking into account global, corpus-level information and generating a different noise distribution for each target word better satisfies the requirements of negative examples for each training word.
Outcome: The proposed approach boosts the word analogy task by about 5% and improves the performance on word similarity tasks by about 11% compared to the baseline.
Building Comparable Corpora for Assessing Multi-Word Term Alignment (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to extract bilingual terminologies from corpora are limited . MWTs pose serious challenges for alignment and machine translation systems .
Approach: They propose an approach to build comparable corpora and bilingual term dictionaries that evaluate bilingual term alignment in comparable corpus.
Outcome: The proposed method is validated on an existing dataset and manually annotated data.
On the Rejection Criterion for Proxy-based Test-time Alignment (2026.acl-short)

Copied to clipboard

Challenge: Recent work suggests that test-time alignment methods rely on a small aligned model as a proxy that guides the generation of a larger base model.
Approach: They propose a rejection criterion based on a conservative confidence bet for test-time alignment methods that use a small aligned model as a proxy to guide the generation of a larger base model.
Outcome: The proposed approach outperforms previous work on several datasets.
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
Embedding Strategies for Specialized Domains: Application to Clinical Entity Recognition (P19-2)

Copied to clipboard

Challenge: Off-the-shelf word embeddings tend to perform poorly on texts from specialized domains such as clinical reports.
Approach: They combine off-the-shelf contextual embeddings with static word2vec embedders trained on a small in-domain corpus built from task data to reach and sometimes outperform representations learned from a large corpus in the medical domain.
Outcome: The proposed embedding strategies outperform representations learned from a large corpus in the medical domain.
Decorate the Examples: A Simple Method of Prompt Design for Biomedical Relation Extraction (2022.lrec-1)

Copied to clipboard

Challenge: Recent research shows that prompt-based learning improves performance on relation extraction tasks.
Approach: They propose a prompt-based learning method that generates comprehensive prompts for biomedical relation extraction using a ChemProt dataset.
Outcome: The proposed method improves fine-tuning on a biomedical relation extraction task with a cloze-test task and fewer training examples to make reasonable predictions.
Re-train or Train from Scratch? Comparing Pre-training Strategies of BERT in the Medical Domain (2022.lrec-1)

Copied to clipboard

Challenge: Recent years have witnessed the widespread use of transfer learning techniques in Natural Language Processing (NLP)
Approach: They train BERT models from scratch using many configurations involving general and medical corpora.
Outcome: The initial corpus only has a weak influence when these are further pre-trained on a medical corpus.
Enriching a Time-Domain Astrophysics Corpus with Named Entity, Coreference and Astrophysical Relationship Annotations (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora for astrophysical natural language processing are limited to Named Entity Recognition tasks, leaving a gap in resource diversity.
Approach: They propose to expand astroECR to cover named entities, coreferences, annotations related to aastrphysical relationships, and normalizing celestial object names.
Outcome: The proposed model extends the time-domain astrophysics corpus to include named entities, coreferences, and annotations related to aastrphysical relationships.
CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters (2020.coling-main)

Copied to clipboard

Challenge: Pre-trained language representations from Transformers have become the most popular choice for building NLP systems.
Approach: They propose a new variant of BERT that drops the wordpiece tokenization system altogether . they propose 'characterBERT' module to represent entire words by consulting their characters .
Outcome: The proposed model improves performance on a variety of medical domain tasks while producing robust, word-level, and open-vocabulary representations.
Three Dimensions of Reproducibility in Natural Language Processing (L18-1)

Copied to clipboard

Challenge: a recent editorial on reproducibility in language processing defined three dimensions of reproducibility . authors had already submitted a correction, but there is no consensus on the definitions .
Approach: They propose an ontology of reproducibility in natural language processing to address these problems . they propose to analyze three dimensions of reproducible in natural languages papers . authors propose to use a 'replicability' term to describe the reproducibility of a conclusion, finding, value .
Outcome: The proposed ontology aims to enhance future research and communication about the topic and retrospective meta-analyses.
KAD: A Framework for Proxy-based Test-time Alignment with Knapsack Approximation Deferral (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have shown that large language models require further alignment to adhere to downstream task requirements and stylistic preferences.
Approach: They propose a proxy-based test-time alignment method to circumvent alignment costs by reducing the token-specific deferral rule to 0-1 knapsack problem.
Outcome: The proposed method improves both task performance and speculative decoding speed.
Handling Entity Normalization with no Annotated Corpus: Weakly Supervised Methods Based on Distributional Representation and Ontological Information (2020.lrec-1)

Copied to clipboard

Challenge: Entity normalization is an important subtask of information extraction . it links entities mentions in text to categories or concepts in a reference vocabulary .
Approach: They propose a method that uses corpus selection, pre-processing and weak supervision strategies to address the scarcity of training data.
Outcome: The proposed method outperforms state-of-the-art methods in terms of accuracy and parametrization . it uses corpus selection, pre-processing and weak supervision strategies .
Automating Document Discovery in the Systematic Review Process: How to Use Chaff to Extract Wheat (L18-1)

Copied to clipboard

Challenge: Systematic reviews address research questions by comprehensively examining the entire published literature.
Approach: They compare the impact of different schemes for choosing positive and negative examples from the different screening stages on the training of automated systems.
Outcome: The proposed ranking system achieves an AUC of 0.803 and 0.768 when relying on gold standard decisions based on title and abstracts of articles, and an AUT of 0.625 and 0.839 when based upon gold standard decision based in full text.
Cross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient’s Perspective (2022.lrec-1)

Copied to clipboard

Challenge: a recent study shows that the class labels of german documents containing ADRs are imbalanced . clinical trials and physicians prescribing medications cannot cover every potential use case.
Approach: They propose to use binary annotated documents from a german patient forum to detect ADRs.
Outcome: The proposed model achieves an F1 score of 37.52 for the positive class on the German patient forum.
Combining rule-based and embedding-based approaches to normalize textual entities with an ontology (L18-1)

Copied to clipboard

Challenge: a method to normalize multi-word terms with concepts from a domain-specific ontology is proposed . a large part of knowledge is expressed in textual form, such as in scientific articles .
Approach: They propose a method to normalize multi-word terms with concepts from a domain-specific ontology.
Outcome: The proposed method outperforms existing methods on a categorization task in bacterial habitats . the results are encouraging, and the proposed method is expected to be widely used in the biomedical/biological field .
A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing clinical corpora mostly revolves around scientific articles in English . existing literature is limited to only a few scientific articles .
Approach: They propose to use user-generated data sources to uncover adverse drug reactions . existing clinical corpora mostly revolves around scientific articles in english . authors provide statistics to highlight certain challenges associated with the corpus .
Outcome: The proposed corpus includes 12 entity types, four attribute types, and 13 relation types . it provides strong baselines for extracting entities and relations between entities .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations