Papers by David Alfter

7 papers
Automatically Generated Definitions and their utility for Modeling Word Meaning (2024.emnlp-main)

Copied to clipboard

Challenge: Modern language models generate semantic representations for words based on context and context based models.
Approach: They propose to use dictionary-like sense definitions to generate sentence embeddings . they evaluate the quality of the generated definitions on existing English benchmarks based on the results of their study .
Outcome: The proposed model sets new state-of-the-art results on lexical semantics tasks compared to baselines .
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) aims to automatically assess the quality of essays.
Approach: They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French.
Outcome: The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam.
More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Word Usage Graphs (WUGs) represent word sense clusters from simple pairwise word use judgments.
Approach: They propose to use a weighted graph to represent human semantic proximity judgments for pairs of word uses to infer word sense clusters from simple pairwise word use judgments.
Outcome: The proposed approach can be applied in a Word Sense Induction (WSI) setting or for Word sense disambiguation (WSD) it is the first and to date largest manually annotated, diachronic WUG dataset.
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)

Copied to clipboard

Challenge: a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence .
Approach: They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora .
Outcome: The proposed toolkit improves performance over standard feature-based readability prediction.
Using Multilingual Resources to Evaluate CEFRLex for Learner Applications (2020.lrec-1)

Copied to clipboard

Challenge: The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities.
Approach: They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels.
Outcome: The results show that the English CEFRLex resource is in accordance with external resources that are gold standard.
Linguistic Corpus Annotation for Automatic Text Simplification Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Evaluating automatic text simplification systems is a difficult task that is performed either by automatic metrics or user-based evaluations.
Approach: They propose to use annotations of the ASSET corpus to analyze SARI’s behavior and to re-evaluate existing ATS systems.
Outcome: The proposed methods can be used to analyze SARI’s behavior and to re-evaluate existing ATS systems.
Is Attention Explanation? An Introduction to the Debate (2022.acl-long)

Copied to clipboard

Challenge: Attention has been used in various tasks of NLP and other fields of machine learning to increase performance and provide some explanations.
Approach: They propose to use attention as an explanation for deep learning models to increase performance . they propose to apply attention weights to queries and queries based on scalar scores .
Outcome: The proposed model can be used to increase performance while providing some explanations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations