Papers by Omar Momen
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets (2026.eacl-long)
Copied to clipboard
| Challenge: | Novel metaphor comprehension involves complex semantic processes and linguistic creativity. |
| Approach: | They propose a cloze-style surprisal method that conditions on full-sentence context. |
| Outcome: | The proposed method shows that LM surprisal yields moderate correlations with scores/labels of metaphor novelty. |
Filling the Temporal Void: Recovering Missing Publication Years in the Project Gutenberg Corpus Using LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Currently, there is no publicly available corpus for diachronic text analysis due to the lack of accurate temporal metadata. |
| Approach: | They propose to add missing temporal metadata to the Gutenberg corpus by using open web, Wikipedia, and Open Library API sources. |
| Outcome: | The proposed corpus includes 53,774 books with a total of 3.8 billion tokens in 11 languages, produced between 1600 and 2000. |
DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification (2023.acl-long)
Copied to clipboard
| Challenge: | Current text simplification research mostly focuses on English and on sentencelevel simplification. |
| Approach: | They propose to use a dataset of parallel, professionally written and manually aligned simplifications in plain German "plain DE" and "Einfache Sprache" they build a web harvester and experiment with automatic alignment methods to facilitate integration of non-aligned and to be-published parallel documents. |
| Outcome: | The proposed dataset of parallel, professionally written and manually aligned simplifications in plain German is extended to 750 document pairs and 3.5k sentence pairs. |