Papers by Kamyar Zeinalipour
Clue-Instruct: Text-Based Clue Generation for Educational Crossword Puzzles (2024.lrec-main)
Copied to clipboard
| Challenge: | Educational crosswords are characterized by less cryptic and more factual clues than traditional puzzles. |
| Approach: | They propose to use a dataset to generate educational clues for Large Language Models (LLMs) they use Wikipedia to gather information associated with relevant keywords and use it to generate clues. |
| Outcome: | The proposed approach generates educational clues from a dataset containing 44,075 examples with text-keyword pairs associated with three distinct crossword clues. |
From Graph to Text and Back: Semantic Fidelity in Automated Industrial Knowledge Graphs (2026.acl-industry)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) often hallucinate entities or omit relations, posing unacceptable liability. |
| Approach: | They propose a self-supervised round-trip pipeline to enforce strict semantic fidelity in KG-to-text generation. |
| Outcome: | The proposed approach improves triple-extraction accuracy and verbalization faithfulness without manual annotation or massive teacher models. |
Elections go bananas: A First Large-scale Multilingual Study of Pluralia Tantum using LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | a large amount of annotated sentences for each feature can be used for in-depth analysis. |
| Approach: | They propose an annotation framework for lexicalization of pluralia tantum . they use an LLM to annotate each instance from the reference corpus . |
| Outcome: | The proposed framework provides useful annotators for semantic, syntactic and sense categories with accuracy ranging from 51% to 89% on a hand-annotated testset. |
PharmaQA.IT: an Italian dataset for Q&A in the pharmaceutical domain (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing medical QA datasets are mostly English and centred on scientific articles or clinical notes. |
| Approach: | They propose an extractive QA dataset built from Riassunti delle Caratteristiche del Prodotto . the final dataset contains 861 high-quality question–answer pairs . |
| Outcome: | The proposed dataset contains 861 high-quality question–answer pairs on indications, contraindications, dosage, warnings, interactions, and pharmacological properties. |