Papers by Sanja Štajner

6 papers
Automatic Assessment of Conceptual Text Complexity Using Knowledge Graphs (C18-1)

Copied to clipboard

Challenge: Existing methods to assess text complexity only at lexical and syntactic levels have not been attempted.
Approach: They propose to automatically estimate conceptual complexity using graph-based measures on a large knowledge base.
Outcome: The proposed measures achieve high discriminative power even in a default setup.
A Detailed Evaluation of Neural Sequence-to-Sequence Models for In-domain and Cross-domain Text Simplification (L18-1)

Copied to clipboard

Challenge: Xu et al., 2016) show that a simple neural architecture can be efficiently used for in-domain and cross-domain text simplification.
Approach: They evaluate neural sequence-to-sequence models for text simplification on Wikipedia and Newsela datasets.
Outcome: The proposed model can generalize across corpora and overcome challenges when tested on Wikipedia and Newsela datasets.
Data-Driven Text Simplification (C18-3)

Copied to clipboard

Challenge: Automatic text simplification is the process of transforming a complex text into an equivalent version which would be easier to read or understand by automatic natural language processors.
Approach: This tutorial provides an overview of automatic text simplification, which is the process of transforming a complex text into an equivalent version.
Outcome: The aim of this paper is to provide a comprehensive overview of past and current research on automatic text simplification.
CATS: A Tool for Customized Alignment of Text Simplification Corpora (L18-1)

Copied to clipboard

Challenge: Existing corpora of original sentences and their manual simplifications are very scarce and small in size, hindering automated text simplification systems.
Approach: They propose a language-independent tool for sentence alignment from parallel/comparable TS resources.
Outcome: The proposed tool performs well on English and Spanish corpora and compares sentences based on their semantic overlap.
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
A Spreading Activation Framework for Tracking Conceptual Complexity of Texts (P19-1)

Copied to clipboard

Challenge: Existing models for assessing conceptual complexity of texts are lacking . conceptual complexity accounts for background knowledge necessary to understand mentioned concepts .
Approach: They propose an unsupervised approach for assessing conceptual complexity of texts based on spreading activation using DBpedia knowledge graph as a proxy to long-term memory.
Outcome: The proposed model outperforms current state of the art in assessing conceptual complexity of texts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations