Papers by Giulia Venturi

8 papers
Profiling-UD: a Tool for Linguistic Profiling of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Profiling–UD is a text analysis tool that can be used to characterize language variation from different perspectives.
Approach: They introduce Profiling–UD, a text analysis tool inspired to the principles of linguistic profiling that can support language variation research from different perspectives.
Outcome: The proposed tool is specifically designed to be multilingual since it is based on the Universal Dependencies framework.
Evaluating Large Language Models via Linguistic Profiling (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) undergo extensive evaluation against various benchmarks collected in established leaderboards to assess their performance across multiple tasks.
Approach: They propose a new evaluation methodology to test LLMs' sentence generation abilities under specific linguistic constraints.
Outcome: The proposed evaluation methodology is based on the 'linguistic profiling' approach and is not intended to be a task-oriented evaluation.
On the Nature of BERT: Correlating Fine-Tuning and Linguistic Competence (2022.coling-1)

Copied to clipboard

Challenge: Several studies on the interpretation of Neural Language Models (NLMs) focus on the linguistic generalization abilities of pre-trained models, but little attention is paid to how the linguistic knowledge of the models changes during fine-tuning.
Approach: They propose to examine whether a wide range of linguistic phenomena are forgotten during fine-tuning and whether it is possible to predict the fine- tuned accuracy solely relying on the assessed linguistic competence.
Outcome: The proposed model can predict the evolution of written language competence of native language learners based on the assessed linguistic competence.
“Voices of the Great War”: A Richly Annotated Corpus of Italian Texts on the First World War (2020.lrec-1)

Copied to clipboard

Challenge: “Voices of the Great War” is the first large corpus of Italian historical texts dating back to the period of First World War.
Approach: "Voices of the Great War" is the first large corpus of Italian historical texts dating back to the period of First World War.
Outcome: The "Voices of the Great War" corpus is the first large corpus of Italian historical texts dating back to the period of First World War.
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
Is this Sentence Difficult? Do you Agree? (D18-1)

Copied to clipboard

Challenge: a crowdsourcing-based approach to model sentence complexity is proposed . word-level predictors shown to correlate with greater processing difficulties are e.g. word frequency, age of acquisition, root frequency effect, orthographic neighbourhood frequency .
Approach: They propose a crowdsourcing-based approach to model human perception of sentence complexity using a corpus of sentences rated with judgments of complexity for two typologically-different languages.
Outcome: The proposed model predicts agreement among annotators independently from the assigned judgment and the perception of sentence complexity in Italian and English.
Linguistic Knowledge Can Enhance Encoder-Decoder Models (If You Let It) (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has shown that pre-trained NLMs can capture syntax- and semantic-sensitive phenomena.
Approach: They investigate whether fine-tuning pre-trained models with linguistic knowledge improves their performance in a target task.
Outcome: The proposed enhancements improve models' performance in a target task, the authors show . the study includes models in Italian and English, and multilingual models in English and Italian .
Linguistic Profiling of a Neural Language Model (2020.coling-main)

Copied to clipboard

Challenge: Neural Language Models (NLMs) have become a central component in NLP systems over the last few years, showing outstanding performance and improving the state-of-the-art on many tasks.
Approach: They use a wide set of probing tasks to investigate how a Neural Language Model learns linguistic properties before and after a fine-tuning process.
Outcome: The proposed model can encode a wide range of linguistic characteristics but loses this information when trained on specific downstream tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations