Papers by João Silva

8 papers
Universal Grammatical Dependencies for Portuguese with CINTIL Data, LX Processing and CLARIN support (2022.lrec-1)

Copied to clipboard

Challenge: a new collection of quality language resources is presented for the computational processing of the Portuguese language . the framework for the mapping between linguistic form and meaning is centered on the notion of grammatical relation .
Approach: They propose a new set of quality language resources for the computational processing of the Portuguese language under the Universal Dependencies framework.
Outcome: The proposed framework provides for the mapping between linguistic form and meaning representations.
A Shared Task of a New, Collaborative Type to Foster Reproducibility: A First Exercise in the Area of Language Science and Technology with REPROLANG2020 (2020.lrec-1)

Copied to clipboard

Challenge: Scientific knowledge is grounded on falsifiable predictions and therefore its credibility and raison d'être rely on the possibility of repeating experiments and getting similar results as originally obtained and reported.
Approach: They propose a collaborative task which is collaborative rather than competitive and supports reproduction of research results.
Outcome: The proposed task is called REPROLANG-The Shared Task on the Reproduction of Research Results in Science and Technology of Natural Language Processing (LREC2020).
Semantic Equivalence Detection: Are Interrogatives Harder than Declaratives? (L18-1)

Copied to clipboard

Challenge: Semantic Text Similarity (STS) tasks are often not seen as similar to semantic equivalence detection tasks.
Approach: They propose to assess the performance of different approaches to STS over different types of textual segments.
Outcome: The proposed methods differ in performance over different types of textual segments, including declaratives and interrogatives, under conditions of comparability.
Browsing and Supporting Pluricentric Global Wordnet, or just your Wordnet of Interest (L18-1)

Copied to clipboard

Challenge: a wordnet browser that allows to consult wordnet content is presented in this paper . the paper presents a browser that meets design requirements and complies with the most ample range of design features.
Approach: They propose a wordnet browser that meets design requirements for wordnets . they use existing browsers to analyze their functionalities and build a new browser .
Outcome: The proposed browser meets design requirements and complies with the most ample range of design features.
Reproduction and Revival of the Argument Reasoning Comprehension Task (2020.lrec-1)

Copied to clipboard

Challenge: Reproduction of scientific results is essential for scientific development across all disciplines.
Approach: They evaluate scientific reproduction of arguments reasoning comprehension systems . they find reproducing results of previous work is a basic requirement for validating hypothesis .
Outcome: The proposed systems were compared with the revised data set and scored in line with the results of the argument reasoning comprehension task.
Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense Reasoning (2021.emnlp-main)

Copied to clipboard

Challenge: a recent study has found that commonsense reasoning models are learning transferable generalizations . commonsensibility is a human capacity that has been a core challenge to Artificial Intelligence since its inception.
Approach: They conduct an analysis of benchmarks that involve commonsense reasoning . they find that most datasets experimented with are problematic . commonsensence is a quintessential human capacity .
Outcome: The proposed model is able to perform well on commonsense reasoning tasks . the model is not learning transferable generalizations or taking advantage of shortcuts .
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities .
Approach: They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese.
Outcome: The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities .
The MWN.PT WordNet for Portuguese: Projection, Validation, Cross-lingual Alignment and Distribution (2020.lrec-1)

Copied to clipboard

Challenge: Lexical semantic networks are pervasive in natural language processing . Lexical ontologies play a key role in virtually all major applications .
Approach: The present paper presents the MWN.PT WordNet for Portuguese . it is the largest high quality, manually validated and cross-lingually integrated wordnet of Portuguese based on the Princeton WordNet of English .
Outcome: The MWN.PT WordNet for Portuguese includes 41,000 concepts expressed by 38,000 lexical units.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations