| Challenge: | Contemplata is dedicated to the annotation of constituency trees. |
| Approach: | They propose to use Contemplata to build treebanks and treebank enrichment with relations between syntactic nodes. |
| Outcome: | The proposed solution is dedicated to the annotation of constituency trees and provides a balanced strategy between automatic parsing and manual revision. |
Similar Papers
Active DOP: A constituency treebank annotation tool with online learning (C18-2)
Copied to clipboard
| Challenge: | a new language-independent treebank annotation tool supports rich annotations with discontinuous constituents and function tags. |
| Approach: | They propose a language-independent treebank annotation tool supporting rich annotations with discontinuous constituents and function tags. |
| Outcome: | The proposed tool supports rich annotations with discontinuous constituents and function tags. |
ODIL_Syntax: a Free Spontaneous Spoken French Treebank Annotated with Constituent Trees (2020.lrec-1)
Copied to clipboard
| Challenge: | ODIL Syntax is a French treebank built on spontaneous speech transcripts . the structure of every speech turn is represented by constituent trees . |
| Approach: | They propose a French treebank built on spontaneous speech transcripts with a constituency tree representation. |
| Outcome: | The proposed treebank is based on the French TreeBank, with some annotation guidelines . the proposed tree bank will be freely distributed by January 2020 under a Creative Commons licence . |
Training a Swedish Constituency Parser on Six Incompatible Treebanks (2020.lrec-1)
Copied to clipboard
| Challenge: | Syntactic parsing is a widely used intermediate step in several natural language processing tasks. |
| Approach: | They propose to use a function-tagged constituent treebank for Swedish which includes discontinuous constituents to improve the accuracy. |
| Outcome: | The proposed parser can be trained on additional treebanks that use other annotation models. |
Parsing Headed Constituencies (2024.lrec-main)
Copied to clipboard
| Challenge: | Using constituency and dependency trees, syntactic representations are preferred for tasks such as nominal phrase extraction and identification of terminology. |
| Approach: | They propose a parsing technique that generates headed constituency trees which combine information typically contained in constituency and dependency trees. |
| Outcome: | The proposed method generates headed constituency trees with discontinuities and can generate constituency tree with discontinuity. |
When Collaborative Treebank Curation Meets Graph Grammars (2020.lrec-1)
Copied to clipboard
| Challenge: | Arborator-Grew is a collaborative annotation tool for treebank development. |
| Approach: | They present a collaborative annotation tool for treebank development that combines the features of Arborator and Grew. |
| Outcome: | The proposed tool is a complete redevelopment and modernization of Arborator, replacing its internal database storage by a new Grew API. |
TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations (L18-1)
Copied to clipboard
| Challenge: | TREEANNOTATOR is a browser-based tool for annotating tree-like structures . it provides a wider range of formats and provides graphical annotations . |
| Approach: | They evaluate TREEANNOTATOR, a browser-based tool for annotating tree-like structures, in particular structures that jointly map dependency relations and inclusion hierarchies, as used by Rhetorical Structure Theory. |
| Outcome: | The GUI interface is user-friendly and provides two visualization modes. |
A Gold Standard Dependency Treebank for Turkish (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, Turkish treebanks are limited due to the limited number of annotated sentences in the domains of Wikipedia and ITU Web Treebanks. |
| Approach: | They propose to annotate Turkish web and Wikipedia sentences for segmentation, morphology, part-of-speech and dependency relations using tagsets and a Wikipedia section. |
| Outcome: | The proposed treebank is the largest publicly available morpho-syntactic treebank in terms of word count and has a dedicated Wikipedia section. |
Heads-up! Unsupervised Constituency Parsing via Self-Attention Heads (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing approaches to analyze syntactic knowledge of pre-trained language models have been limited. |
| Approach: | They propose an unsupervised method that extracts constituency trees from PLM attention heads. |
| Outcome: | The proposed method outperforms existing approaches if no development set is present. |
Seshat: a Tool for Managing and Verifying Annotation Campaigns of Audio Data (2020.lrec-1)
Copied to clipboard
Hadrien Titeux, Rachid Riad, Xuan-Nga Cao, Nicolas Hamilakis, Kris Madden, Alejandrina Cristia, Anne-Catherine Bachoud-Lévi, Emmanuel Dupoux
| Challenge: | Seshat is a software for the automated management of annotation campaigns for audio/speech data. |
| Approach: | They propose a system for the automated management of annotation campaigns for audio/speech data which addresses these challenges. |
| Outcome: | The proposed system computes an associated inter-annotator agreement with the gamma measure taking into account the categorisation and segmentation discrepancies. |
A Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish Language (2022.lrec-1)
Copied to clipboard
Büşra Marşan, Oğuz K. Yıldız, Aslı Kuzgun, Neslihan Cesur, Arife B. Yenice, Ezgi Sanıyar, Oğuzhan Kuyrukçu, Bilge N. Arıcan, Olcay Taner Yıldız
| Challenge: | a team of linguists manually annotated a set of constituency trees. |
| Approach: | They propose to create the first Turkish-based dependency-to-constituency conversion algorithm using a morphologic analyser and feature-based machine learning model. |
| Outcome: | The proposed algorithm can be used to generate new constituency treebanks and training data for NLP resources like constituency parsers. |