Contemplata, a Free Platform for Constituency Treebank Annotation (2020.lrec-1)

Copied to clipboard

Challenge: Contemplata is dedicated to the annotation of constituency trees.
Approach: They propose to use Contemplata to build treebanks and treebank enrichment with relations between syntactic nodes.
Outcome: The proposed solution is dedicated to the annotation of constituency trees and provides a balanced strategy between automatic parsing and manual revision.

Similar Papers

Active DOP: A constituency treebank annotation tool with online learning (C18-2)

Copied to clipboard

Challenge: a new language-independent treebank annotation tool supports rich annotations with discontinuous constituents and function tags.
Approach: They propose a language-independent treebank annotation tool supporting rich annotations with discontinuous constituents and function tags.
Outcome: The proposed tool supports rich annotations with discontinuous constituents and function tags.
ODIL_Syntax: a Free Spontaneous Spoken French Treebank Annotated with Constituent Trees (2020.lrec-1)

Copied to clipboard

Challenge: ODIL Syntax is a French treebank built on spontaneous speech transcripts . the structure of every speech turn is represented by constituent trees .
Approach: They propose a French treebank built on spontaneous speech transcripts with a constituency tree representation.
Outcome: The proposed treebank is based on the French TreeBank, with some annotation guidelines . the proposed tree bank will be freely distributed by January 2020 under a Creative Commons licence .
Training a Swedish Constituency Parser on Six Incompatible Treebanks (2020.lrec-1)

Copied to clipboard

Challenge: Syntactic parsing is a widely used intermediate step in several natural language processing tasks.
Approach: They propose to use a function-tagged constituent treebank for Swedish which includes discontinuous constituents to improve the accuracy.
Outcome: The proposed parser can be trained on additional treebanks that use other annotation models.
Parsing Headed Constituencies (2024.lrec-main)

Copied to clipboard

Challenge: Using constituency and dependency trees, syntactic representations are preferred for tasks such as nominal phrase extraction and identification of terminology.
Approach: They propose a parsing technique that generates headed constituency trees which combine information typically contained in constituency and dependency trees.
Outcome: The proposed method generates headed constituency trees with discontinuities and can generate constituency tree with discontinuity.
When Collaborative Treebank Curation Meets Graph Grammars (2020.lrec-1)

Copied to clipboard

Challenge: Arborator-Grew is a collaborative annotation tool for treebank development.
Approach: They present a collaborative annotation tool for treebank development that combines the features of Arborator and Grew.
Outcome: The proposed tool is a complete redevelopment and modernization of Arborator, replacing its internal database storage by a new Grew API.
TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations (L18-1)

Copied to clipboard

Challenge: TREEANNOTATOR is a browser-based tool for annotating tree-like structures . it provides a wider range of formats and provides graphical annotations .
Approach: They evaluate TREEANNOTATOR, a browser-based tool for annotating tree-like structures, in particular structures that jointly map dependency relations and inclusion hierarchies, as used by Rhetorical Structure Theory.
Outcome: The GUI interface is user-friendly and provides two visualization modes.
A Gold Standard Dependency Treebank for Turkish (2020.lrec-1)

Copied to clipboard

Challenge: Currently, Turkish treebanks are limited due to the limited number of annotated sentences in the domains of Wikipedia and ITU Web Treebanks.
Approach: They propose to annotate Turkish web and Wikipedia sentences for segmentation, morphology, part-of-speech and dependency relations using tagsets and a Wikipedia section.
Outcome: The proposed treebank is the largest publicly available morpho-syntactic treebank in terms of word count and has a dedicated Wikipedia section.
Heads-up! Unsupervised Constituency Parsing via Self-Attention Heads (2020.aacl-main)

Copied to clipboard

Challenge: Existing approaches to analyze syntactic knowledge of pre-trained language models have been limited.
Approach: They propose an unsupervised method that extracts constituency trees from PLM attention heads.
Outcome: The proposed method outperforms existing approaches if no development set is present.
Seshat: a Tool for Managing and Verifying Annotation Campaigns of Audio Data (2020.lrec-1)

Copied to clipboard

Challenge: Seshat is a software for the automated management of annotation campaigns for audio/speech data.
Approach: They propose a system for the automated management of annotation campaigns for audio/speech data which addresses these challenges.
Outcome: The proposed system computes an associated inter-annotator agreement with the gamma measure taking into account the categorisation and segmentation discrepancies.
A Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish Language (2022.lrec-1)

Copied to clipboard

Challenge: a team of linguists manually annotated a set of constituency trees.
Approach: They propose to create the first Turkish-based dependency-to-constituency conversion algorithm using a morphologic analyser and feature-based machine learning model.
Outcome: The proposed algorithm can be used to generate new constituency treebanks and training data for NLP resources like constituency parsers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations