A Lightweight Modeling Middleware for Corpus Processing (L18-1)

Copied to clipboard

Challenge: Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources.
Approach: They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools.
Outcome: The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface.

Similar Papers

The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences.
Approach: They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs.
Outcome: The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering .
Graph Matching and Graph Rewriting: GREW tools for corpus exploration, maintenance and conversion (2021.eacl-demos)

Copied to clipboard

Challenge: Graph Rewriting is a mathematical formalism that can be used to describe rule-based transformations on linguistic structures.
Approach: They propose to use graph rewriting to describe rule-based transformations on linguistic structures.
Outcome: The proposed tools can be used to compute rule-based transformations on linguistic structures.
Corpus Considerations for Annotator Modeling and Scaling (2024.naacl-long)

Copied to clipboard

Challenge: Recent trends in natural language processing and annotation tasks emphasize individual perspectives . annotator models that rely on a single ground truth may disregard valuable minority perspectives omissions .
Approach: They propose a composite embedding approach to investigate annotator modeling techniques . they show that the commonly used user token model consistently outperforms more complex models .
Outcome: The proposed model outperforms more complex models on a given dataset.
Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Grounding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods fragment document parsing into pipeline of separated subtasks, resulting in incomplete semantics and error propagation.
Approach: They propose an end-to-end document parsing framework that leverages vision-language priors of MLLMs.
Outcome: The proposed method surpasses existing methods significantly in document parsing . it leverages the vision-language priors of MLLMs to decouple parse and layout grounding based on visual information.
Lightweight Grammatical Annotation in the TEI: New Perspectives (L18-1)

Copied to clipboard

Challenge: a small set of descriptive devices have been made available for lightweight linguistic annotation . merit of a predefined TEI tagset is the homogeneity of tagging and better interoperability of simple linguistic resources encoded in the TE.
Approach: They propose a new attribute class that would gather token-level attributes facilitating simple linguistic annotation.
Outcome: The proposed attribute class addresses community feedback on the lack of a specific tagset for lightweight linguistic annotation within the TEI.
End-to-end Parsing of Procedural Text into Flow Graphs (2024.lrec-main)

Copied to clipboard

Challenge: Existing flow graph parsers lack sufficient annotated data to train them . a lack of annotation can cause costly training, and poor flow graph training results in a large improvement.
Approach: They propose a multi-task framework that performs tagging and graph generation simultaneously . they take advantage of the abundance of unlabelled recipes and generate noisy silver annotations .
Outcome: The proposed model can unify the input representation and use compact encoders, resulting in small models with significantly fewer parameters than existing models.
A Parser for LTAG and Frame Semantics (L18-1)

Copied to clipboard

Challenge: Existing parsers for Lexicalized Tree Adjoining Grammars and frame semantics are difficult to use due to the size of the resources to develop.
Approach: They propose a parser which uses Lexicalized Tree Adjoining Grammars and frame semantics to combine them.
Outcome: The proposed grammars are based on Lexicalized Tree Adjoining Grammars and frame semantics.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.
Syntactic Search by Example (2020.acl-demos)

Copied to clipboard

Challenge: a new system allows a user to search a large linguistically annotated corpus using syntactic patterns over dependency graphs.
Approach: They propose a query language that allows a user to search a large linguistically annotated corpus using syntactic patterns over dependency graphs.
Outcome: The proposed system searches the English wikipedia and English pubmed abstracts at a rapid speed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations