Challenge: KonText is a corpus query interface built on top of core NoSketch Engine core libraries.
Approach: KonText is built on top of core NoSketch Engine core libraries . it provides integration capabilities allowing connection of basic corpus search service with other languages .
Outcome: The proposed interface is built on top of core libraries of the open-source corpus search engine NoSketch Engine (NoSkE) it overcomes some limitations and provides integration capabilities with other language resources.

Similar Papers

To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)

Copied to clipboard

Challenge: a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics .
Approach: They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development .
Outcome: This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development .
Facilitating Corpus Usage: Making Icelandic Corpora More Accessible for Researchers and Language Users (2020.lrec-1)

Copied to clipboard

Challenge: Gigaword corpus is a large text corpus used in natural language processing . large corpora are needed to achieve better performance in the field of NLP .
Approach: They propose a set of tools to facilitate the use of the Icelandic Gigaword Corpus . they provide n-grams based on the corpus, and a variety of pre-trained word embeddings models .
Outcome: The proposed tools facilitate the use of the Icelandic Gigaword corpus in the field of Natural Language Processing and other fields.
OpusFilter: A Configurable Parallel Corpus Filtering Toolbox (2020.acl-demos)

Copied to clipboard

Challenge: OpusFilter is a toolbox for filtering parallel corpora using noisy training data.
Approach: They propose a toolbox for filtering parallel corpora with heuristic filters, language identification libraries, character-based language models and word alignment tools.
Outcome: The proposed tool outperforms a similar tool on a Finnish-English news translation task using noisy web crawls.
Syntactic Search by Example (2020.acl-demos)

Copied to clipboard

Challenge: a new system allows a user to search a large linguistically annotated corpus using syntactic patterns over dependency graphs.
Approach: They propose a query language that allows a user to search a large linguistically annotated corpus using syntactic patterns over dependency graphs.
Outcome: The proposed system searches the English wikipedia and English pubmed abstracts at a rapid speed.
A Lightweight Modeling Middleware for Corpus Processing (L18-1)

Copied to clipboard

Challenge: Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources.
Approach: They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools.
Outcome: The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface.
Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrAC (2024.lrec-main)

Copied to clipboard

Challenge: OntoLex is the dominant community standard for machine-readable lexical resources . it is currently extended with a designated module for Frequency, Attestations and Corpus-based Information .
Approach: They propose a module for Frequency, Attestations and Corpus-based Information for OntoLex . the module enables RDF-based web services to exchange corpus queries dynamically .
Outcome: The proposed module addresses the incorporation of corpus queries for linking dictionaries with corpus engines and enabling RDF-based web services to exchange corpus query data dynamically.
BERT-QE: Contextualized Query Expansion for Document Re-ranking (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to expand query use pseudo relevance feedback (PRF) but they are under-equipped to evaluate the relevance of information pieces used for expansion.
Approach: They propose a query expansion model that leverages the BERT model to select relevant document chunks for expansion.
Outcome: The proposed model significantly outperforms existing models on the TREC Robust04 and GOV2 test collections.
NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System (L18-1)

Copied to clipboard

Challenge: NL2Bash is a new semantic parsing problem for mapping English sentences to Bash commands.
Approach: They propose a dataset of English commands and expert-written Bash commands to map English sentences to Bash.
Outcome: The proposed methods are significantly larger (from two to ten times) than most existing benchmarks.
Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language Resources (2023.emnlp-main)

Copied to clipboard

Challenge: Existing research on the Somali language information retrieval relies on query translation . lack of digital resources is key obstacle to advancing language technologies .
Approach: They develop an annotated corpus for Somali information retrieval using query expansion technique.
Outcome: The proposed corpus comprises 2335 documents collected from well-known online sites . it can be used for text classification-related tasks and question-answering research purposes.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations