GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration (2023.acl-demo)
Copied to clipboard
Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki, Akintunde Oladipo, Xinyu Zhang, Hailey Schoelkopf, Stella Biderman, Martin Potthast, Jimmy Lin
| Challenge: | Using the mature and well-tested methods from the domain of Information Retrieval (IR) we propose to integrate Pyserini with Hugging Face to provide qualitative analysis tools for NLP researchers. |
| Approach: | They propose to integrate Pyserini with Hugging Face to provide qualitative analysis tools for NLP researchers. |
| Outcome: | The proposed tools can be integrated with the Hugging Face ecosystem of open-source AI libraries and artifacts. |
Similar Papers
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face (2023.emnlp-demo)
Copied to clipboard
Christopher Akiki, Odunayo Ogundepo, Aleksandra Piktus, Xinyu Zhang, Akintunde Oladipo, Jimmy Lin, Martin Potthast
| Challenge: | a toolkit for reproducible information retrieval research is available for free. |
| Approach: | They present a tool that integrates Pyserini and Hugging Face to enable the seamless construction and deployment of interactive search engines. |
| Outcome: | The proposed tool makes state-of-the-art retrieval models more accessible to non-IR practitioners while minimizing deployment effort. |
GAIA: A Fine-grained Multimedia Knowledge Extraction System (2020.acl-demos)
Copied to clipboard
Manling Li, Alireza Zareian, Ying Lin, Xiaoman Pan, Spencer Whitehead, Brian Chen, Bo Wu, Heng Ji, Shih-Fu Chang, Clare Voss, Daniel Napierski, Marjorie Freedman
| Challenge: | Open source knowledge extraction tools are used for many real-world applications, but there is no comprehensive system for KE. |
| Approach: | They propose a multimedia knowledge extraction system that takes multimedia data from various sources and languages as input and creates a coherent, structured knowledge base. |
| Outcome: | The system achieves top performance at the recent NIST TAC SM-KBP2019 evaluation. |
TutorialBank: A Manually-Collected Corpus for Prerequisite Chains, Survey Extraction and Resource Recommendation (P18-1)
Copied to clipboard
Alexander Fabbri, Irene Li, Prawat Trairatvorakul, Yijiao He, Weitai Ting, Robert Tung, Caitlin Westerfield, Dragomir Radev
| Challenge: | TutorialBank is a publicly available dataset that aims to facilitate NLP education and research . a google search of "Natural Language Processing" returns over 100 million hits with papers, tutorials, 1 http://aan.how blog posts, codebases and other related online resources. |
| Approach: | They have manually collected and categorized over 5,600 resources on NLP . they have created a search engine and command-line tool to search the corpus . |
| Outcome: | The tutorial bank dataset is the largest manually-picked corpus of resources intended for NLP education . it includes lists of research topics, relevant resources for each topic, prerequisite relations among topics . |
Datasets: A Community Library for Natural Language Processing (2021.emnlp-demo)
Copied to clipboard
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, Thomas Wolf
| Challenge: | Contemporary NLP systems use many different datasets at significantly varying scale and level of annotation. |
| Approach: | a community library for contemporary NLP is available at https://github.com/datasets . the library includes more than 650 unique datasets and has more than 250 contributors a year after its initial development . |
| Outcome: | the library includes more than 650 unique datasets and has more than 250 contributors . it supports a variety of cross-dataset research projects and shared tasks . |
COMBO: State-of-the-Art Morphosyntactic Analysis (2021.emnlp-demo)
Copied to clipboard
| Challenge: | COMBO is an end-to-end NLP system for accurate part-of-speech tagging, morphological analysis, and (enhanced) dependency parsing. |
| Approach: | They propose a fully neural NLP system for accurate part-of-speech tagging, morphological analysis, lemmatisation, and (enhanced) dependency parsing. |
| Outcome: | The proposed system predicts categorical morphosyntactic features whilst also exposes their vector representations, extracted from hidden layers. |
NLP Scholar: An Interactive Visual Explorer for Natural Language Processing Literature (2020.acl-demos)
Copied to clipboard
| Challenge: | aCL Anthology and Google Scholar provide a single dataset of NLP papers and their meta-information . authors describe interactive visualizations that present various aspects of the data . |
| Approach: | They propose to use citation data from the ACL Anthology and Google Scholar to create a unified dataset of NLP papers and their meta-information. |
| Outcome: | The proposed dataset includes papers published in the area of their interest and by specified authors. |
NeuralQA: A Usable Library for Question Answering (Contextual Query Expansion + BERT) on Large Datasets (2020.emnlp-demos)
Copied to clipboard
| Challenge: | Existing tools for Question Answering (QA) have challenges that limit their use in practice. |
| Approach: | They propose a library that integrates with existing infrastructure and offers helpful defaults for QA subtasks. |
| Outcome: | NeuralQA integrates well with existing infrastructure and offers helpful defaults for QA subtasks. |
From Text to Context: Contextualizing Language with Humans, Groups, and Communities for Socially Aware NLP (2024.naacl-tutorials)
Copied to clipboard
Adithya V Ganesan, Siddharth Mangalik, Vasudha Varadarajan, Nikita Soni, Swanie Juhng, João Sedoc, H. Andrew Schwartz, Salvatore Giorgi, Ryan L Boyd
| Challenge: | This tutorial will cover the latest techniques and libraries for doing so at each level of analysis. |
| Approach: | This tutorial will cover the latest techniques and libraries for doing so at each level of analysis. |
| Outcome: | The tutorial covers human-centered techniques that provide benefit to traditional document- or word-level NLP tasks. |
Applying BERT to Document Retrieval with Birch (D19-3)
Copied to clipboard
| Challenge: | Birch is an open-source document retrieval system that integrates with the Anserini information retrieval toolkit to demonstrate end-to-end search over large document collections. |
| Approach: | They propose to integrate Anserini with a BERT-based document ranking model that provides an end-to-end open-source search engine. |
| Outcome: | The proposed system outperforms existing approaches to document retrieval and question answering on standard newswire and social media test collections. |
STREAM: Simplified Topic Retrieval, Exploration, and Analysis Module (2024.acl-short)
Copied to clipboard
| Challenge: | Topic modeling is a widely used technique to analyze large document corpora. |
| Approach: | They propose a module for topic retrieval, exploration, and analysis that implements multiple intruder-word based topic evaluation metrics. |
| Outcome: | The proposed module implements multiple intruder-word based topic evaluation metrics and extends existing datasets. |