A Multilingual Approach to Question Classification (L18-1)

Copied to clipboard

Challenge: Existing work on questions has focused on understanding the structure of questions per se . a few approaches explicitly focus on information-seeking questions, but this work is either based on big data or crowdsourcing.
Approach: They propose a dependency-parsed, parallel multilingual corpus of information-seeking and non-information-seeing questions . they employ a linguistically motivated rule-based system that uses linguistic cues from one language to help classify questions across other languages.
Outcome: The proposed system correctly classifies questions in 79% of cases, compared to other systems.

Similar Papers

XOR QA: Cross-lingual Open-Retrieval Question Answering (2021.naacl-main)

Copied to clipboard

Challenge: a dataset of 40k information-seeking questions across seven languages is used to answer multilingual question answering tasks.
Approach: They propose a task framework that allows questions from one language to be answered via answer content from another language.
Outcome: The proposed framework can be used to answer questions from one language to another . the dataset was built on 40K questions across 7 languages, but could not find same-language answers .
A Corpus for Multilingual Document Classification in Eight Languages (L18-1)

Copied to clipboard

Challenge: a subset of the Reuters corpus volume 2 is used to evaluate cross-lingual document classification . current best practice is to evaluate document classification on resources in one language and transfer it to another without additional resources.
Approach: They propose to use a subset of the Reuters corpus to evaluate cross-lingual document classification . they propose to add Italian, Russian, Japanese and Chinese to the subset .
Outcome: The proposed subset of the Reuters corpus has balanced class priors for eight languages.
A System for Answering Simple Questions in Multiple Languages (2023.acl-demo)

Copied to clipboard

Challenge: Existing knowledge graph question answering systems are limited to simple questions, but they can be used to answer complex questions.
Approach: They propose a multilingual Knowledge Graph Question Answering technique that orders potential responses based on the distance between the question’s text embeddings and the answer’s graph embedds.
Outcome: The proposed method consistently outperforms baseline systems, including seq2seq QA models and complex rule-based pipelines.
MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering (2021.tacl-1)

Copied to clipboard

Challenge: Existing multilingual QA datasets lack linguistic diversity and comparable evaluation between languages.
Approach: They propose a multilingual question-answer evaluation set with 10k English queries and human translations of them into 25 additional languages and dialects.
Outcome: The proposed model is based on a multilingual knowledge questions and answers evaluation set with 26 languages.
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages (2020.tacl-1)

Copied to clipboard

Challenge: Existing models for multilingual modeling are based on a set of typological features that are used to express meaning in languages such as English.
Approach: They present a question-answer-typed question-referenced dataset that covers 11 typologically diverse languages with 204K question-and-answered pairs.
Outcome: The proposed dataset covers 11 typologically diverse languages with 204K question-answer pairs.
What Are the Implications of Your Question? Non-Information Seeking Question-Type Identification in CNN Transcripts (2024.lrec-main)

Copied to clipboard

Challenge: Non-information seeking questions capture subtle dynamics of human discourse . authors use dataset of over 1,500 information-seeking questions and NISQs as benchmark .
Approach: They use a dataset of over 1,500 information-seeking question(ISQ) and NISQ to evaluate human and machine performance on classifying fine-grained NISq types.
Outcome: The proposed corpus is the first publicly available for annotation of non-information seeking questions . it evaluates human and machine performance on classifying fine-grained questions based on models .
A Practical Toolkit for Multilingual Question and Answer Generation (2023.acl-demo)

Copied to clipboard

Challenge: Generating questions and answers from text is a challenging task due to the expected structured output.
Approach: They propose an online service for multilingual QAG along with a python package for model fine-tuning, generation, and evaluation.
Outcome: The proposed model is available in eight languages and can be used online or locally via lmqg.
WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Frame-semantic annotations exist for a tiny fraction of the world’s languages, however, Wikidata provides a common, distant supervision signal for semantic parsers.
Approach: They propose a multilingual resource with partial semantic dependency structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch.
Outcome: The proposed resource can be used to augment pre-existing resources or reduce the annotation effort for low-resource languages.
Cross-Lingual Training for Automatic Question Generation (P19-1)

Copied to clipboard

Challenge: Automatic question generation is a challenging problem in natural language understanding . manual curating a dataset of comparable size for a new language is tedious and expensive.
Approach: They propose to reuse available large QG dataset in a secondary language to learn a QG model for a primary language.
Outcome: The proposed model outperforms baseline models in Hindi and Chinese.
Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Existing automatic question generation datasets focus on English, resulting in data gaps for other languages.
Approach: They propose a cross-lingual transfer method that allows models to generate questions in low-resource languages.
Outcome: The proposed method outperforms other models and achieves comparable performance across languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations