Challenge: "naturalistic" stimuli are now offering a new way to study language comprehension in the brain, in synergy with natural language processing tools.
Approach: They propose to use a set of datasets from a story in English to test new linguistic and computational hypotheses about natural language comprehension in the brain.
Outcome: The Alice Datasets are a set of datasets based on magnetic resonance and electrophysiological data, collected while participants heard a story in English.

Similar Papers

English Machine Reading Comprehension Datasets: A Survey (2021.emnlp-main)

Copied to clipboard

Challenge: a survey of English Machine Reading Comprehension datasets is carried out . the aim is to provide a concise yet informative overview of the landscape .
Approach: They survey 60 English Machine Reading Comprehension datasets to provide a resource for other researchers interested in this problem.
Outcome: The proposed survey covers 60 English MRC datasets with a view to providing a resource for other researchers interested in the problem.
The Language of Brain Signals: Natural Language Processing of Electroencephalography Reports (2020.lrec-1)

Copied to clipboard

Challenge: Clinical electroencephalography (EEG) is an excellent tool for probing neural function.
Approach: They propose to use EEG to capture brain signals and its correlations with pathologies by a corpus of EEG reports to provide examples of EMG-specific concepts.
Outcome: The proposed method provides examples of EEG-specific and clinically relevant concepts and exemplifies a self-attention joint-learning model to predict similar annotations in the EEG report corpus.
Model-based analysis of brain activity reveals the hierarchy of language in 305 subjects (2021.findings-emnlp)

Copied to clipboard

Challenge: a popular approach to decompose the neural bases of language requires large and costly data sets to obtain.
Approach: They propose a model-based approach to decompose the neural bases of language that can be used to correlate brain responses to different stimuli.
Outcome: The proposed model-based approach replicates the seminal study of Lerner et al. (2011), which revealed the hierarchy of language areas by comparing the functional-magnetic resonance imaging (fMRI) of seven subjects listening to 7min of both regular and scrambled narratives.
Mapping Brains with Language Models: A Survey (2023.findings-acl)

Copied to clipboard

Challenge: accumulated evidence for brain and language model activations remains ambiguous, but correlations with model size and quality provide grounds for cautious optimism.
Approach: They examine the evidence accumulated by 30 studies spanning 10 datasets and 8 metrics to determine whether there is any overlap between brain and language model activations.
Outcome: The findings suggest that representations extracted from NLP models can (partially) explain the signal found in neural data.
Encoding and Decoding Language in the Brain with Language Models (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and brain-based fine-caching with language models.
Approach: This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and scaling with language models.
Outcome: This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and decoding with language models.
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for text comprehension only cover 30 languages, but lack of labeled data is a major obstacle to building functional systems in most languages.
Approach: They present a multiple-choice machine reading comprehension dataset spanning 122 languages . they use it to evaluate the capabilities of multilingual masked language models and large language models .
Outcome: The proposed dataset enables the evaluation of text models in high-, medium- and low-resource languages.
ZuCo 2.0: A Dataset of Physiological Recordings During Natural Reading and Annotation (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset of eye-tracking and electroencephalography captures language understanding . eye movement data provides millisecond-accurate records of where humans look when reading .
Approach: They recorded and preprocessed eye-tracking and electroencephalography data during natural reading and during annotation.
Outcome: The study combines eye-tracking and electroencephalography to capture the reading process . the data can be used to evaluate state-of-the-art machine learning systems .
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge (L18-1)

Copied to clipboard

Challenge: Various approaches for script knowledge extraction and processing have been proposed in recent years.
Approach: They propose a dataset to evaluate natural language understanding approaches based on commonsense knowledge.
Outcome: The proposed dataset provides test cases for the broader natural language understanding community.
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English (2024.lrec-main)

Copied to clipboard

Challenge: Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research.
Approach: They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse.
Outcome: The proposed dataset is the largest of its kind that is publicly accessible.
Comprehensive Multi-Dataset Evaluation of Reading Comprehension (D19-58)

Copied to clipboard

Challenge: Recent research aims to facilitate training and evaluation on several reading comprehension datasets at the same time.
Approach: They propose an evaluation server that reports performance on seven diverse reading comprehension datasets and includes synthetic augmentations to test models' ability to handle out-of-domain questions.
Outcome: The evaluation server performs on seven reading comprehension datasets, and collects and includes synthetic augmentations for these datasets to test models' ability to handle out-of-domain questions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations