The Alice Datasets: fMRI & EEG Observations of Natural Language Comprehension (2020.lrec-1)
Copied to clipboard
| Challenge: | "naturalistic" stimuli are now offering a new way to study language comprehension in the brain, in synergy with natural language processing tools. |
| Approach: | They propose to use a set of datasets from a story in English to test new linguistic and computational hypotheses about natural language comprehension in the brain. |
| Outcome: | The Alice Datasets are a set of datasets based on magnetic resonance and electrophysiological data, collected while participants heard a story in English. |
Similar Papers
English Machine Reading Comprehension Datasets: A Survey (2021.emnlp-main)
Copied to clipboard
| Challenge: | a survey of English Machine Reading Comprehension datasets is carried out . the aim is to provide a concise yet informative overview of the landscape . |
| Approach: | They survey 60 English Machine Reading Comprehension datasets to provide a resource for other researchers interested in this problem. |
| Outcome: | The proposed survey covers 60 English MRC datasets with a view to providing a resource for other researchers interested in the problem. |
The Language of Brain Signals: Natural Language Processing of Electroencephalography Reports (2020.lrec-1)
Copied to clipboard
| Challenge: | Clinical electroencephalography (EEG) is an excellent tool for probing neural function. |
| Approach: | They propose to use EEG to capture brain signals and its correlations with pathologies by a corpus of EEG reports to provide examples of EMG-specific concepts. |
| Outcome: | The proposed method provides examples of EEG-specific and clinically relevant concepts and exemplifies a self-attention joint-learning model to predict similar annotations in the EEG report corpus. |
Model-based analysis of brain activity reveals the hierarchy of language in 305 subjects (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a popular approach to decompose the neural bases of language requires large and costly data sets to obtain. |
| Approach: | They propose a model-based approach to decompose the neural bases of language that can be used to correlate brain responses to different stimuli. |
| Outcome: | The proposed model-based approach replicates the seminal study of Lerner et al. (2011), which revealed the hierarchy of language areas by comparing the functional-magnetic resonance imaging (fMRI) of seven subjects listening to 7min of both regular and scrambled narratives. |
Mapping Brains with Language Models: A Survey (2023.findings-acl)
Copied to clipboard
| Challenge: | accumulated evidence for brain and language model activations remains ambiguous, but correlations with model size and quality provide grounds for cautious optimism. |
| Approach: | They examine the evidence accumulated by 30 studies spanning 10 datasets and 8 metrics to determine whether there is any overlap between brain and language model activations. |
| Outcome: | The findings suggest that representations extracted from NLP models can (partially) explain the signal found in neural data. |
Encoding and Decoding Language in the Brain with Language Models (2026.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and brain-based fine-caching with language models. |
| Approach: | This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and scaling with language models. |
| Outcome: | This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and decoding with language models. |
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants (2024.acl-long)
Copied to clipboard
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa
| Challenge: | Existing benchmarks for text comprehension only cover 30 languages, but lack of labeled data is a major obstacle to building functional systems in most languages. |
| Approach: | They present a multiple-choice machine reading comprehension dataset spanning 122 languages . they use it to evaluate the capabilities of multilingual masked language models and large language models . |
| Outcome: | The proposed dataset enables the evaluation of text models in high-, medium- and low-resource languages. |
ZuCo 2.0: A Dataset of Physiological Recordings During Natural Reading and Annotation (2020.lrec-1)
Copied to clipboard
| Challenge: | a new dataset of eye-tracking and electroencephalography captures language understanding . eye movement data provides millisecond-accurate records of where humans look when reading . |
| Approach: | They recorded and preprocessed eye-tracking and electroencephalography data during natural reading and during annotation. |
| Outcome: | The study combines eye-tracking and electroencephalography to capture the reading process . the data can be used to evaluate state-of-the-art machine learning systems . |
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge (L18-1)
Copied to clipboard
| Challenge: | Various approaches for script knowledge extraction and processing have been proposed in recent years. |
| Approach: | They propose a dataset to evaluate natural language understanding approaches based on commonsense knowledge. |
| Outcome: | The proposed dataset provides test cases for the broader natural language understanding community. |
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English (2024.lrec-main)
Copied to clipboard
| Challenge: | Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research. |
| Approach: | They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse. |
| Outcome: | The proposed dataset is the largest of its kind that is publicly accessible. |
Comprehensive Multi-Dataset Evaluation of Reading Comprehension (D19-58)
Copied to clipboard
| Challenge: | Recent research aims to facilitate training and evaluation on several reading comprehension datasets at the same time. |
| Approach: | They propose an evaluation server that reports performance on seven diverse reading comprehension datasets and includes synthetic augmentations to test models' ability to handle out-of-domain questions. |
| Outcome: | The evaluation server performs on seven reading comprehension datasets, and collects and includes synthetic augmentations for these datasets to test models' ability to handle out-of-domain questions. |