• An open-access EEG dataset for speech decoding is available (Nature), supporting benchmarking and transfer for EEG-based speech BCIs.1
  • The dataset enables exploration of the role of articulation and coarticulation in decoding pipelines.1
  • It directly supports near-term speech prosthesis and BCI methods development.1 1

Weekly enrichment (2026-07-20)

  • The source paper (Moreira et al., Scientific Data, Nature, 2025) releases two validated EEG datasets (N = 8 and N = 16) for speech decoding at the phoneme and word level and by articulatory phoneme properties.2
  • EEG was recorded from 64 channels while subjects listened to and repeated 6 consonants (/b/, /p/, /d/, /t/, /s/, /z/) and 5 vowels (/i/, /ɛ/, /ɑ/, /u/, /oʊ/), yielding 11 phonemes, 40 consonant–vowel pairs, 20 real words, and 20 pseudowords to capture coarticulation.2
  • The stimulus set is scaffolded by articulatory features (place, manner, voicing) so classifiers can be tested on progressively more complex, naturalistic phonetic environments; consonants were chosen as unique feature combinations (bilabial/alveolar, stop/fricative, voiced/unvoiced).2
  • Uniquely, trials were collected both in a control condition and during transcranial magnetic stimulation (TMS) targeted to motor-cortex articulator regions to test whether neuromodulation augments the speech-related EEG signal—the authors state it is the first neural speech-decoding dataset to combine neuromodulation with coarticulation-structured stimuli.2
  • TMS was individualized: resting motor threshold set via first-dorsal-interosseous MEPs (≥50 μV in 5/10 trials), with stimulation at 110% rMT following the reference protocol; the two collection rounds occurred in 2019 and 2021.2
  • The two datasets (different time points, same stimulus types) are designed for external validation: the authors recommend developing models on the larger set and validating on the smaller, stricter set to reduce overfitting.2
  • The public release (OpenNeuro/NEMAR ds006104) contains recordings from 24 participants, 61 EEG channels, two sessions, ~309 files (~83.8 GB), under a permissive license, with analysis code (EEGLAB-based) on OSF and GitHub.3
  • Context: independent honest benchmarking on ds006104 (16 subjects, 256 Hz) under strict leave-one-subject-out found five-class vowel decoding only ~24–25% accurate (chance 20%, p < 0.001), with classical ML matching deep learning—signal is real but weak.4
  • Context: a Conformer-based benchmark (CIPHER) on ds006104 showed binary articulatory tasks reach near-ceiling but are explained by acoustic-onset and TMS-target confounds; the confound-controlled 11-class CVC phoneme task had a best word error rate of ~0.671, far from practical free-form decoding.5
  • Context: reviews of EEG imagined/covert-speech decoding emphasize that low signal-to-noise ratio and limited data are the core bottlenecks, motivating shared naturalistic datasets like this one.6

Footnotes

  1. https://news.google.com/rss/articles/CBMiX0FVX3lxTFBkcTNOanJFZWdOaDBwZkFEWnNCMDJvSUZwYXNjc0tmc0NVSU1zV0ZCZ0dUZjZCQ2pXejlJYlE1dlA5b0hIUDRCR3JPeUVmQWY1eGgwblFLXy1FYUVsaEw4?oc=5 2 3 4

  2. https://www.nature.com/articles/s41597-025-05187-2 2 3 4 5 6

  3. https://nemar.org/dataexplorer/detail?dataset_id=ds006104

  4. https://arxiv.org/html/2605.00865v1

  5. https://www.arxiv.org/pdf/2604.02362

  6. https://www.frontiersin.org/journals/human-neuroscience/articles/10.3389/fnhum.2022.867281/full