• A new open ECoG dataset supports natural language comprehension research and modeling of neural activity during ecologically valid listening.1
  • The dataset enables reproducible neural NLP and speech prosthesis research with high-quality iEEG.1
  • It is published in Scientific Data (Nature) and is decision-useful for the next 12 months.1 1

Gardner updates

  • The open ECoG dataset enables reproducible neural NLP and speech prosthesis research with high-quality iEEG during ecologically valid listening (Scientific Data/Nature). 1

Weekly enrichment (2026-07-20)

  • The “Podcast” ECoG dataset (Zada, Nastase, Goldstein, Devinsky, Flinker, Hasson and colleagues; published in Nature’s Scientific Data, 2025) shares human intracranial recordings from 9 participants with 1,330 electrodes (grid, strip, and depth), collected while they listened to a single ~30-minute audio story containing over 5,000 words.2 3
  • Data were recorded during clinical epilepsy monitoring at NYU’s Comprehensive Epilepsy Center; three participants used FDA-approved hybrid grids that interleave 64 standard 2 mm contacts with 64 additional 1 mm contacts for denser sampling.2
  • Signals were acquired with clinical amplifiers (NicoletOne at 512 Hz or a NeuroWorks Quantum at 2,048 Hz) and, in the shared preprocessed version, downsampled to 512 Hz, cleaned (bad-channel removal, despiking, common-average referencing, notch filtering), and reduced to high-gamma band power via a 70–200 Hz band-pass and Hilbert envelope.2
  • The stimulus is the “This American Life” segment “Monkey in the Middle”; word onsets/offsets were estimated with the Penn Phonetics Lab Forced Aligner and then manually corrected, yielding a high-resolution aligned transcript.2
  • The release includes five linguistic feature spaces: 80 mel-spectral acoustic bins, a 44-dimensional phoneme representation (plus articulation groupings), syntactic features (50 part-of-speech tags and 45 dependency relations), non-contextual embeddings (e.g., GloVe, 300-dimensional), and contextual large language model embeddings.2
  • In electrode-wise linear encoding analyses, contextual embeddings from GPT-2 XL explained the most variance across nearly all tested electrodes, echoing prior findings of alignment between LLM internal representations and human language-related neural activity.2
  • The dataset follows the BIDS-iEEG standard, ships raw data in EDF plus preprocessed data in MNE .fif format, includes de-faced T1 MRIs with MNI-registered electrode coordinates, and is released under a CC0 license.2
  • It is openly available on OpenNeuro (ds005574) with reproducible tutorials for preprocessing, feature extraction, and encoding analyses, making it a reusable resource for naturalistic language decoding and speech-BCI research.2 4

Footnotes

  1. https://news.google.com/rss/articles/CBMiX0FVX3lxTFBQWHZQX293TUszQkoyWlBIWEhoUHhJUERuUTAzb3d3a1UyYUJDMDJrVFk0S29sTjk5SlVacUNRQTM1b2JBNFZUejBNWmJZUENqU21zNEVzMG1qYlVxTmk4?oc=5 2 3 4 5

  2. https://www.nature.com/articles/s41597-025-05462-2 2 3 4 5 6 7 8

  3. https://doi.org/10.1038/s41597-025-05462-2

  4. https://doi.org/10.18112/openneuro.ds005574.v1.0.2