• An open EEG dataset supports multimodal semantic alignment across reading and listening.1
  • The dataset enables reproducible decoders and language/BCI pipelines for EEG-based BCI and neural decoding (Nature source; tier-1 for immediate use).1 1

Weekly enrichment (2026-07-20)

  • The dataset is ChineseEEG-2 (Chen, Li, He et al.), published in Nature’s Scientific Data (2026; DOI 10.1038/s41597-025-06466-8) with a preprint at arXiv:2508.04240, from the NCC Lab at SUSTech.2 3
  • It is described as the first high-density, token-aligned Chinese EEG benchmark enabling cross-modal semantic alignment between reading aloud and passive listening within a unified Chinese corpus, addressing the English-centric bias of existing language EEG datasets.4 3
  • The dataset totals ~32.4 hours of EEG from 12 participants: 4 readers recorded during ~10.7-10.8 hours of reading aloud, whose audio was then replayed to 8 listeners to collect ~21.6 hours of passive-listening EEG aligned to identical semantic content.2 4
  • EEG was recorded with a 128-channel EGI Geodesic Sensor Net (GSN-HydroCel-128, GES 400 series) at 250 Hz for reading-aloud and 1000 Hz for the passive-listening tasks.4
  • Beyond raw EEG, raw text, and audio, the release provides derivatives: preprocessed EEG plus precomputed semantic embeddings from Wav2Vec2 (audio) and BERT-base-Chinese (text), enabling direct brain-to-LLM/MLLM alignment benchmarking.2 4
  • Raw EEG is distributed in standardized BrainVision formats (.vhdr header, .vmrk marker, .eeg signal), and the full dataset is openly available on Science Data Bank (DOI 10.57760/sciencedb.CHNNeuro.00001).2 4
  • The GitHub repository ships full experimental code, preprocessing pipelines, and analysis tools including inter-subject correlation (ISC), source reconstruction, and stimulus decoding, complementing the earlier silent-reading ChineseEEG dataset to jointly cover speaking, listening, and reading modalities.4 3

Footnotes

  1. https://news.google.com/rss/articles/CBMiX0FVX3lxTE1RU1RkaExpYlNrdncxTkxMc0p2aDJMTmdMWUU2WW5HMG9RUko0RmZhZjlxLVhkbUxsQjM4NzUwV3dFRmdLU1l2czltVzA0ZGZIemlFZllleDdtOXI3cXdB?oc=5 2 3

  2. https://www.nature.com/articles/s41597-025-06466-8 2 3 4

  3. https://arxiv.org/abs/2508.04240 2 3

  4. https://github.com/ncclab-sustech/ChineseEEG-2 2 3 4 5 6