• The 3M-CPSEED dataset provides EEG recordings for Chinese Pinyin production across overt, mouthed, and imagined speech.1
  • It supports speech prosthesis and decoding research, including for speech BCI.1
  • The dataset is published in Nature and has clear BCI/speech relevance.1 1

Weekly enrichment (2026-07-20)

  • 3M-CPSEED records the same Chinese Pinyin production in three modes—overt speech, silently articulated (“mouthed”) speech, and imagined speech—so researchers can compare neural activity across production modalities within one dataset.23
  • The dataset comprises EEG from 20 healthy, right-handed native Mandarin speakers (mean age 24.55 years, SD 2.58; 11 female, 9 male), each completing four experimental blocks in a single day.24
  • It yields 1,800 validated trials built from a compact, articulation-balanced syllable set: finals “a, i, u, ü” and initials “m, f, j, l, k, ch,” chosen to span distinct articulatory positions.24
  • The design rationale is that most existing Chinese EEG speech datasets use full sentences and a single paradigm, whereas Pinyin is the phonetic foundation of Chinese characters, enabling decoding of individual character components and transfer learning to other alphabetic languages.2
  • The data are openly released on OpenNeuro as accession ds006465 under a CC0 license, totaling roughly 80 recordings, about 36.2 hours, and 8.2 GB, sampled at 500 Hz in EDF format.34
  • Channel counts vary by participant (32-channel for most, plus 126- and 33-channel subsets), reflecting different montages across the cohort.3
  • The dataset is programmatically loadable via the EEGDash Python library (canonical aliases CPSEED_3M / CPSEED), which streams it into braindecode/PyTorch datasets on demand.3
  • It was published in Scientific Data (2025), a Nature Portfolio data journal, rather than the main Nature journal noted in the original bullets.2

Footnotes

  1. https://news.google.com/rss/articles/CBMiX0FVX3lxTE5OM3FJajBNek1GQUM0RkxLQXZ1OVVMNGVfVzc3UGh4Rkw0ZlBiS25XRUpQMWdsbll0WWdBQXlwcmxLZ3BRWUwxZ2h5OTNGeEtMRXUyLXhBYnIwaVpoWVBr?oc=5 2 3 4

  2. https://www.nature.com/articles/s41597-025-06346-1 2 3 4 5

  3. https://huggingface.co/datasets/EEGDash/ds006465 2 3 4

  4. https://nemar.org/dataexplorer/detail?dataset_id=ds006465 2 3