• ArEEG is an open-access EEG dataset for Arabic inner speech, directly supporting EEG-based BCI and speech decoding research.1
  • The dataset enables reproducible methods and expands language diversity in inner-speech BCI; it is published in a Nature source with a clear implementation path and is tier-1.1 1

Weekly enrichment (2026-07-20)

  • ArEEG is the first open-access EEG dataset capturing Arabic inner (imagined) speech, published in Scientific Data (Nature portfolio, 2025; Metwalli et al., DOI 10.1038/s41597-025-05387-w).2
  • Recordings come from 12 native Arabic speakers with balanced gender and ages 17–25, using an 8-channel Unicorn Hybrid Black+ headset at a 250 Hz sampling rate.2
  • The task uses five inner-speech command classes — Up, Down, Left, Right, Select — deliberately exceeding the typical four-class inner-speech datasets while keeping the low, cost-effective electrode count.2
  • The corpus totals 4,650 trials: each participant recorded 375 trials across 15 sessions on separate days, except subject 3 who contributed 525 trials over 21 sessions; class order was randomized per trial with each class presented equally.2
  • A core design goal is cost-effective BCI: the authors argue that reliance on high-electrode-count rigs blocks affordable systems, so an 8-electrode consumer headset lowers the barrier for Arabic-speaking regions.2
  • The dataset is hosted publicly on OpenNeuro (accession ds005262) and Kaggle, with open-source Python loading/preprocessing and ML pipelines (NumPy, Pandas, Scikit-learn, MNE) provided for reproducibility.23
  • Inner-speech decoding remains hard on this data: a 2025 benchmark using EEGNet on raw signals reached only ~27.1% per-subject accuracy, and a KNN classifier with Relative Wavelet Energy plus Gabor-transform features (ANOVA-selected) reached 31.03% in subject-dependent analysis (chance is 20% for five classes).4
  • ArEEG complements the related ArEEG_Words dataset (22 participants, 14-channel Emotiv Epoc X, 16 Arabic words, 15,360 segmented 250 ms signals), together broadening non-English EEG resources for speech BCIs.5
  • BCI implication: language-specific inner-speech corpora like ArEEG target assistive communication for motor/communication-impaired Arabic speakers and enable reproducible benchmarking of low-cost, non-invasive speech-decoding pipelines.2

Footnotes

  1. https://news.google.com/rss/articles/CBMiX0FVX3lxTE5aS2NIcWNrMmc1YTFDNHVfMVo5SEtoTGNtcm15dDY0VnM5NUFLbDcxQ01Fajh3NnFfR292UUo0X1BRaGpkMEluX0NuUTk5ZTZ4d0Z1NWVlRnFHWFZZekln?oc=5 2 3

  2. https://www.nature.com/articles/s41597-025-05387-w 2 3 4 5 6 7

  3. https://openneuro.org/datasets/ds005262/

  4. https://doi.org/10.1109/iceeng64546.2025.11031298

  5. https://arxiv.org/pdf/2411.18888