- ArEEG is an open-access EEG dataset for Arabic inner speech, directly supporting EEG-based BCI and speech decoding research.1
- The dataset enables reproducible methods and expands language diversity in inner-speech BCI; it is published in a Nature source with a clear implementation path and is tier-1.1 1
Weekly enrichment (2026-07-20)
- ArEEG is the first open-access EEG dataset capturing Arabic inner (imagined) speech, published in Scientific Data (Nature portfolio, 2025; Metwalli et al., DOI 10.1038/s41597-025-05387-w).2
- Recordings come from 12 native Arabic speakers with balanced gender and ages 17–25, using an 8-channel Unicorn Hybrid Black+ headset at a 250 Hz sampling rate.2
- The task uses five inner-speech command classes — Up, Down, Left, Right, Select — deliberately exceeding the typical four-class inner-speech datasets while keeping the low, cost-effective electrode count.2
- The corpus totals 4,650 trials: each participant recorded 375 trials across 15 sessions on separate days, except subject 3 who contributed 525 trials over 21 sessions; class order was randomized per trial with each class presented equally.2
- A core design goal is cost-effective BCI: the authors argue that reliance on high-electrode-count rigs blocks affordable systems, so an 8-electrode consumer headset lowers the barrier for Arabic-speaking regions.2
- The dataset is hosted publicly on OpenNeuro (accession ds005262) and Kaggle, with open-source Python loading/preprocessing and ML pipelines (NumPy, Pandas, Scikit-learn, MNE) provided for reproducibility.23
- Inner-speech decoding remains hard on this data: a 2025 benchmark using EEGNet on raw signals reached only ~27.1% per-subject accuracy, and a KNN classifier with Relative Wavelet Energy plus Gabor-transform features (ANOVA-selected) reached 31.03% in subject-dependent analysis (chance is 20% for five classes).4
- ArEEG complements the related ArEEG_Words dataset (22 participants, 14-channel Emotiv Epoc X, 16 Arabic words, 15,360 segmented 250 ms signals), together broadening non-English EEG resources for speech BCIs.5
- BCI implication: language-specific inner-speech corpora like ArEEG target assistive communication for motor/communication-impaired Arabic speakers and enable reproducible benchmarking of low-cost, non-invasive speech-decoding pipelines.2