Skip to main navigation Skip to search Skip to main content

VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in Tonal Language

  • Tianyu He
  • , Mingyi Wei
  • , Ruicong Wang
  • , Renzhi Wang
  • , Shiwei Du
  • , Siqi Cai*
  • , Wei Tao*
  • , Haizhou Li*
  • *Corresponding author for this work
  • The Chinese University of Hong Kong, Shenzhen
  • Shenzhen University

Research output: Contribution to journalArticlepeer-review

Abstract

Speech BCIs based on implanted electrodes hold significant promise for enhancing spoken communication through high temporal resolution and invasive neural sensing. Despite the potential, acquiring such data is challenging due to its invasive nature, and publicly available datasets, particularly for tonal languages, are limited. In this study, we introduce VocalMind, a stereotactic electroencephalography (sEEG) dataset focused on Mandarin Chinese, a tonal language. This dataset includes sEEG-speech parallel recordings from three distinct speech modes, namely vocalized speech, mimed speech, and imagined speech, at both word and sentence levels, totaling over one hour of intracranial neural recordings related to speech production. This paper also presents a baseline model as the reference model for future studies, at the same time, ensuring the integrity of the dataset. The diversity of tasks and the substantial data volume provide a valuable resource for developing advanced algorithms for speech decoding, thereby advancing BCI research for spoken communication.

Original languageEnglish
Article number657
JournalScientific Data
Volume12
Issue number1
DOIs
StatePublished - Dec 2025
Externally publishedYes

Fingerprint

Dive into the research topics of 'VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in Tonal Language'. Together they form a unique fingerprint.

Cite this