- CNN-derived tuning subspaces from single-neuron auditory cortex data yield a sparse, efficient coding framework for natural sounds (Nature Neuroscience).1
- Subspaces distinguish functional properties between neuronal subtypes and are directly relevant to neural decoding and interpretable models for sensory BCIs.1
- Authors Jereme C. Wingert, Satyabrata Parida, Sam V. Norman-Haignere, and Stephen V. David (Oregon Health and Science University / University of Rochester) recorded from 2,874 single units across 67 sites in auditory cortex (A1 and PEG) of 4 ferrets using Neuropixels probes and 64-channel silicon arrays.1
- A two-block population CNN was trained on 484,375 natural sound segments drawn from Audioset and PSE corpora; the subspace model retained 95.4% of CNN variance on average (median prediction correlation r = 0.585 vs. CNN r = 0.600, P = 1.0 × 10⁻⁷), confirming functional equivalence.1
- The derivative spectrotemporal receptive field (dSTRF) method extracts each neuron’s linear tuning subspace by computing the Jacobian of CNN output relative to the input spectrogram; on average 11 principal components explained 81% of dSTRF variance for A1 neurons (range 3–16).1
- Local neural populations sparsely tile a shared stimulus subspace: within-site subspace similarity index (SSI) was 0.55 vs. 0.35 between sites (P = 2.5 × 10⁻¹¹), while high-activity SSRFs overlapped only ~4% more than chance and 71% less than full overlap.1
- Narrow-spiking (putative inhibitory) neurons covered larger stimulus subspace areas than regular-spiking (excitatory) neurons (mean SSRF area 0.115 vs. 0.076, P = 4.4 × 10⁻⁵) and showed stronger tuning symmetry differences in layer 4 (χ² = 62.0, P = 3.4 × 10⁻¹⁵), implicating cell-type-specific roles in cortical circuit computations.1
- The framework provides a principled method for interpreting deep-learning encoding models and is applicable to designing auditory sensory BCIs and cortical stimulation strategies targeting specific local circuit elements.1
- The CNN outperformed the classical linear–nonlinear (LN) model (median r = 0.600 vs. 0.416, P = 1.8 × 10⁻¹⁴), and the subspace approach preserved most of that gain while adding interpretability, bridging the gap between predictive accuracy and mechanistic understanding.1