In the previous part
In Part 6 I used an electrode at O1 to record a repeatable 10 Hz rhythm that increased whenever I closed my eyes. This established that the homemade ADS1299 recorder could preserve a condition-dependent EEG signal. It did not establish that the smaller changes produced by left- and right-hand motor imagery would be separable.
The first motor-imagery recording
After the alpha test, I moved the electrodes to C3, C4, and Cz. These were the 3 positions selected from the public-dataset benchmark in Part 4, and they were now connected to the recorder that had passed the physiological signal test. The remaining question was whether I could produce a motor-imagery response that was useful for steering.
I recorded 20 balanced left- and right-hand trials at 250 Hz. After removing 2 artifact trials, the conventional channel-by-channel analysis left one plausible lead: during right-hand imagery, low-beta power at C3 changed in the expected contralateral direction. The effect was not repeated clearly for left-hand imagery, and it disappeared in the bipolar sensitivity analysis - suggesting a false positive. The textbook 8-13 Hz lateralized ERD result was negative.
A 3-channel CSP check did not solve the problem. Depending on the frequency band, it correctly classified between 27.8% and 61.1% of the 18 retained trials, with every permutation test remaining compatible with chance (50%). There was measurable synchronization activity around 10 Hz, but it did not change in the stable, lateralized, pattern required to distinguish the imagined hand.
Why use CSP if textbook ERD failed
The 2 analyses do not test the same claim. ERD is a predefined contrast: 8-13 Hz power at C3 against C4, decreasing contralateral to the imagined hand. A negative result rejects that one contrast and nothing wider. CSP is not told where to look - it learns channel weightings from the class covariance matrices, so any other consistent spatial pattern can emerge as a usable feature even when the lateralized contrast is flat.
Note: the argument only runs in one direction. A weak ERD result still leaves the spatial model worth trying, but CSP failing on the same 18 trials rules out the predefined contrast and the emergent one at once.
Why I added Oz
The clear 10 Hz synchronization created another possible explanation for the failed classification. Alpha modulation is strongest over the posterior part of the head, but broadly distributed alpha can also appear in C3, C4, and Cz. My hypothesis was that this shared modulation was overshadowing the smaller left-right motor-imagery difference.
I therefore added Oz as a 4th electrode. It was not intended to control the car and I did not expect Oz by itself to identify the imagined hand. It was a control measurement of posterior alpha that CSP could combine with C3, C4, and Cz. If a component appeared similarly at the posterior and central electrodes, a spatial filter could reduce it while retaining a difference that changed across the motor channels.
This was only a proposed mechanism. CSP learns channel combinations that separate labelled classes; it does not know which component is motor activity and which is alpha. The test therefore had to compare the same retained trials with and without Oz, and also check whether Oz alone contained class information.
Four-channel protocol and analysis
I recorded C3, C4, Cz, and Oz at 250 Hz. Each of the 20 trials contained 3 seconds of rest, a 0.5-1.0 second cue with randomized duration, and 5 seconds of imagery while I kept looking at the same fixation point. The final 2.5 seconds of rest formed the baseline, and the predefined imagery window covered 0.5-4.5 seconds after imagery onset.
The filter-bank CSP model used 5 bands from 8 to 30 Hz, 20% covariance shrinkage, and the first and last CSP log-variance feature from every band. CSP, feature scaling, and the class centroids were refitted inside every validation fold so the held-out trial could not influence its own spatial filters.
A promising first session
The first 4-channel session retained 18 balanced trials. Filter-bank CSP using C3/C4/Cz/Oz classified 14/18 correctly: 77.8% balanced accuracy with a 500-permutation p-value of .032. The matched C3/C4/Cz model classified only 8/18, or 44.4%.
Oz alone reached 55.6%, and adding Oz as another ordinary band-power value also reached 55.6%. The improvement appeared only when CSP could use Oz jointly with the 3 motor channels. Trial-to-trial 8-13 Hz changes at Oz also shared approximately 37-39% of their variance with the changes at C3, C4, and Cz.
This pattern was consistent with the original hypothesis: CSP may have used Oz to reduce shared posterior or global alpha while preserving class-related spatial covariance. It did not prove that mechanism, and the classical C3/C4 ERD result in the same recording was still weak. The classifier was using multiband spatial covariance, not a clear textbook difference in motor-band amplitude.
I also inspected several imagery intervals. The best was 2-4 seconds, where the 4-channel result reached 83.3%. Because I selected that interval after looking at the session, it was exploratory. The unbiased result remained 77.8% in the predefined 0.5-4.5 second window.
One session with 18 clean trials was enough to produce a hypothesis, but not enough to call the system reliable. I froze the montage, filter bank, artifact rules, CSP procedure, and classifier before collecting more data. The practical target was at least 70% balanced accuracy on a session that had not contributed any labels to training or adjustment.
The first replication failed
Next day, the second recording matched the same C3/C4/Cz/Oz protocol and retained 19/20 trials. I trained the frozen model on session 1 and applied it unchanged to session 2. In the prospectively selected 2-4 second window, the 4-channel model reached 55.0% balanced accuracy. The matched 3-channel version reached 48.3%.
The 55.0% figure was also less useful than it first appeared. The model predicted 18 of the 19 trials as left and recalled only 10% of the right-hand trials. It had transferred mainly as a single-class decision rule, not as a usable steering signal.
A session offset I have not tested yet
One candidate explanation for the single-class behavior is a session-level gain change. The features are log-variance, so a gain common to all channels multiplies every variance by the same factor and adds the same constant to every feature. The session 2 cloud translates, the session 1 centroids stay where they were, and every trial lands on the same side of the boundary.
The fix for that would be to recenter session 2 on its own unlabelled feature mean and apply the unchanged centroids. Refitting the centroids on session 2 labels would also work, but it would turn the transfer test into calibration. TODO for next analyses.
I then combined sessions 1 and 2 to see whether a personalized model could learn both known recording conditions. In the primary 2-4 second pooled test, the 4-channel model remained at 51.5%. The full 0.5-4.5 second model reached 64.9% in one fixed split, but across 100 session-balanced fold assignments it averaged 62.4%, and only 13/100 assignments reached 70%.
This suggested that the full imagery window carried a weak and more repeatable signal than the late window, but it did not establish generalization. I fitted a combined sessions 1+2 model and selected the full 0.5-4.5 second interval as the prospective hypothesis for an untouched third session. The earlier 2-4 second model remained as a legacy comparison.
Untouched session 3
Session 3 contained 51,050 recorded samples with 0.098% missing. The frozen artifact rule retained 18/20 trials, balanced at 9 left and 9 right. I trained on sessions 1+2 and applied both models to session 3 without using any session 3 labels for training or adjustment.
| Model | 0.5-4.5 s accuracy | p | 2-4 s accuracy | p |
|---|---|---|---|---|
| C3/C4/Cz | 61.1% (11/18) | .246 | 61.1% | .228 |
| C3/C4/Cz/Oz | 66.7% (12/18) | .064 | 66.7% | .132 |
The 4-channel model classified 12/18 trials correctly. This was 5.6 percentage points better than the matched 3-channel model, but it did not meet the optimal 70% target. Pooling the first 2 sessions had not solved unseen-session generalization.
That is the main result of this phase. The first session’s 77.8% was a promising pilot, while 66.7% is the prospectively tested estimate on the untouched third session. The permutation p-value of .064 did not pass .05, but the practical decision did not depend on that threshold: the classifier missed the 70% target that had been chosen before the recording was opened.
Pairwise transfer into session 3
These comparisons show which earlier recording session 3 resembled, but they are diagnostic rather than replacements for the prespecified sessions 1+2 model.
| Training data | C3/C4/Cz | C3/C4/Cz/Oz |
|---|---|---|
| Session 1 | 72.2% (13/18) | 66.7% (12/18) |
| Session 2 | 50.0% (9/18) | 55.6% (10/18) |
Session 3 internal decoding
Leave-one-out decoding within session 3 was secondary to the untouched transfer result because it trained on other trials from the same recording.
| Model | 0.5-4.5 s | p | 2-4 s | p |
|---|---|---|---|---|
| C3/C4/Cz | 66.7% | .126 | 66.7% | .120 |
| C3/C4/Cz/Oz | 55.6% | .407 | 61.1% | .255 |
What remained consistent across sessions
I repeated the same full-window transfer analysis with each session held out in turn. This was not a new prospective test for sessions 1 and 2, because their data had already been inspected, but it showed whether the Oz difference was confined to one convenient train-test direction.
| Untouched test session | Train on | C3/C4/Cz | C3/C4/Cz/Oz |
|---|---|---|---|
| Session 1 | Sessions 2+3 | 44.4% (p=.764) | 61.1% (p=.120) |
| Session 2 | Sessions 1+3 | 52.8% (p=.407) | 59.4% (p=.078) |
| Session 3 | Sessions 1+2 | 61.1% (p=.246) | 66.7% (p=.064) |
The data supports keeping the extra electrode because it repeatedly helped the spatial model at the cost of one additional contact.
Pooled calibration was not replication
After completing the untouched test, I pooled all 3 sessions. This produced 55 clean trials, with 27 left and 28 right. Every 5-fold training set contained trials from all 3 sessions, and so did each held-out fold.
The full-window 4-channel model reached 63.6% in the fixed split, compared with 54.4% for 3 channels. Across 100 session-balanced fold assignments, the 4-channel model averaged 67.4%, with a median of 67.3% and a central 95% split range of 58.2-74.5%. It reached 70% in 30/100 assignments. A repeated-fold permutation test gave 67.4% with p=.005.
Fixed pooled comparison
| Model | 0.5-4.5 s | p | 2-4 s | p |
|---|---|---|---|---|
| C3/C4/Cz | 54.4% (30/55) | .311 | 47.2% | .627 |
| C3/C4/Cz/Oz | 63.6% (35/55) | .008 | 54.6% | .144 |
These numbers show that the pooled data contain class-related structure that the model can use when every recording session is represented during training. They are calibration estimates, not independent confirmation. A held-out trial from session 3 can still be trained on other session 3 trials, so the model has already seen that session’s distribution.
Preparing the model for the next test
I fitted final 3- and 4-channel models for both imagery windows using all 55 clean trials and saved their CSP filters, feature normalization, and class centroids. Their training-set accuracy is not a performance estimate. Their only valid next test is another untouched recording.
I also replayed the raw session 3 file through the same sequence buffer and quality gate intended for live mode. It classified 19/20 trials, rejected trial 2, and reproduced all 18 stored predictions for the offline-clean trials exactly. A sample-block replay also matched every offline prediction. This verified that the inference implementation agreed with the analysis code and could be used during live classification.
Decision
I will retain Oz for the next sessions. Its proposed role as a measurement of posterior alpha remains a hypothesis, but the controlled comparisons showed a consistent cross-session advantage and one additional electrode is still a reasonable hardware cost.
Engineering options
This project is engineering-first, so a missed statistical target is cheap - I am not defending a claim about physiology. In the end, live control also has degrees of freedom the offline test does not: the car can accumulate several classifications and act only when a confidence interval excludes chance, and the threshold can stay conservative, since standing still costs less than driving the wrong way. If that is not enough, a short calibration window at the start of each session can recompute the centroids for that recording.
Failing to reproduce the textbook ERD is a disappointment, but it was never the objective. The goal was a model that can read my mind well enough to steer a car, and for that, any method works, if it works.