Designing an experiment that can fail
Hypotheses, counterbalancing, jitter, triggers, sample size, pre-registration, consent, and the pilot. The parts of a human experiment that decide whether the result means anything, before any data exist.
You are skimming: the title, the first figure, and the short version. Switch to Read in the header for the full page, or Deep to open every deep dive.
An experiment is a question arranged so that the world can answer no. Most of what goes wrong in student experiments goes wrong before the first participant arrives: the question was vague, the conditions differed in more than one way, the timing was periodic, the sample was too small, and the analysis was chosen after seeing the data. Each of those has a fix that costs an hour. Here they are.
One question, stated so it can be wrong
“Does face processing produce an N170” is not a question. “Mean amplitude at P8 from 150 to 190 ms is more negative for faces than for houses, by at least 1 µV, in 16 participants” is. It names the measure, the window, the channel, the direction, the size, and the sample. Write it down before collecting anything. If the result comes out the other way, you have learned something; if you wrote nothing down, you will find a way to explain whatever happened.
Conditions that differ in one thing
Faces versus houses differ in what they depict. They also differ, unless you work at it, in brightness, contrast, spatial frequency, size, and how interesting they are. Any of those could produce a difference at 170 ms. Match the images on everything you can measure, and randomize what you cannot. This is the whole art of stimulus design and it is where most of the time goes.
Randomize and counterbalance
Trial order random, so that nothing about position in the session (fatigue, learning, electrode drying) lines up with a condition. Response mappings (left key for faces) swapped across participants, so a motor difference does not masquerade as a perceptual one. Block order counterbalanced when blocks are unavoidable.
Jitter the timing
A fixed inter-stimulus interval lets the participant predict the next stimulus, which produces its own ERP (the contingent negative variation) and lets rhythms entrain. It also lets any periodic artifact (heartbeat, alpha, mains) align with the stimulus and survive averaging. Draw each interval from a range: 1000 to 1500 ms, uniformly.
Triggers you can trust
Markers pushed at the flip, a photodiode check once, LSL for everything. The infrastructure page covers it. A 30 ms systematic offset is invisible in a single study and a disaster when you compare latencies to the literature.
Your pilot participant shows a beautiful effect. Should you run the full study with the same analysis?
Use the pilot to fix the procedure and fix the analysis in writing, then run the study without the pilot’s data. The pilot’s job is to find broken timing, confusing instructions, and an analysis window that makes sense. Once you have looked at pilot data and chosen windows, that data cannot be part of the test, or you have chosen the analysis to fit it.
Sample size
Effect sizes for classic ERPs are known: the N170 face effect is large (a Cohen’s d around 1), so twelve to sixteen participants give good power. The P300 oddball effect is larger still. New effects have unknown sizes, and the honest move is a pilot to estimate it and a power calculation to set N. A study with six participants that finds nothing has found nothing; a study with six that finds something has probably found noise.
Pre-register
Write the hypothesis, conditions, exclusion rules, measure, window, channel, and test, in a document with a date, before collecting the real data. Put it in your repository. It costs an hour and it converts “we found” into “we predicted and found,” which is a different kind of claim. Labs increasingly require it; do it before they ask.
Consent, done as a conversation
The IRBInstitutional review board (IRB)The committee that must approve any research on people before it starts, to protect the participants. Glossary entry approves a consent form; the consent is the conversation. Explain what the electrodes do and do not do (record, never stimulate), how long it takes, that they can stop any time without explanation, what happens to the data, and who to contact. Ask if they have questions. Watch for the person who is uneasy and has not said so; offer the exit again. Then the signature. A participant who understood is a participant whose data you can use with a clear conscience.
The pilot
Run yourself, then one friend, through the whole thing, timing included. Check the marker offsets, the epoch counts, the number of rejected trials, and whether the instructions were understood. Fix everything. Then freeze the protocol and begin.
Deep dive Exclusion rules, written first 2 min
Which participants will you exclude (too few clean trials, did not follow instructions, technical failure) and which trials (amplitude threshold, response too fast or too slow)? Decide the thresholds before the data exist. Report how many were excluded and why. An exclusion rule invented after looking at results is a way of choosing the result.
Deep dive Within-subject versus between-subject designs 2 min
Every participant seeing every condition (within-subject) is far more powerful, because the huge differences between people cancel in the comparison. Almost every EEG study is within-subject for this reason. Between-subject designs (one group gets feedback, another does not) need several times the participants and are unavoidable only when the manipulation cannot be undone.
Explain what this page was about to your roommate in three sentences. No jargon they would not know.