Two binocular 2AFC arms, within participant — the structure, the three
candidate versions, the stimuli, the novel obscuring mechanism, the best-practice spine, and the
working agreement (roles & milestones).
Contract for discussion · 2026-07-23
Note on § numbering: section numbers are global across the
single design-contract source document, not per-page. This page carries §2–8; the gap at
§6 (Power) is intentional — it lives on the Power page. Numbers are
not renumbered per page so cross-references stay stable.
2 · Experimental structure — two arms, within participant
Each participant completes two arms of a forced binary choice task. Presentation is
binocular (both eyes, no dichoptic rivalry) — despite “CFS” in the
project name, Part 1 is deliberately the pre-CFS validation baseline: the flashing
Mondrians are a graded masking/degradation of a binocularly-viewed face, not interocular
suppression. Both arms yield choice + RT on every trial → DDM.
Two arms, one participant, one pipeline. Both the random-dot-motion arm and the masked-face arm feed the same EZ+HDDM pipeline, yielding the identical parameter set v, a, t₀, z on every trial.
The scientific tests
T1 — graded within-subject drop in v across obscuring (does obscuring lower drift?).
T2 — parameter selectivity (is the effect drift-specific — v moves, a/t₀ don’t?).
T3 — cross-arm equivalence (TOST) — is the coherence→v slope quantitatively equivalent to the obscuring→v slope in the same person? This is the claim that certifies the pipeline as a reusable tool.
T4 — individual-differences convergent validity — do motion-drift and face-drift covary across people?
🎧 Contextual deep-dive
b-CFS framing: why Part 1 is deliberately not yet CFS
A spoken explainer on why the project name says “bCFS” but Part 1 uses plain binocular
graded masking, not interocular suppression — and what breaking-CFS (b-CFS) would add in a
later part.
The whole design in one figure. One participant spans both arms; each arm has a graded difficulty ladder feeding a central drift-diffusion model that outputs the four fitted parameters. The four tests (T1–T4) are annotated, with T3 linking the two arms’ drift slopes. Standalone file: assets/flagship-design.svg.
3 · The three design versions
Stimulus set for all versions: RaFD (Radboud Faces Database) — on disk, 8,040
images, 67 IDs × 8 emotions × 5 camera angles × 3 gaze directions
(left / frontal / right). Frontal camera angle used throughout.
V1 · Gaze
Left vs. right gaze direction
Literature anchor: none — no published gaze-discrimination
drift estimate exists. Elegant structural parallel to RDM motion-direction.
Highest novelty; pilot-dependent.
V2 · Emotion · flagship
Angry vs. happy
Literature anchor: Williams 2023PDF (masking→v on RaFD) — the closest published anchor the field offers, though its occlusion geometry differs from our graded noise mask (see Power caveats).
Best-anchored available; recommended flagship.
V3 · Factorial
Gaze × emotion (2×2)
Literature anchor: partial (emotion anchored, gaze not).
Richest, most publishable single study; largest N. PI leans here.
Flagship recommendation: run V2 (emotion) as the flagship and a
~10-person V1 (gaze) pilot in parallel to obtain the missing gaze drift anchor; keep V3
costed and ready as the ambitious variant.
Example stimuli — the manipulated dimensions
Gaze LEFT
Gaze RIGHT
ANGRY
HAPPY
Schematic of the two manipulated dimensions — gaze (columns) and emotion (rows).
Radboud Faces Database (RaFD), Langner et al. 2010PDF — academic use, kept local. Frontal camera (Rafd090). These are the real manipulated dimensions the task discriminates.
A separate in-house synthetic face set (assets/faces/ex_*.png) exists only as an illustrative pipeline-test set — the study stimuli are RaFD.
4 · Stimuli, apparatus, task — and where each choice comes from
RaFD faces — used by our anchor paper (Williams 2023, Study 1), so it is not
novel to the corpus; using it is justified on RaFD’s own validation (Langner et al. 2010PDF) and
norming and its 3 gaze directions (needed for V1/V3), not on face-DDM precedent. Low-level confounds (luminance, contrast, spatial frequency)
controlled per Purcell 1996, Becker 2011, Calvo 2016.
Graded binocular Mondrian masking0/19 precedent
— only 1 distinct precedent masks at all (Williams 2023; Hartmann 2021 is its preprint — not an
independent second study, so the corpus is effectively ≤18 distinct studies), and it is binary. A continuously
graded mask is novel — justified from the RDM/CFS graded-noise tradition (Newsome &
Paré 1988 for per-subject coherence/noise-dot calibration; Tsuchiya & Koch 2005 for 10 Hz Mondrian dynamics), not
from a face-DDM precedent. Mask levels titrate discriminability, mirroring coherence in Arm A.
Adaptive per-subject calibration — QUEST/Ψ0/19 precedent — not one corpus paper uses
adaptive staircase calibration of face-task difficulty. Imported wholesale from psychophysics
(Watson & Pelli 1983PDF; Kontsevich & Tyler 1999) to equate difficulty across participants and
across the two arms. A genuine methodological contribution for face-DDM.
Trials: target ≥100 correct trials/cell for the EZ baseline; ~20–40/cell
adequate under HDDM shrinkage (Lerche, Voss & Nagler 2017PDF; Wiecki 2013PDF). Williams ran
~18/emotion×mask cell (sparse) — our design exceeds this.
Trial timeline. Fixation → stimulus (motion coherence or a graded face mask) → response, one row per arm. Both arms share the same three-stage structure and record choice + RT.
🎧 Contextual deep-dive
The RaFD face stimulus set
A spoken walkthrough of the Radboud Faces Database — its validation, the frontal/gaze/emotion
subsetting this design needs, and why using it is justified on its own norming merits — RaFD is
used by our anchor paper (Williams 2023), so it is not novel to the corpus (see Comparison §3.5(a)).
The obscuring mechanism — a novel contribution
The single largest methodological gap the golden-standard comparison surfaces (see
Comparison) is that no paper in the 19-paper corpus grades
face-masking severity continuously — occlusion/masking has only 1 distinct precedent
(Williams 2023; Hartmann 2021 is its preprint, not an independent study), and it is binary (masked vs.
unmasked). A continuously graded occlusion/masking manipulation (percentage of
face area obscured, or a stepped noise-mask series) has zero precedent
in the reviewed literature. This design's graded binocular flashing-Mondrian mask — mirroring the
coherence ladder in the RDM arm — must be justified from the RDM/b-CFS graded-noise literature
(with Newsome & Paré 1988 cited specifically for per-subject coherence/noise-dot
calibration) rather than from a face-specific precedent, since none exists.
Matched graded difficulty. The motion-coherence ladder (Arm A) and the graded flashing-Mondrian mask ladder (Arm B), matched in number of levels and titrated per subject — coherence ↔ mask.
🎧 Contextual deep-dive
RDM & DDM best practices behind this design
A spoken walkthrough of the RDM/psychophysics tradition this design borrows from: coherence-based
graded difficulty, adaptive per-subject calibration, and the EZ+HDDM estimator pairing.
5 · Analysis pipeline (open-source)
Per-subject baseline: EZ-diffusion (Wagenmakers 2007PDF) — fast, closed-form
0/19 precedent; chosen from best-practices, not corpus.
Confirmatory:HDDM (Wiecki, Sofer & Frank 2013PDF) —
hierarchical Bayesian; the plurality choice among corpus papers that name a method (4/6). Required
because EZ recovers a/t₀ poorly (see Power §1 caveat).
Parameter recovery validated by simulation before fitting real data
(r_v=0.99; r_a, r_t₀≈0.55–0.59 under EZ — motivates
HDDM for the selectivity claim).
Equivalence: TOST on standardized difficulty→v slopes across
arms, pre-specified margin (see Power §3.3).
Everything pre-registered; task code, analysis, and (permitted) data open.
The drift-diffusion model. A noisy evidence path starts at z, drifts at rate v across the gap a between two boundaries after a non-decision time t₀, producing a choice and a response time.
7 · Best-practice spine (golden standards)
Pre-registration of hypotheses, N, exclusion rules, and the equivalence margin (before data).
Adaptive per-subject calibration (QUEST/Ψ) to equate difficulty across arms — novel for face-DDM.
Graded difficulty in both arms (coherence ↔ mask), matched in number of levels.
EZ (screen) + HDDM (confirm); report parameter recovery.
Low-level confound controls on RaFD stimuli; counterbalancing of arm order, response mapping.
Distributional RT analysis — never mean-RT (Tipples 2023PDF: conclusions flip with outlier/model choice).
Open task, analysis, and data; local-first tooling.
EZ + HDDM pipeline; parameter-recovery checks; QUEST/Ψ calibration code; power re-estimation
from pilot variance; TOST equivalence; figures.
Milestones (relative to 2026-07-23)
#
Milestone
Owner
Target
M0
This contract approved
all
today
M1
RaFD subsets + graded Mondrian masks built
techs
+2 wk
M2
PsychoPy two-arm task (RDM + face) running
techs + analysts
+4 wk
M3
V1 gaze + V2 emotion pilot (~10) → real drift anchors, mask-level slopes
all
+6 wk
M4
Pre-registration (final N, margin, exclusions) from pilot variance
PI + analysts
+8 wk
M5
Main data collection (flagship V2, N≈50)
techs
+8–16 wk
M6
DDM analysis, equivalence, writeup
analysts + PI
+16–20 wk
Deliverables of this KB effort: the contract; this interactive KB website; the
illustrated slide deck (main acts + annex); the golden-standard comparison; the grounded power analysis.