1 / 51
First draft — preliminary contract for discussion
Scientific project contract meeting

RDMxbCFS (Neuronenas) — Part 1

Within-participant DDM: random-dot motion vs masked-face decisions

Date: 2026-07-23
PI: L. A. Barradas-Chacón
Random-Dot-Motion × binocular Continuous Flash Suppression (full project name) · Part 1 manipulation: graded visual masking (pre-CFS) · Part 1 of N
Press or Space to advance · to go back
§1 · Thesis

The one-line thesis

Goal of Part 1: stand up and validate our in-house DDM data pipeline by measuring, within the same participant, whether perceptual decisions about random-dot motion and about masked RaFD faces are governed by comparable drift-diffusion signatures.

Two very different-looking 2AFC tasks — one classically "low-level" (motion), one "high-level" (faces) — run through the same pipeline, the same participant, the same parameters (v, a, t₀, z). If the drift-diffusion signatures line up, the pipeline is validated as a reusable, domain-general tool.

Roadmap

The argument in five acts

A map of the whole deck — from the design, through stimuli and the model, to the evidence and the power decision.

ACT I

Design

  • One-line thesis
  • Two binocular 2AFC arms
  • Versions V1 / V2 / V3
ACT II

Stimuli & obscuring

  • RaFD faces
  • Graded Mondrian mask
  • Matched difficulty ladder
ACT III

DDM & four tests

  • DDM primer
  • Flagship design figure
  • Tests T1–T4
ACT IV

Evidence & novelty

  • 19-paper corpus
  • Zero-precedent tally
  • Retrieved effect sizes
ACT V

Power & decisions

  • Four claim tiers
  • Decisions D1–D5 on the table
  • Simulated power & N

The decision for the room lands in Act V: the five explicit choices D1–D5 — design version (§3), obscuring mechanism (§4), power anchor, claim tier and equivalence margin (§6) — each argued with its cost, and backed by the simulated power distributions. Everything before it is the evidence for that call.

§1 · Honest framing

Reproduction vs. novel — stated plainly

The face arm is not a strict reproduction. Across the 19-paper face-2AFC-DDM corpus, no study combines our elements (RaFD + graded masking + strict 2AFC + full v/a/t₀/z fit) on either axis.

Reproducible

What is established

  • The RDM arm — Palmer, Huk & Shadlen 2005, well-established proportional-rate diffusion model
  • The direction of the face effect — masking lowers drift (Williams et al. 2023)
Novel

What is new

  • The combination: RaFD + graded masking + strict 2AFC + full v/a/t₀/z fit — 0/19 precedent for this exact bundle, on either axis

Part 1 therefore reproduces established directions and methods while being, in combination, a novel design — which is precisely where the publishable contribution sits.

§2 · Experimental structure

Two binocular 2AFC arms, within participant

Presentation is binocular (both eyes, no dichoptic rivalry) — Part 1 is deliberately the pre-CFS validation baseline.

Arm A — RDM

Binary choice
left vs right motion
Difficulty ladder
~5 coherence levels {3.2, 6.4, 12.8, 25.6, 51.2}%
DDM prediction
drift v ↑ with coherence (µ′ = k·x, k ~ N(21,6))
Anchor
Palmer, Huk & Shadlen 2005

Arm B — Faces (RaFD)

Binary choice
see design versions V1/V2/V3
Difficulty ladder
~5 graded binocular flashing-Mondrian mask levels
DDM prediction
drift v ↓ with obscuring (Δv ≈ −0.38 … −1.12)
Anchor
Williams et al. 2023 (RaFD, DDM)

Both arms yield choice + RT on every trial → DDM. Dependent measures are identical: v, a, t₀, z.

§4 · Arm A detail

Random-dot motion — the coherence ladder

Stimulus & task

  • Coherence ladder: {3.2, 6.4, 12.8, 25.6, 51.2}%
  • Left vs right motion, forced 2AFC
  • Built on the PI's own open PsychoPy RDK (abcsds/RDM, 2024)

Drift model

µ′ = k·x

Drift scales linearly with coherence proportion x; sensitivity modeled as k ~ N(21,6) between subjects (range ≈ 9–28) across observers/experiments (Palmer, Huk & Shadlen 2005, Table 2).

Anchor: Palmer, Huk & Shadlen (2005) — normalized bound A′ ≈ 0.71, non-decision time tR ≈ 300–420 ms, threshold ratio ≈ 3.0–3.8, validating the proportional-rate model across 5 experiments.

§4 · Arm B detail

Masked faces — graded binocular Mondrian obscuring

The manipulation

  • ~5 graded binocular flashing-Mondrian mask levels, titrating discriminability — mirrors coherence in Arm A
  • Mask levels degrade the RaFD face image progressively, not a single fixed occlusion

Why "pre-CFS"

Despite "CFS" in the project name, Part 1 is deliberately the pre-CFS validation baseline: the flashing Mondrians are a graded masking/degradation of a binocularly-viewed face, not interocular suppression.

No rivalry, both eyes = pre-CFS baseline.

Anchor: Williams et al. 2023 (Sci. Reports) — face-mask occlusion lowers drift rate v; two independent pre-registered studies (N=228 RaFD, N=264 RADIATE).

§4 · Task

One trial, both arms — the timeline

Fixation → stimulus (coherence / mask) → response. Identical structure across arms; every trial yields choice + RT.

Arm A — Random-dot motion Fixation ~0.5 s Motion stimulus until response coherence 3.2–51.2% Response LEFT RIGHT choice + RT recorded Arm B — Masked RaFD face Fixation ~0.5 s Face + Mondrian mask until response 5 graded mask levels Response ANGRY / HAPPY or gaze L / R choice + RT recorded time Both arms yield choice + RT on every trial → drift-diffusion model (v, a, t₀, z)
§3 · Design versions

Three design versions — flagship = V2 + V1 pilot

VersionFace binary choiceLiterature anchorNotes
V1 — Gazeleft vs right gaze directionnone no published gaze-discrimination drift estimateElegant structural parallel to RDM motion-direction. Highest novelty; pilot-dependent.
V2 — Emotionangry vs happyWilliams 2023 masking→v on RaFDBest-anchored; recommended flagship.
V3 — Factorialgaze × emotion (2×2)partial emotion anchored, gaze notRichest, most publishable single study; largest N. PI leans here.

Flagship recommendation: run V2 (emotion) as the flagship and a ~10-person V1 (gaze) pilot in parallel to obtain the missing gaze drift anchor; keep V3 costed and ready as the ambitious variant.

§4 · Stimuli

RaFD, not FACES

RaFD (chosen)

  • Radboud Faces Database — 8040 images on disk
  • 67 IDs (adults + children) × 8 expressions (anger, disgust, fear, happiness, sadness, surprise, contempt, neutral)
  • × 5 camera angles × 3 gaze directions (left / frontal / right)
  • Frontal camera angle used for this study
  • Free for academic use (Langner et al. 2010)

FACES (Ebner 2010) — not chosen

  • 171 models across 3 age groups
  • 6 expressions (neutrality, sadness, disgust, fear, anger, happiness)
  • Frontal shots only — no gaze-direction variation
  • Would rule out V1 (gaze) and V3 (factorial) entirely

RaFD is not novel to this corpus — our anchor paper (Williams 2023, Study 1) uses it. The choice is justified on RaFD's own validation (Langner et al. 2010) and its 3 gaze directions (needed for V1/V3), not as a zero-precedent novelty, with low-level confounds (luminance, contrast, spatial frequency) controlled per Purcell 1996, Becker 2011, Calvo 2016.

§3–4 · Stimuli

Real RaFD stimuli — the two manipulated dimensions

Gaze LEFT
Gaze RIGHT
ANGRY
RaFD angry face, gaze leftRaFD angry face, gaze right
HAPPY
RaFD happy face, gaze leftRaFD happy face, gaze right
GAZE (V1: left vs right) Gaze LEFT Gaze RIGHT EMOTION (V2: angry vs happy) ANGRY HAPPY V3 = the full 2 × 2 factorial (gaze × emotion) Schematic depiction of the two manipulated dimensions — the real stimuli are photographic RaFD faces.
§4 · Stimuli

The obscuring mechanism — graded Mondrian masking

Novel · 0/19 precedent for graded severity

Only 1 distinct corpus study masks at all (Williams 2023; Hartmann 2021 is its preprint, not an independent second study), and it is binary (masked vs. unmasked). A continuously graded mask — mirroring the coherence ladder in Arm A — has zero precedent in the face-DDM literature.

Justified from

  • Newsome & Paré 1988 — coherence/noise-dot fraction logic
  • Tsuchiya & Koch 2005 — 10 Hz Mondrian dynamics

Not justified from

  • Any face-DDM precedent — none exists
  • The RDM/CFS graded-noise tradition is the source, not the face literature

Mask levels titrate discriminability, mirroring coherence in Arm A — imported wholesale from the RDM/CFS tradition rather than replicated from a face-DDM template.

§4 · Parallel difficulty

Matched graded difficulty — coherence ↔ mask

Parallel graded difficulty — coherence ↔ mask Arm A: motion coherence Arm B: face-mask level 51.2% 25.6% 12.8% 6.4% 3.2% easiest hardest drift v ↑ with coherence (µ′ = k·x, k ~ N(21,6)) Palmer, Huk & Shadlen 2005 level 1 least obscured level 2 level 3 level 4 level 5 most obscured drift v ↓ with obscuring (Δv ≈ −0.38 … −1.12) Williams et al. 2023 EASIER ↑ HARDER ↓ matched in number of levels titrated per-subject (QUEST / Ψ)
§2 · Scientific tests

The four scientific tests, T1–T4

T1

Graded within-subject drop in v across obscuring — does obscuring lower drift?

T2

Parameter selectivity — is the effect drift-specific (v moves, a/t₀ don't)?

T3 ★

Cross-arm equivalence (TOST) — is the coherence→v slope quantitatively equivalent to the obscuring→v slope, in the same person?

T4

Individual-differences convergent validity — do motion-drift and face-drift covary across people?

T3 is the claim that certifies the pipeline as a reusable tool — quantitative equivalence between a "low-level" and a "high-level" decision would mean the DDM is reading out a domain-general accumulation process.

Background

DDM 60-second primer

The drift-diffusion model — accumulation to a bound a noisy evidence signal races from a start point to one of two boundaries upper boundary → choice A lower boundary → choice B a t₀ z v response time (RT) = t₀ + decision time v drift rate evidence quality / how fast good evidence arrives a boundary separation caution — the speed vs accuracy trade-off t₀ non-decision time encoding + motor time, outside accumulation z start point prior bias — where accumulation begins

Why choice + RT, not mean RT: a single mean-RT number collapses four dissociable processes into one. Tipples (2023): an emotion-DDM conclusion can flip with outlier-removal / RT-model choices — never trust mean RT alone.

The centerpiece

Flagship — the entire design at a glance

The whole design in one figure One participant · two 2AFC arms · one drift-diffusion readout (v, a, t₀, z) · four tests T1–T4 ● ONE PARTICIPANT both arms, within-subject, same session ARM A — Random-dot motion ◀ left vs right ▶ (2AFC) ARM B — Masked RaFD face angry / happy (V2) · gaze L / R (V1) difficulty ladder — motion coherence (%) 3.2 6.4 12.8 25.6 51.2 harder easier → difficulty ladder — 5 graded mask levels L5 L4 L3 L2 L1 harder easier → coherence → v (slope A) v difficulty PHS05: µ′=k·x, k~N(21,6) T1 obscuring → v (slope B) v difficulty Williams 2023: Δv −0.38…−1.12 T1 DRIFT-DIFFUSION MODEL upper bound → choice A lower bound → choice B a t₀ z v response time (RT) = t₀ + decision T3 slope A ≡ slope B (TOST) choice + RT fit v,a,t₀,z v drift rate evidence quality — how fast good evidence arrives a boundary sep. caution / threshold — speed–accuracy trade-off t₀ non-decision encoding + motor time, outside accumulation z start point prior bias — where begins accumulation THE FOUR TESTS T1 Graded slope Does obscuring lower drift? v changes monotonically across the ladder — each arm. T2 Selectivity Is it drift-specific? v moves while a and t₀ stay put. T3 Cross-arm equivalence (TOST) Is slope A ≡ slope B in the SAME person? Certifies the pipeline. T4 Across people Do motion-drift and face-drift covary across participants (r)?
§5 · Analysis pipeline

EZ (screen) + HDDM (confirm)

EZ-diffusion

Wagenmakers 2007 — fast, closed-form per-subject baseline. 0/19 precedent — chosen from best-practices, not the corpus.

HDDM

Wiecki, Sofer & Frank 2013 — hierarchical Bayesian. Plurality choice among corpus papers naming a method (4/6).

Why HDDM is required

Simulated parameter recovery at our trial budget:

rv = 0.99  |  ra ≈ 0.55–0.59  |  rt₀ ≈ 0.53–0.55

Drift is recovered cleanly under EZ; boundary and non-decision time are not. Any "drift-specific" (T2) claim needs HDDM and/or higher trial counts before it can be trusted.

§7 golden standards · comparison

The 19-paper corpus vs. 8 inclusion criteria

Every candidate paper is scored against 8 criteria (2AFC · face stimuli · gaze/emotion axis · graded masking · full DDM fit · within-subject · RaFD-class stimulus set · adequate trials/cell). All eight required for a strict PASS.

0

strict PASS

14

PARTIAL

5

FAIL

Headline: no paper in this set is a strict PASS on all eight criteria simultaneously. No study has run the RDM-equivalent, per-subject-calibrated, graded-masking 2AFC+DDM design on faces — our project would be the first to combine all eight elements at once, for either axis.

§7 · Key slide

The tally — where our design has zero precedent

Full combination (this design)
0/19
Graded masking (any axis)
0/19
Adaptive staircase (QUEST/Ψ)
0/19
EZ-diffusion used
0/19
Emotion as discrimination axis
12/19
Gaze as discrimination axis
2/19

Emotion outnumbers gaze 6:1 among papers that judge the face's own content — and neither of the two gaze papers combines good power with a direct gaze-2AFC-plus-DDM design. These gaps are not weaknesses to hide; they are the design's selling points.

§7 · What the corpus teaches

Best-practice lessons from the 19 papers

Each lesson is argued from a specific source paper — every paper below links to its bibliography entry and full text.

Never trust mean RT

A single mean-RT number collapses four dissociable processes. Tipples shows an angry-vs-happy conclusion can flip with outlier-removal / RT-model choice — so we pre-commit to distributional RT + DDM.

Tipples 2023bibpdf

DDM reveals what mean RT hides

The founding demonstration: a drift-rate difference (threat → v↑ in high-anxious readers) is invisible to mean RT / accuracy but recovered by the full diffusion model — the reason our readout is v,a,t₀,z, not RT.

White, Ratcliff, Vasey & McKoon 2010bibpdf

Pre-register + anchor masking→drift

Two pre-registered studies (N=228 RaFD, N=264) show occluding the expressive face region lowers v (b=−0.38…−1.12). Our masking→drift anchor and our pre-registration discipline come from here.

Williams et al. 2023bibpdf

Energy-match the controls

To prove an emotion→v effect is not a low-level artifact, Sawada contrasts expressions against energy-matched anti-expressions. Our low-level confound controls (luminance, contrast, SF) inherit this logic.

Sawada et al. 2022bibdoi

Gaze cueing is a t₀ effect — a warning

Best-powered gaze DDM (N=171, 139,001 trials): the cueing effect loads on non-decision time, not drift (v inclusion 3–13%). A red flag for assuming our gaze arm (V1) will show a clean drift effect — hence the pilot.

Alister, McKay, Sewell & Evans 2023bibdoi

Right task, but no DDM

The one direct left/right gaze 2AFC with parametrically graded head×eye angle — but threshold-only, no v/a/t₀ fit. It seeds our gaze stimulus calibration (cone ½-diff ≈10°), not a drift anchor.

Palmer, Caruana & Clifford 2018bibdoi

Two lessons repeat across the corpus and become non-negotiable in our pre-reg: (1) distributional RT + full DDM, never mean RT (Tipples, White); (2) low-level confound control (Sawada). The gaze pair (Alister, Palmer) jointly warn that the gaze arm is the fragile one.

Effect sizes retrieved

What the literature actually reports, numerically

Williams et al. 2023 — masking → v

Collapsed-across-emotion, lower-mask vs. none: b = −0.38 [−0.41,−0.34] (Study 1, RaFD, N=228) → b = −0.65 [−0.71,−0.59] (Study 2, RADIATE, N=264). Per-emotion up to b = −1.12 (happiness/sadness).

PHS05 — coherence → v (RDM)

µ′ = k·x, sensitivity modeled as k ~ N(21,6) between subjects (range ≈ 9–28); non-decision tR ≈ 300–420 ms.

Sawada et al. 2022 — emotion → v

GAP Paywalled — no numeric value retrievable. Direction only: v larger, t₀ shorter, a larger for normal vs. anti-expressions.

Gaze → v

GAP No drift-rate anchor exists. Alister 2023: gaze is a t₀ effect (v inclusion prob. 3–13%). Palmer/Caruana/Clifford/Seymour 2018: no DDM fit at all.

§6 · Power

The claim-tier decision matrix

Four claim tiers — ambition sets the minimum N sample size is driven by participants, not trials per cell 1 Validation T1 + T2 min N 12–16 Pipeline reproduces masking→drift, drift-specific 2 + Convergent validity + T4 (r=0.5) min N 50–65 Arms covary across people — recommended default 3 Strong equivalence T3 (TOST) min N 50 · 130 · 260 "Same process in a person" — margin is a PI call 4 Full factorial V3 gaze×emotion min N ~80 Interaction on drift (optimistic effects) more ambitious claim, more participants →
§6 · Power

The equivalence-margin trade-off

The equivalence-margin trade-off (T3) a tighter margin you defend costs more participants N 50 130 260 50 participants Δ = 0.5 lenient 130 participants Δ = 0.3 moderate 260 participants Δ = 0.2 strict equivalence margin Δ (SD units) — narrower = stronger claim
Power curves across sample size for T1, T2, T4, and T3 equivalence margins
Power curves — minimum N by test and equivalence margin (power_curves.png)
§6 · Recommendation

Sample-size recommendation + PI leaning

PI's instinct

Leans toward tiers ① (cheap validation) and ④ (full factorial) — the two ends of the ambition spectrum.

Grounded compromise path

  • Flagship V2 at N≈50 (tier ②, best-anchored arm)
  • V1 gaze pilot (~10) in parallel, to obtain the missing drift anchor
  • Pre-register ③ equivalence at Δ=0.3 as an honestly under-powered secondary outcome
  • Upgrade to ④ factorial, N≈130 in a committed follow-up

This is the decision the team makes in the room today — the tier choice determines what Part 1 can honestly claim.

§6 · The choices for the room

Decisions on the table

Five explicit decisions. For each: the question, every option with its argument for and its cost, and the current leaning — marked as a leaning, not decided.

  • D1Discrimination axis / design version — V1 gaze vs V2 emotion vs V3 factorial
  • D2Obscuring mechanism — graded Mondrian mask vs noise degradation vs dichoptic b-CFS
  • D3Power anchor — Williams 2023 vs Sawada 2022 vs both, bracketed
  • D4Claim tier — validation vs +convergent vs strong equivalence vs factorial
  • D5Equivalence margin (if tier ③) — Δ = 0.5 vs 0.3 vs 0.2
Decision D1 · Design version

QWhich face-discrimination axis is the flagship — and why RaFD, not FACES or synthetic?

V1 — Gaze (L / R) 0/19

FOR — elegant structural mirror of RDM motion-direction (left/right ↔ left/right); the single highest-novelty arm.
COSTno published drift anchor exists (Alister: gaze cueing is a t₀ effect; Palmer: no DDM). Sample size is unknowable until a pilot runs.

V2 — Emotion (angry / happy) anchored

FOR — best-anchored (Williams masking→v on RaFD; emotion 12/19). Powerable today; the safe flagship.
COST — direct angry-vs-happy magnitude (Sawada) is paywalled; less structurally parallel to motion than gaze.

V3 — Factorial (gaze × emotion)

FOR — richest, most publishable single study; the only design that tests an interaction on drift.
COST — largest N (≈80, interaction fragile — 79% power only under optimistic effects); inherits the un-anchored gaze arm.

Why RaFD, not FACES: RaFD ships 3 gaze directions; FACES is frontal-only, which would rule out V1 and V3 entirely. RaFD is justified on its own validation (Langner 2010); it is used by our anchor paper (Williams 2023, Study 1), so it is not novel to the corpus.

Why not synthetic: an in-house synthetic set is used only for pipeline/recovery tests — it lacks the ecological validity and norming a substantive face claim requires.

Leaning, not decided: V2 (emotion) as flagship + a ~10-person V1 (gaze) pilot in parallel to buy the missing drift anchor; keep V3 costed and ready.
Decision D2 · Obscuring mechanism

QHow do we grade the face's discriminability across the difficulty ladder?

Graded binocular Mondrian mask novel · 0/19

FOR — mirrors the coherence ladder in Arm A (Newsome & Paré 1988 for per-subject coherence/noise-dot calibration; Tsuchiya & Koch 2005 10 Hz Mondrian dynamics). Sets up the pre-CFS baseline the parent project builds on.
COSTzero face-DDM precedent for graded severity (only 1 distinct study masks at all — Williams; Hartmann is its preprint — and it is binary). The per-level mask→v slope is assumed linear — must be re-estimated from pilot.

Contrast / phase-noise degradation

FOR — psychophysically standard, continuously parametric, trivially energy-matched; easiest to calibrate.
COST — departs from the Mondrian/CFS lineage that motivates the project; weaker bridge to the eventual b-CFS arm.

Dichoptic b-CFS out of scope

FOR — the project's namesake and ultimate target: true interocular suppression.
COSTout of scope for Part 1 — rivalry/awareness confounds are exactly what the binocular pre-CFS baseline is designed to exclude first.
Leaning, not decided: graded binocular Mondrian mask, explicitly flagged NOVEL (0/19 precedent) and justified from the RDM/CFS graded-noise tradition, not a face-DDM template.
Decision D3 · Power anchor

QWhich effect size anchors the masking→drift power analysis?

Williams 2023 only

FOR — the only numerically retrieved masking→v effect (b = −0.38 … −1.12), pre-registered, RaFD, N=228/264.
COST — masking is binary, not graded; a single-source anchor with no independent magnitude to cross-check.

Sawada 2022 only

FOR — the cleanest direct angry-vs-happy control (energy-matched anti-expressions) — closest to our actual task.
COST — magnitude is paywalled / unavailable; only the direction is known — it cannot seed a number.

Both, bracketed chosen

FOR — brackets the plausible range (conservative Δv = −0.38 → optimistic −1.12); power is reported as an honest band, not a false point estimate.
COST — wider N band; because Sawada gives direction only, the bracket ultimately rests on Williams' magnitudes.
Chosen (not merely leaning): anchor on both, bracketed — Williams supplies the numeric range; Sawada corroborates direction. Note the Sawada magnitude is paywalled/unavailable and remains a flagged gap.
Decision D4 · Claim tier

QHow strong a claim does Part 1 make — and at what participant cost?

① Validation N≈12–16

FOR — cheap, fast, essentially guaranteed (99.8% power at N=16); proves the pipeline reproduces masking→drift and it's drift-specific.
COST — modest claim; T2 selectivity leans on a/t₀ that EZ recovers poorly (needs HDDM); says nothing about cross-domain generality.

② +Convergent N≈50–65

FOR — affordable and publishable; motion-drift and face-drift covary across people on the best-anchored arm (V2).
COST — correlation attenuated by recovery noise (needs N≈50 for r=0.5); it is convergence, not equivalence.

③ Strong equivalence N≈50–260

FOR — the flagship scientific claim: same accumulation process within a person → the DDM as a domain-general readout.
COST — expensive and margin-driven (N≈130 at Δ=0.3, 260 at Δ=0.2); at N≈50 it is honestly underpowered.

④ Full factorial N≈80

FOR — richest single study; an interaction on drift; the most publishable one-shot.
COST — most fragile (79% power only under optimistic d=0.5; not reached under the conservative bracket); inherits the un-anchored gaze arm.
Leaning, not decided: the PI is drawn to ④ and ① (the two ends); the analyst recommendation is ② — flagship V2 at N≈50 (validation + convergent), with ③ equivalence pre-registered at Δ=0.3 as an honestly-underpowered secondary, upgrading to ④/N≈130 in a committed follow-up.
Decision D5 · Equivalence margin (only if tier ③)

QIf we pursue strong equivalence (T3 / TOST), how tight a margin do we defend?

Δ = 0.5 — lenient N≈50

FOR — affordable (same N as tier ②); claims the two slopes are “not wildly different.”
COST — a weak claim — a 0.5 SD band is wide; a skeptic can dismiss the equivalence as vacuous.

Δ = 0.3 — moderate N≈130

FOR — a genuinely defensible “practically equivalent” claim (N≈130; 160 for 90% power).
COST — ≈2.6× the cost of Δ=0.5; infeasible as a first in-house study without a follow-up commitment.

Δ = 0.2 — strict N≈260

FOR — a strong, hard-to-dismiss equivalence claim — the tightest band anyone would ask for.
COST — ≈260 participants (>260 for 90%) — infeasible for Part 1 by a wide margin.
Leaning, not decided: pre-register Δ = 0.3 as a secondary outcome now (honestly underpowered at N≈50), and commit the full N≈130 only in a follow-up that makes equivalence the headline.
§6 · How we choose N

Simulated power — how each tier's N was set

The N's in D4/D5 are not rules of thumb: each is read off a Monte-Carlo simulation of this exact pipeline.

How to read the next five figures

  • Each is the empirical sampling distribution of the actual test statistic, 3,200 iterations per arm, at the recommended N.
  • The H0 curve = no effect; the H1 curve = the true (bracketed) effect.
  • The 5% α critical value and the power region (target 80%) are shaded directly on the curves.

Which figure backs which decision

  • Scenario 1 → tier (D4) · V2 / T1
  • Scenario 2 → tier (D4) · V2 / T4
  • Scenario 3 → tier (D4) + D5 margin · T3 TOST
  • Scenario 4 → tier (D4) · V3 interaction
  • Scenario 5 → V1 gaze pilot (D1) · extrapolated
How we choose N · Scenario 1 → Tier ①

Validation — N = 16 (V2, test T1)

Power scenario 1: H0 vs H1 distributions for the validation test T1 at N=16, with 5% alpha and 80% power regions
H0 (no masking effect) vs H1 (conservative Δv = −0.38, Williams Study 1) at N=16, 100 trials/cell, α=0.05 two-sided

Because the masking effect is huge relative to N=16 sampling noise, the two distributions barely overlap — achieved power 99.8%. “Obscuring lowers drift” is essentially guaranteed to be detected.

The honest reading: N=16 is already more than this test needs. The real constraint on tier ① is practical (recruit a small cohort), not statistical.

How we choose N · Scenario 2 → Tier ②

Convergent validity — N = 50 (V2, test T4)

Power scenario 2: one-sided H0 vs H1 distributions for the cross-arm correlation T4 at N=50
Cross-arm drift correlation, true r=0.5, N=50, 100 trials/cell; one-sided α=0.05 (H1: r>0)

Recovered r is centered near 0.35–0.40 — attenuated well below the true 0.5 by EZ-recovery noise. That attenuation is exactly why N=50, not the textbook ~28, is required for 80% power.

Achieved power 81.0% — matches the report's headline N=50 recommendation for the recommended default tier.

How we choose N · Scenario 3 → Tier ③ + margin D5

Strong equivalence — N = 130, Δ = 0.3 (test T3, TOST)

Power scenario 3: TOST equivalence geometry at N=130 with a 0.3 SD acceptance interval
TOST cross-arm slope difference: H1 = true equivalence (diff=0); H0 = boundary of non-equivalence (diff=+0.3 SD). Dashed lines = the Δ=0.3 acceptance interval, N=130

Mass of the true-equivalence curve inside the acceptance interval = power (87.6%); the boundary-null curve inside it = Type-I error (≈5%).

The single most expensive claim in the design. N=130 buys only the moderate Δ=0.3 margin (D5); a tighter, more defensible Δ=0.2 needs roughly double.

How we choose N · Scenario 4 → Tier ④

Full factorial — N = 80 (V3, gaze × emotion interaction)

Power scenario 4: H0 vs H1 distributions for the gaze-by-emotion interaction at N=80 under an optimistic effect
Gaze×emotion interaction contrast on v, N=80, 100 trials/cell, α=0.05 two-sided; effect = extrapolated d=0.5 (optimistic bracket, paired with Δv=−1.12)

Achieved power 79.1% — right at the 80% line, but only under the optimistic effect. No published anchor exists for this interaction.

The design's most fragile confirmatory claim: under the conservative bracket it is not reached at all without N well beyond 80.

How we choose N · Scenario 5 → V1 gaze (D1)

Gaze arm — N = 10 pilot (V1, EXTRAPOLATED)

Power scenario 5: gaze-arm T1 at N=10 with an inset sensitivity band sweeping baseline drift
Masking→drift (T1) applied to the gaze arm, N=10, 60 trials/cell. Main panel: baseline v=2.0 (85.9% power). Inset: sensitivity band sweeping the un-anchored baseline v ∈ {1.0,1.5,2.0,2.5}

The Williams Δv=−0.38 slope is borrowed onto a gaze task with no published drift anchor (Alister: gaze cueing loads on t₀, not v). So the power claim is a range (71–94%), not a point.

The band straddles the 80% line — which is the argument for running the ~10-subject pilot (D1) before committing a full V1 sample.

§7–8 · Plan

Best-practice spine + roles & milestones

  1. Pre-registration (hypotheses, N, exclusions, margin)
  2. Adaptive per-subject calibration (QUEST/Ψ) — novel for face-DDM
  3. Graded difficulty in both arms, matched levels
  4. EZ (screen) + HDDM (confirm); report recovery
  5. Low-level confound controls; counterbalancing
  6. Distributional RT analysis — never mean-RT
  7. Open task, analysis, and data
M0
Contract approved
today
M1
RaFD subsets + masks built
+2 wk
M2
Two-arm PsychoPy task running
+4 wk
M3
V1+V2 pilot (~10)
+6 wk
M4
Pre-registration
+8 wk
M5
Main collection (N≈50)
+8–16 wk
M6
DDM analysis + writeup
+16–20 wk
§9 · Where the science is

Research gaps — the decision to be made today

  1. No drift-rate estimate exists for gaze-direction discrimination — V1 answers this.
  2. Graded masking of faces + full DDM has never been run — T3 answers this.
  3. Within-person equivalence of a "high-level" and "low-level" decision is untested — a domain-general accumulation claim.
  4. Reliability/individual-differences of face-DDM parameters is essentially unstudied (T4) — a tool for clinical/developmental work.
  5. Sawada 2022 magnitude gap — a concrete literature-completion task.

The decision for the room: which claim tier (§6) and which design version (§3) do we commit to for Part 1?

Listen

Two audio deep-dives

Spoken walkthroughs — a general overview and a focused dive on the design decisions and DDM parameters.

General deep-dive

The whole design contract — two-arm structure, the three versions, the golden-standard literature comparison, the power analysis, and the open questions.

Design decisions & parameters

A focused dive on the design choices and the four DDM parameters (v, a, t₀, z) — why each option was taken and what it buys.

Annex

Supporting material — full tables, checklists, figures, and references

  • A1Full 19-paper comparison table
  • A1·Which design each paper uses — linked (bib + PDF)
  • A2The 8-criterion checklist
  • A3Per-test / per-version N tables
  • A4Equivalence-margin table (T3, full)
  • A5Parameter recovery — recovery_scatter.png
  • A6Effect-size provenance table
  • A7Full milestone timeline + roles
  • A8RaFD dataset structure
  • A9Glossary of DDM terms
  • A10References — key anchors & corpus
Annex A1

Full 19-paper comparison table

PaperDOIVerdictOne-line reason
Sawada, Sato, Nakashima & Kumada 202210.1016/j.cognition.2022.105235PARTIALFace-in-crowd detection, not single-face 2AFC; no graded masking; best precedent for anger/happy→v direction.
Brennan & Baskin-Sommers 202010.1177/0956797620904157PARTIALEmotion-identification (likely >2 alternatives); no masking; individual-difference design.
Brennan & Baskin-Sommers 202110.1037/per0000473PARTIAL3-way blend/context grading, not occlusion; >2 response options.
White, Ratcliff, Vasey & McKoon 201010.1037/a0019474FAILThreat words, not faces — fails stimulus criterion outright.
Williams, Haque, Mai & Venkatraman 202310.1038/s41598-023-35381-4PARTIALBest masking precedent; but 6-way choice and binary (not graded) masking. Strongest structural match overall.
Hartmann et al. 2021 (preprint)10.31234/osf.io/a8yxfPARTIAL (low conf.)Duplicate of Williams 2023 design; no independent verifiable data.
Ozturk et al. 202410.1016/j.bpsgos.2023.07.005PARTIALGraded manipulation is cue uncertainty, not masking of the face itself.
Nagrodzki et al. 202510.1037/emo0001499FAILFace is a task-irrelevant incidental prime, not the judged stimulus.
Klein & Todd 202410.3758/s13423-024-02526-zFAIL2AFC axis is weapon-vs-tool identification; face is priming context.
Nan et al. 202410.1016/j.psyneuen.2023.106948PARTIAL (strong)Morph-continuum grading, not occlusion/masking; pharmacological RCT design.
Schreiber, Hall, Parr & Hallquist 202510.1017/S0033291725000595PARTIALDifficulty via congruent/conflicting emotion words, not masking.
Haller et al. 202410.1093/scan/nsae034PARTIAL (strong)Morph-continuum labeling, closest emotion-axis analog; small N=44 (fMRI).
Schrader, Habel, Jo, Walter & Wagels 202310.1016/j.concog.2023.103493PARTIAL (strong)Graded exposure duration — genuine masking-like manipulation; but 3-way choice.
Yang, Brunet-Gouet, Burca, Kalunga & Amorim 202010.3389/fnhum.2020.00340PARTIALUpright/inverted × photo/sketch degradation; binary factorial, not graded continuum.
Evans et al. 202510.1080/02699931.2025.2533382FAIL (axis)Approach-vs-avoid choice, not emotion/gaze identification.
Maka, Chrustowicz & Okruszek 202310.1111/psyp.14406FAIL (axis)Dot-probe task — response is about probe location, not the face.
Tipples 202310.1037/emo0001098PARTIALMethods-critique paper, not a new effect-size report; no masking.
Alister, McKay, Sewell & Evans 202310.1177/17470218231181238PARTIALBest-powered gaze precedent (N=171, 139,001 trials); but cueing task, not direct gaze judgment.
Palmer, Caruana & Clifford 201810.1098/rsos.180885PARTIALClosest gaze-axis task structure; but no DDM fit at all — threshold-only.
Annex A1 · continued

Which design does each paper use — linked

Stimulus set · axis · degradation · DDM software · titration · N. Each row links to its bibliography entry (bib) and full text (pdf local, or doi where no PDF is held).

PaperStimulus setAxisDegradationDDM softwareTitrationN
Sawada et al. 2022bibdoinot stated (face-in-crowd)Emotion (angry/happy vs anti-expr.)anti-expression control (not graded mask)custom (Ratcliff-style)nonen/r
Brennan & Baskin-Sommers 2020bibdoinot statedEmotion (anger)none (individual-difference)not stated (DDM)none90
Brennan & Baskin-Sommers 2021bibdoinot statedEmotion (anger/happy/fear blends)blend + threat context (not mask)not statedfixed blend levels92
White, Ratcliff, Vasey & McKoon 2010bibpdfthreat words (not faces)lexical decisionnoneRatcliff full-DDMnonen/r
Williams, Haque, Mai & Venkatraman 2023bibpdfRaFD (S1) / RADIATE (S2)Emotion (6-way, incl. anger)mask occlusion (binary)DDM (not named)none (binary)228 / 264
Hartmann et al. 2021 (preprint)bibpdfas WilliamsEmotion (6-way)mask occlusion (binary)not statednoneas Williams
Ozturk et al. 2024bibdoinot statedThreat vs neutralpre-stimulus cue uncertainty (not face mask)HDDM + SDTnone55
Nagrodzki et al. 2025bibdoinot statedAngry/neutral (task-irrelevant)none (incidental prime)DDMnone134
Klein & Todd 2024bibpdfBlack/White faces (prime)weapon ID (face incidental)expression salience factordiffusion decision modelnone546
Nan et al. 2024bibpdfnot statedEmotion (anger/fear morphs)morph continuum (not occlusion)DDM + SDTfixed morph levels120
Schreiber, Hall, Parr & Hallquist 2025bibpdfnot statedEmotion decodingemotion-word conflict (not mask)HDDMnone86
Haller et al. 2024bibdoinot statedEmotion (happy-angry morph)morph continuumDDM (sens./bias)fixed continuum steps44
Schrader, Habel, Jo, Walter & Wagels 2023bibdoinot statedEmotion (sad/neutral/happy, 3-way)graded exposure duration (8.3/16.7/25 ms)HDDM3 fixed durations40
Yang et al. 2020bibpdfphoto / sketch versionsEmotion recognitioninversion + sketch (not occlusion)DDM + ERPnone (binary factorial)n/r
Evans et al. 2025bibdoihappiness/anger morphsapproach/avoid (face incidental)morph continuumDDMfixed morph levelsn/r
Maka, Chrustowicz & Okruszek 2023bibdoinot stateddot-probe (face incidental)noneDDM + N2pcnone52
Tipples 2023bibpdfnot statedEmotion (angry vs happy, 2AFC)none (methods critique)recommends DDM / ex-Gaussiannonen/r
Alister, McKay, Sewell & Evans 2023bibdoireal face photos (gaze-cueing)Gaze (cue; response = target)none (congruent/incongruent)diffusion / LBA, hierarchical Bayesnone171
Palmer, Caruana & Clifford 2018bibdoiown gaze photos (not RaFD)Gaze (left/right, direct 2AFC)head × eye angle (graded angle, not occlusion)none (threshold-only)per-participant threshold22 SZ / 27 HC
Annex A2

The 8-criterion checklist

#CriterionOperational test
C1Forced binary choice (2AFC)Exactly two response options — not go/no-go, not >2-way categorization, not a rating scale.
C2Face stimuliDiscriminated stimulus is a face image — not words, dot arrays, or objects.
C3Axis = gaze OR emotion of the face itselfJudgment is about the face's own gaze/emotion — not a downstream cued target or orthogonal task.
C4Graded difficulty via masking/obscuringDiscriminability parametrically manipulated by occlusion/noise/masking — not fixed-intensity prototypes only.
C5Full DDM fit reportedDrift rate v recovered via a diffusion/sequential-sampling model — not threshold-only or SDT d′ alone.
C6Within-participant repeated measuresSame subjects run across multiple difficulty/condition levels.
C7RaFD or comparably validated stimulus setRaFD, KDEF, NimStim, or FACES with reported norms — not unvalidated in-house photos.
C8Adequate trial counts for per-cell recoveryN and trials/condition sufficient in principle for stable v estimation (≥100/cell EZ, ≥20–40/cell HDDM).

All eight required for a strict PASS; satisfying task-family intent but failing structural items = PARTIAL; failing the core task family = FAIL.

Annex A3

Per-test / per-version minimum-N tables

T1 / T2 — validation core

TestV1V2V3
T1 (masking lowers v)≤8≤8≤8
T2 (drift-specific)8–108–108–10

T4 — convergent validity

True rmin N 80%min N 90%
r = 0.55065
r = 0.3>100>100

V3 factorial — main effects & interaction (optimistic bracket)

Contrastmin N 80%min N 90%
Gaze main effect5065
Emotion main effect5065
Gaze × emotion interaction80>80

Conservative bracket: V3 main effects/interaction not reached within the simulated grid — factorial version viable only under optimistic effects or N ≥ ~80.

Annex A4

T3 equivalence — full margin table

Target powerMargin Δ (SD)Min N
80%0.2 (strict)260
80%0.3 (moderate)130
80%0.5 (lenient)50
90%0.2 (strict)>260 (not reached)
90%0.3 (moderate)160
90%0.5 (lenient)65

This is the binding constraint and the PI's decision. Trials/cell (60/100/150) barely move these N's — the cost is entirely in participants. Williams' masking is binary; our graded severity is an interpolated assumption, to be re-estimated from pilot data.

Annex A5

Parameter recovery — recovery_scatter.png

Simulated parameter recovery scatter plots for v, a, t0
True vs. recovered v, a, t₀ under EZ-diffusion, n=2000 synthetic subjects
rv
0.9873 — excellent
ra
0.5894 — weak
rt₀
0.5317 — weak
RT-variance inflation
≈ 3.95×
Drift rescaling
SCALE_K = 0.32552 (PHS05 k=21.0 → top-coherence v=3.5)

n=2000 synthetic subjects, seed 20260723. Drift is recovered cleanly; boundary and non-decision time are not — motivates HDDM confirmatory fitting for the selectivity claim.

Annex A6

Effect-size provenance table

SourceDesignValueStatus
Williams et al. 2023, Study 1RaFD, N=228, DDM (software not named), 648 trials/ppmasking→v: b=−0.38 [−0.41,−0.34]Retrieved (PMC full text)
Williams et al. 2023, Study 2RADIATE, N=264, DDM (software not named), 324 trials/ppmasking→v: b=−0.65 [−0.71,−0.59]; happiness/sadness up to b=−1.12Retrieved (PMC full text)
Williams et al. 2023 (boundary a)samea ≈ 1.67–1.80 across mask/emotion conditionsRetrieved; no t₀/z reported anywhere
Sawada et al. 2022face-in-crowd detectiondirection only: v↑, t₀↓, a↑ for normal vs. anti-expressionGAP — paywalled, no OA/preprint found
Palmer, Huk & Shadlen 2005 (PHS05)6 observers, coherence 3.2–51.2%µ′=k·x; mean k=20±3 (Exp.1), k=21±1/22±1 (Exp.3); A′≈0.6–0.86; tR≈300–420ms; threshold ratio≈3.0–3.8Retrieved (author PDF, PyMuPDF extraction)
Alister, McKay, Sewell & Evans 20233 gaze-cueing datasets, N=171, 139,001 trialsv has lowest model-inclusion probability (3–13%) of the 3 DDM parameters; no numeric v/z shift reportedGAP — v/z magnitudes not numerically reported
Palmer, Caruana, Clifford & Seymour 2018gaze 2AFC, N=22 SZ / 27 HCcone-model half-difference: SZ 9.94° (sd 9.58°); HC 10.68° (sd 6.81°)Retrieved; no DDM fit exists in this paper
Annex A7

Full milestone timeline + roles

#MilestoneOwnerTarget
M0This contract approvedalltoday
M1RaFD subsets + graded Mondrian masks builttechs+2 wk
M2PsychoPy two-arm task (RDM + face) runningtechs + analysts+4 wk
M3V1 gaze + V2 emotion pilot (~10) → real drift anchors, mask-level slopesall+6 wk
M4Pre-registration (final N, margin, exclusions) from pilot variancePI + analysts+8 wk
M5Main data collection (flagship V2, N≈50)techs+8–16 wk
M6DDM analysis, equivalence, writeupanalysts + PI+16–20 wk

Roles

PI
theory, design sign-off, pre-registration, equivalence-margin decision, writeup lead
Lab techs
RaFD subsetting, graded Mondrian mask generation, PsychoPy task build, apparatus, data collection
Analysts
EZ+HDDM pipeline, parameter-recovery checks, QUEST/Ψ code, power re-estimation, TOST, figures

Deliverables

  • This contract
  • Interactive KB webpage (report/)
  • This illustrated deck — main acts + annex (presentation/)
  • Golden-standard comparison
  • Grounded power analysis
Annex A8

RaFD dataset structure

Radboud Faces Database (Langner et al. 2010)

67 IDs × 8 expressions × 5 camera angles × 3 gaze = 8040 images

Models
67 (adults + children)
Expressions (8)
anger, disgust, fear, happiness, sadness, surprise, contempt, neutral
Camera angles (5)
full range; frontal angle used for this study
Gaze directions (3)
left / frontal / right
Access
free for academic use; on disk locally
Used for
V1 (gaze L/R), V2 (angry/happy), V3 (2×2 factorial)
Annex A9

Glossary of DDM terms

TermMeaning
2AFCTwo-alternative forced choice — exactly two response options per trial
v (drift rate)Average rate of evidence accumulation; the signal-quality parameter
a (boundary separation)Distance between the two decision thresholds; speed–accuracy tradeoff
t₀ / Ter (non-decision time)Encoding + motor time outside the accumulation process
z (starting point / bias)Where accumulation begins relative to the two boundaries
µ′ = k·xProportional-rate diffusion model: drift scales linearly with stimulus strength x, sensitivity k
EZ-diffusionFast, closed-form per-subject DDM estimator (Wagenmakers 2007)
HDDMHierarchical Bayesian DDM fitting package (Wiecki, Sofer & Frank 2013)
TOSTTwo one-sided tests — the standard statistical procedure for testing equivalence
QUEST / ΨAdaptive psychophysical staircase procedures for per-subject difficulty calibration
Parameter recoverySimulation check: can the fitting procedure recover known true parameter values?
Annex A10

References — key anchors & corpus

Key anchors

  • Palmer, Huk & Shadlen 2005 — J. Vision, DOI 10.1167/5.5.1
  • Williams et al. 2023 — Sci. Reports, DOI 10.1038/s41598-023-35381-4
  • Ratcliff & McKoon 2008 — DDM foundations
  • Wiecki, Sofer & Frank 2013 — HDDM
  • Wagenmakers et al. 2007 — EZ-diffusion
  • Langner et al. 2010 — RaFD validation, DOI 10.1080/02699930903485076
  • Tsuchiya & Koch 2005 — Mondrian CFS dynamics
  • Stein, Hebart & Sterzer 2011 — b-CFS background
  • Tipples 2023 — Emotion, DOI 10.1037/emo0001098

Full bibliography: data/biblio/references.yaml (140 entries).

Corpus & provenance sources

  • 19-paper golden-standard comparison: data/analysis/ddm-face-comparison.md
  • Retrieved effect sizes: data/analysis/effect-sizes-retrieved.md
  • Power analysis: data/power/power_report.md, power_analysis.py
  • Contract source: design/2026-07-23-rdmxbcfs-part1-design.md
  • Alister, McKay, Sewell & Evans 2023 — QJEP, DOI 10.1177/17470218231181238
  • Palmer, Caruana, Clifford & Seymour 2018 — R Soc Open Sci, DOI 10.1098/rsos.180885
  • Sawada, Sato, Nakashima & Kumada 2022 — Cognition, DOI 10.1016/j.cognition.2022.105235