RDMxbCFS · Part 1 · §6

Power & Sample Size

The grounded decision matrix: four claim tiers, the equivalence-margin trade-off, per-version minimum N, and the parameter-recovery caveat that governs what can honestly be claimed.

Simulation: power_analysis.py · seed 20260723 · effect sizes traced to effect-sizes-retrieved.md · 2026-07-23

Read this first: N is driven by participants, not trials. Trials/cell (60→100→150) barely move any of the minimum-N figures below. The expensive dial is how many people you recruit, and the equivalence claim (§3.3) is the one that scales hardest with how tight a margin you defend.

The claim-tier decision matrix

Claim tierTest(s)Min NCharacter
① ValidationT1 + T2~12–16Pipeline reproduces masking→drift, drift-specific. Cheap, modest claim.
② + Convergent validity+ T4 (r=0.5)~50–65Arms covary across people. Affordable, publishable. Recommended default.
③ Strong equivalenceT3 (TOST)50 (Δ=0.5) · 130–160 (Δ=0.3) · 260 (Δ=0.2)"Same process within a person." Margin is a PI judgment.
④ Full factorialV3 gaze×emotion~80 (optimistic)Interaction on drift. PI leans here. Needs optimistic effects.
Four claim tiers — ambition sets the minimum N sample size is driven by participants, not trials per cell 1 Validation T1 + T2 min N 12–16 Pipeline reproduces masking→drift, drift-specific 2 + Convergent validity + T4 (r=0.5) min N 50–65 Arms covary across people — recommended default 3 Strong equivalence T3 (TOST) min N 50 · 130 · 260 "Same process in a person" — margin is a PI call 4 Full factorial V3 gaze×emotion min N ~80 Interaction on drift (optimistic effects) more ambitious claim, more participants →
Four claim tiers. Ambition sets the minimum N: validation (T1+T2) needs ~12–16; adding convergent validity (T4) ~50–65; strong equivalence (T3) 50/130/260 depending on the margin; the full factorial ~80 (optimistic). Sample size scales with participants, not trials.

Recommendation, by scientific goal

If the headline claim is…Recommended NWhich version
"Our pipeline reproduces masking→drift and it's drift-specific" (validation)12–16V1 or V2
"Motion and face drift covary within people" (convergent validity)~50–65V2 (best-anchored)
"The two arms are quantitatively equivalent" (strong)130–160 (Δ=0.3)V2, or V3 with N≥80 for interaction
Full factorial gaze × emotion~80 (optimistic)V3
Suggested default for a first study: target the validation + convergent-validity tier at N ≈ 50 on the best-anchored arm (V2, emotion), run a V1 gaze pilot (~10) in parallel to obtain the missing drift anchor, and pre-register the equivalence test at a Δ = 0.3 margin as an honestly under-powered secondary outcome — upgrading to N≈130 only if a follow-up commits to the strong equivalence claim.

The equivalence-margin trade-off (T3)

This is the binding constraint and the PI's decision. A tight, defensible equivalence margin (0.2 SD) needs ~260 participants — infeasible for a first in-house study. A lenient margin (0.5 SD) is affordable (~50) but is a weak claim.

Equivalence margin Δ (SD units)min N (80% power)min N (90% power)
Δ = 0.5 (lenient)5065
Δ = 0.3 (moderate)130160
Δ = 0.2 (strict)260> 260
The equivalence-margin trade-off (T3) a tighter margin you defend costs more participants N 50 130 260 50 participants Δ = 0.5 lenient 130 participants Δ = 0.3 moderate 260 participants Δ = 0.2 strict equivalence margin Δ (SD units) — narrower = stronger claim
The equivalence-margin trade-off (T3). A margin of Δ=0.5 SD needs ~50 participants; Δ=0.3 needs ~130; Δ=0.2 needs ~260. The tighter the margin you defend, the more people the equivalence claim costs.
🎧 Contextual deep-dive

RDM & DDM best practices: where these numbers come from

A spoken walkthrough of the RDM/psychophysics literature (Palmer, Huk & Shadlen 2005PDF) that seeds the RDM-arm drift parameters used in this simulation.

Per-version N tables

Cheap tests (T1, T2) — the pipeline-validation core

TestV1V2V3Notes
T1 (masking lowers v)≤ 8≤ 8≤ 8floor-limited (grid starts at N=8); effect is very large
T2 (drift-specific)8–108–108–10see recovery caveat below

Reproducing "obscuring lowers drift" and showing it is drift-specific needs only a handful of participants — a pipeline-validation study is inexpensive. But see the caveat below: T2's low nominal N is necessary but not sufficient.

Convergent validity (T4) — do the two arms track each other across people?

True cross-arm rmin N (80%)min N (90%)
r = 0.55065
r = 0.3> 100 (not reached in grid)> 100

Attenuated by measurement error in recovered drift, so higher than the textbook N for a raw correlation.

Factorial (V3) main effects and interaction

Contrast (optimistic effects)min N (80%)min N (90%)
Gaze main effect5065
Emotion main effect5065
Gaze × emotion interaction80> 80

Under the conservative bracket, V3 main effects/interaction are not reached within the grid — the factorial version is only viable under optimistic effects or with N ≥ ~80.

How we choose N — H0 vs H1

Each figure below is a genuine Monte Carlo simulation (3,200 iterations per arm, power_figs.py, re-using power_analysis.py's exact Euler–Maruyama DDM simulator, EZ-diffusion forward/inverse equations, and calibrated fast sampler — no new generative model was introduced). Each panel plots the sampling distribution of the actual test statistic under the null (H0 — no effect) and under the alternative (H1 — the effect at its assumed size), at the N the report recommends for that claim tier. The α = 0.05 critical value marks where the null distribution is cut off — the shaded tail beyond it is the false-positive rate the test accepts; the shaded power region is the area of the H1 curve beyond that same cutoff — the probability of correctly detecting the effect at that N. Where a test has no single null/alternative pair in the usual sense (the TOST equivalence test, scenario 3), the geometry is adapted and explained in that figure's caption.

Tier 1 Validation: null and alternative sampling distributions of the per-subject masking-to-drift slope t-statistic at N=16, alpha=0.05 two-sided, achieved power 99.8%
Scenario 1 · Tier ① Validation — Design V2 (emotion), Test T1, N = 16, 100 trials/cell, α = 0.05 two-sided. Does obscuring lower drift rate v? The effect is Williams et al. 2023's conservative bracket (Δv = −0.38, collapsed masking slope, Study 1, RaFD, N=228). The null distribution of the one-sample t-statistic on each subject's OLS slope of recovered drift against masking level (centered at t=0) is overlaid against the alternative distribution under the true conservative effect, shifted far into the negative tail. Because the effect is this large relative to the N=16 sampling noise, the two distributions barely overlap: achieved power is 99.8% — this pipeline-validation claim is essentially guaranteed to be detected at this N. The honest reading: N=16 is already far more than this test needs, and the real constraint on this tier is practical (recruit a small, cheap validation cohort), not statistical.
Tier 2 Convergent Validity: null and alternative sampling distributions of the cross-arm correlation at N=50, one-sided alpha=0.05, achieved power 81.0%
Scenario 2 · Tier ② Convergent validity — Design V2 (best-anchored arm), Test T4, true cross-arm r = 0.5, N = 50, 100 trials/cell. Do RDM-arm and face-arm drift covary within the same person? Because T4's confirmatory hypothesis is directional (H1: r > 0), the critical value shown is the one-sided α = 0.05 threshold, not two-sided — a two-sided template would understate the power this N=50 recommendation is based on. The null distribution (r centered at 0) and the alternative distribution (recovered r centered near 0.35–0.4, attenuated below the true r=0.5 by EZ-recovery measurement noise) are shown with the one-sided rejection region shaded. Achieved power is 81.0%, matching the report's headline N=50 / 80%-power claim.
Tier 3 Strong Equivalence TOST: true-equivalence and boundary-null sampling distributions of the standardized cross-arm difference at N=130, margin 0.3 SD, achieved power 87.6%
Scenario 3 · Tier ③ Strong equivalence (TOST) — cross-arm equivalence test (T3) of the standardized difficulty→drift slope, RDM vs. face, within participant, N = 130, 100 trials/cell, margin Δ = 0.3 SD, α = 0.05 per one-sided test (two one-sided tests). This test has no single null/alternative pair, so the geometry is adapted: the “H1” curve is the mean standardized cross-arm difference simulated under true equivalence (true difference = 0); the “H0” curve is the same statistic simulated at the boundary of non-equivalence (true difference = +0.3 SD, the worst case TOST must guard against). The dashed vertical lines mark the TOST acceptance interval; equivalence is concluded whenever the observed mean difference falls inside it. That region evaluated under the true-equivalence curve is power (87.6%); the same region evaluated under the boundary-null curve is the Type-I error rate (≈5%) — wrongly concluding equivalence when the true difference sits exactly at the margin. This is the single most expensive claim in the whole design: N=130 buys only a moderate (0.3 SD) margin; a tighter, more defensible 0.2 SD margin needs roughly double the sample.
Tier 4 Factorial: null and alternative sampling distributions of the gaze by emotion interaction t-statistic at N=80, alpha=0.05 two-sided, achieved power 79.1%
Scenario 4 · Tier ④ Full factorial — Design V3 (gaze × emotion × 5 masking levels), gaze×emotion interaction on drift v, N = 80, 100 trials/cell, α = 0.05 two-sided. No published anchor exists for this interaction, so the effect is an extrapolated Cohen's d = 0.5 (optimistic bracket, paired with Williams's optimistic masking Δv = −1.12); under the conservative d = 0.3 bracket this contrast is not reachable within the simulated N grid. The null distribution (interaction contrast t centered at 0) and alternative distribution (shifted positive under the injected interaction effect) are shown with both two-sided critical values. Achieved power is 79.1%, essentially at the report's N=80 / 80%-power threshold for the optimistic bracket — this is the design's most fragile confirmatory claim: it only clears 80% power under the optimistic effect size, and is not reached at all under the conservative one without increasing N well beyond 80.
V1 gaze arm, extrapolated: null and alternative sampling distributions at N=10 with a sensitivity inset sweeping baseline drift, achieved power 85.9% at the central baseline, ranging 71 to 94% across the swept baseline
Scenario 5 · V1 gaze arm (pilot-dependent, EXTRAPOLATED) — same masking→drift test (T1) as Scenario 1, but applied to gaze discrimination, N = 10, 60 trials/cell, α = 0.05 two-sided. The Williams 2023 Δv = −0.38 masking slope is extrapolated onto a gaze-discrimination task for which no published drift-rate anchor exists (Alister 2023 finds gaze cueing loads mainly on non-decision time t₀, not drift v). The main panel shows the null vs. alternative distributions at the central baseline v_easy = 2.0 (85.9% power); the inset sweeps the unanchored baseline drift across {1.0, 1.5, 2.0, 2.5}, with achieved power ranging 71–94%, dipping below the conventional 80% threshold at the highest baseline. Because this arm's effect size is borrowed rather than measured, the honest power claim is a range, not a point estimate — straddling the 80% line is exactly the argument for running the pilot (~8–10 subjects, see below) first.

Figures

Power curves across sample size for T1, T2, T4, and T3 equivalence at three margins
Power curves. Statistical power as a function of sample size N for the confirmatory tests (T1 masking→v, T2 drift-specificity, T4 convergent validity) and the T3 equivalence test swept across the Δ=0.2/0.3/0.5 margins. The steep, cheap curves (T1/T2) sit far to the left; the equivalence curves fan out to the right as the margin tightens.
Parameter recovery scatterplots for v, a, and t0 under EZ-diffusion
Parameter recovery. True vs. recovered DDM parameters under EZ-diffusion at the simulated trial budget (n=2000 synthetic subjects). Drift rate v recovers tightly (r=0.99); boundary separation a and non-decision time t₀ are visibly noisier (r≈0.53–0.59) — the empirical basis for the recovery caveat below.

The recovery caveat — EZ recovers v well, a/t0 poorly

Parameter recovery at the simulated trial budget (n=2000 synthetic subjects, EZ recovery): r_v = 0.99 (excellent), r_a = 0.59, r_t0 = 0.53 (both weak). RT-variance inflation ≈ 3.95.

Drift is recovered cleanly; boundary and non-decision are not, at these trial counts with EZ. T2's "drift-specific" claim leans on a/t₀ staying null — a parameter EZ recovers poorly here. To defend drift-specificity credibly, raise trials/cell and/or fit hierarchically (HDDM) rather than per-subject EZ. T2's low nominal N is necessary but not sufficient — recovery quality, not detection power, is its real constraint.

The gaze arm (V1) is pilot-dependent

There is no published drift-rate estimate for gaze-direction discrimination (Alister 2023 shows gaze cueing is a non-decision-time effect, not drift; Palmer, Caruana, Clifford & Seymour 2018 fit no DDM at all). V1's numbers above assume the masking→v mechanism transfers to a gaze-discrimination task — plausible (it is a perceptual discrimination like RDM) but unverified. A short pilot (~8–10 subjects) to pin baseline gaze-discrimination v and confirm masking lowers it is a prerequisite before committing V1's sample size. This is the highest-novelty arm precisely because the number does not exist yet — see Gaps & open questions §1.

Assumptions and their provenance

QuantityValue usedSource
RDM drift vs. coherenceμ′ = k·x, k ~ N(21, 6) between subjects; coherence x ∈ {3.2, 6.4, 12.8, 25.6, 51.2}%Palmer, Huk & Shadlen 2005PDF, Table 2 (k range 9–28)
RDM bound / non-decisionA′ ≈ 0.71, t₀ ≈ 0.35 sPalmer, Huk & Shadlen 2005PDF
Face masking → drift (max drop)Δv = −0.38 (conservative) … −1.12 (optimistic) across ~5 graded levels; Δv = −0.25 sensitivity floor (see row below)Williams et al. 2023PDF: collapsed lower-mask b=−0.38 (Study 1, N=228) → −0.65 (Study 2); happy/sad up to −1.12. Cross-manipulation extrapolation: Williams' manipulation is binary garment occlusion (removes a face region), not the graded Mondrian visual-noise mask used here (adds noise / lowers contrast). Same DDM construct (signal-to-noise → v), different geometry — to be re-estimated from pilot. Triangulated with the closest in-kind precedents: Kalhan et al. 2022 (CFS/Mondrian on faces → measurable drift) and Schrader et al. 2023 (graded degradation → graded v drop).
Face masking → drift (shallow-slope sensitivity)Δv = −0.25 per collapsed contrast (below Williams' −0.38 conservative bracket)Covers the possibility that a graded Mondrian noise mask produces a smaller per-level drift drop than Williams' binary occlusion. Powering the Tier-1 validation slope test on Δv = −0.25 keeps N ≈ 16 comfortably >80% (the effect is still large relative to N=16 sampling noise); the convergent-validity and equivalence tiers are unaffected because they key off the cross-arm correlation, not the absolute Δv. Rationale: anchor-search report §5(b).
Face bound / non-decisiona ≈ 1.75; t₀ ≈ 0.35 s (assumed)Williams 2023PDF (a); t₀ assumed — not reported
Baseline (unmasked) discrimination driftv_easy ~ N(2.0, 0.5); V1 swept over {1.0, 1.5, 2.0, 2.5}No literature anchor — see Gaze pilot section
Emotion arm (V2) masking slopesame as WilliamsSawada 2022 paywalled/unavailable — direction only
Gaze arm (V1) driftno anchor — pilot-dependentAlister 2023 finds gaze cueing is a t₀ effect (v inclusion 3–13%); Palmer/Caruana/Clifford/Seymour 2018 give stimulus calibration only (cone ½-diff ≈ 10°)

Caveats