RDMxbCFS · Part 1 · Golden Standard

Golden-Standard Comparison

Our design (the benchmark): within-participant 2AFC on RaFD faces, judging gaze direction (left/right) or emotion (angry/happy), with faces graded in difficulty by a masking/obscuring manipulation, analyzed with a full DDM fit (choice + RT → v, a, t₀, z) — scored against 19 published face-2AFC-DDM papers.

Headline finding: no paper in this set is a strict PASS on all eight criteria simultaneously. Our project would be the first to combine all eight elements at once, for either axis.

1 · The eight inclusion/exclusion criteria

A study is scored against eight checklist items. Satisfying all eight is a strict PASS; satisfying the task-family/axis intent but failing one or more structural items is PARTIAL; failing the core task family (wrong response type, wrong stimulus modality, or no DDM fit at all) is FAIL.

Note on C4 vs. morph continua: several papers grade difficulty via emotion-morphing (neutral↔angry blends) rather than occlusion/masking. That is a legitimate difficulty manipulation but is not the masking/obscuring manipulation our design specifies, so it is flagged separately rather than silently counted as a C4 pass.

2 · Per-paper verdicts

PaperVerdictReason
Sawada, Sato, Nakashima & Kumada 2022 · Cognition · DOIPARTIALFace-in-crowd visual search/detection, not single-face 2AFC; no graded masking (categorical normal-vs-anti-expression); best precedent for the direction of the anger/happy→v effect.
Brennan & Baskin-Sommers 2020 · Psych Science · DOIPARTIALEmotion-identification task (likely >2 alternatives), incarcerated men; no masking manipulation; individual-difference (aggression) design, not stimulus-difficulty.
Brennan & Baskin-Sommers 2021 · Personality Disorders · DOIPARTIALThree-way anger/happiness/fear blends with contextual-threat manipulation — graded via blend/context not occlusion; >2 response options.
White, Ratcliff, Vasey & McKoon 2010PDF · Emotion · DOIFAILStimuli are threat words (lexical decision), not faces — fails C2 outright. Retained as the methodological template for DDM-reveals-what-mean-RT-misses.
Williams, Haque, Mai & Venkatraman 2023PDF · Sci Reports · DOIPARTIALBest masking/occlusion precedent (upper/lower face masked vs. unmasked → lower v) — but six-emotion categorization (fails strict C1) and masking is binary, not graded (partial C4). Two pre-registered studies, N=228/264 — otherwise the strongest structural match.
Hartmann et al. 2021PDF (preprint) · DOIPARTIAL (low confidence)Same design as Williams 2023PDF; no linked public OSF data found — treat as a duplicate/lower-confidence data point.
Ozturk et al. 2024 · Biol Psychiatry: GOS · DOIPARTIALBinary threat/neutral (2AFC-like) but graded manipulation is pre-stimulus cue uncertainty, not face masking; face-stimulus status uncertain; HDDM used (C5 satisfied).
Nagrodzki et al. 2025 · Emotion · DOIFAILDDM applied to a task-irrelevant sustained-attention measure — the face is an incidental prime, not the discriminated stimulus (fails C3).
Klein & Todd 2024PDF · Psychon Bull Rev · DOIFAIL2AFC dimension is weapon-vs-tool identification; face emotion/race are priming context, not the judged axis (fails C3).
Nan et al. 2024PDF · Psychoneuroendocrinology · DOIPARTIAL (strong)Anger/fear-vs-neutral morph judgment — binary-like response with graded morph intensity (good C1/C6); grading is morph, not occlusion (C4 mismatch). Pharmacological RCT.
Schreiber, Hall, Parr & Hallquist 2025PDF · Psych Medicine · DOIPARTIALHDDM on facial-emotion decoding (C5 satisfied) but difficulty manipulated via congruent/conflicting emotion words, not face masking (fails C4); response set likely >2 emotions.
Haller et al. 2024 · Soc Cogn Affect Neurosci · DOIPARTIAL (strong)Happy-to-angry morph continuum labeling task — closest emotion-axis analog to our design's intent, but graded via morphing not masking; small N=44 (fMRI).
Schrader, Habel, Jo, Walter & Wagels 2023 · Consciousness & Cognition · DOIPARTIAL (strong)Sad/neutral/happy faces at graded exposure durations (8.3/16.7/25 ms) — a genuine masking-like manipulation (best C4 match alongside Williams), HDDM, within-subject — but 3-way choice, not strict 2AFC; N=40.
Yang, Brunet-Gouet, Burca, Kalunga & Amorim 2020PDF · Front Hum Neurosci · DOIPARTIALUpright/inverted × photo/sketch is a form of image degradation (partial C4), DDM+ERP (C5), but binary factorial conditions, not a graded continuum.
Evans et al. 2025 · Cognition & Emotion · DOIFAIL (axis)Choice dimension is approach-vs-avoid, not emotion/gaze identification (fails C3), despite graded happy/anger morph intensity and a strong pre-registered-replication design.
Maka, Chrustowicz & Okruszek 2023 · Psychophysiology · DOIFAIL (axis)Dot-probe task — the 2AFC response is about probe location, not the face itself (fails C3). Valuable as a null/cautionary result, not a design template.
Tipples 2023PDF · Emotion · DOIPARTIALGenuine angry-vs-happy 2AFC axis (C1/C3 satisfied) but a methods-critique/re-analysis paper, not a new effect-size report; no masking manipulation. Retained as the mandatory methodological caution.
Alister, McKay, Sewell & Evans 2023 · QJEP · DOIPARTIALBest-powered DDM+gaze precedent (N=171, 139,001 trials) — but the response is to a cued target's location/identity (Posner gaze-cueing), not a judgment of the face's own gaze direction (fails strict C3); no masking.
Palmer, Caruana & Clifford 2018 · R Soc Open Sci · DOIPARTIALClosest task-structure match on the gaze axis: genuine left/right gaze 2AFC, head+eye angle parametrically varied — but reports thresholds in degrees only, no DDM fit at all (fails C5). A psychophysics paper, not a DDM paper.

2.5 · Result and best-practice lesson, paper by paper

For each of the 19 papers: its key empirical result, and the best-practice lesson it teaches for our own design — argued in one or two sentences, drawn from the verdicts above and the §4 synthesis below. All 19 link to their bibliography entry, and to a local PDF where one exists.

PARTIAL · Cognition

Sawada, Sato, Nakashima & Kumada 2022 DOI ↗

Result: Normal (angry/happy) expressions show larger drift v, shorter non-decision time t₀, and larger boundary a than energy-matched anti-expressions; arousal ratings correlate positively with v.

Lesson for our design: Use an energy-matched anti-expression control (not just a neutral face) to isolate a genuine emotion-specific drift effect from low-level image confounds like contrast and spatial energy.
PARTIAL · Psych Science

Brennan & Baskin-Sommers 2020 DOI ↗

Result: Aggression predicts a v increase specific to anger, with no effect on starting point z or boundary a.

Lesson for our design: A clean drift-vs-bias dissociation is only as strong as the task structure behind it — build a stimulus-driven difficulty ladder in from the start rather than relying on an individual-difference grouping alone.
PARTIAL · Personality Disorders

Brennan & Baskin-Sommers 2021 DOI ↗

Result: Psychopathy predicts a slower non-decision time t₀ (especially on angry-ambiguous trials); externalizing predicts a faster t₀ and larger boundary a.

Lesson for our design: Blend/context manipulations can reveal genuine parameter-selective effects (t₀ vs. a) — but they are not a substitute for a graded masking/occlusion severity ladder; keep the two manipulation types analytically separate.
FAIL · Emotion

White, Ratcliff, Vasey & McKoon 2010PDF DOI ↗

Result: High-anxious participants show a v advantage for threat words that mean RT and accuracy miss entirely.

Lesson for our design: Always report full DDM parameters alongside mean RT/accuracy — the drift-rate effect can be sitting exactly where the raw behavioral summary reads as null.
PARTIAL · Sci Reports

Williams, Haque, Mai & Venkatraman 2023PDF DOI ↗

Result: Masking the expressive (lower) face region lowers v across two pre-registered studies (N=228, N=264); the effect intensified later in the pandemic.

Lesson for our design: Pre-register a two-study design on a validated database (RaFD) — this is the strongest anchor in the corpus precisely because it replicates. Take its binary occlusion as the floor and extend it into a continuously graded ladder.
PARTIAL (low confidence) · OSF preprint

Hartmann et al. 2021 (preprint)PDF DOI ↗

Result: Same masking→v effect as Williams 2023 (same design), reported in a preprint with no independently linked public data.

Lesson for our design: Treat unpublished, data-unlinked preprints as duplicate lower-confidence evidence — don't double-count them when tallying corpus precedent or seeding an effect-size prior.
PARTIAL · Biol Psychiatry: GOS

Ozturk et al. 2024 DOI ↗

Result: High pre-stimulus threat-cue uncertainty raises v (efficiency) and shifts starting point z toward threat; anxiety raises z but not v.

Lesson for our design: Hierarchical DDM (HDDM) cleanly separates a drift effect from a starting-point/bias effect — fit both jointly from the outset rather than assuming a manipulation only moves v.
FAIL · Emotion

Nagrodzki et al. 2025 DOI ↗

Result: Angry (vs. neutral) faces lower v during an orthogonal sustained-attention task, an effect that scales with depression history.

Lesson for our design: If the face is only an incidental prime inside an unrelated task, the resulting DDM fit doesn't speak to face-discrimination drift at all — keep the face itself as the explicit 2AFC discriminandum, never a background stimulus.
FAIL · Psychon Bull Rev

Klein & Todd 2024PDF DOI ↗

Result: Race shifts the starting point z toward “gun” after Black faces; the shift weakens when emotional-expression salience is heightened — the effect lands on z, not v.

Lesson for our design: A large, well-powered two-experiment replication (N=546) can still land its effect on z rather than v — never assume a manipulation targets drift; test parameter selectivity explicitly (our T2).
PARTIAL (strong) · Psychoneuroendocrinology

Nan et al. 2024PDF DOI ↗

Result: Testosterone lowers sensitivity/v specifically for anger (not fear) in a double-blind, placebo-controlled design.

Lesson for our design: A pharmacological RCT plus morph-continuum grading plus SDT+DDM triangulation isolates a genuinely causal sensitivity effect — borrow the triangulation logic even though our own grading is masking, not morphing.
PARTIAL · Psych Medicine

Schreiber, Hall, Parr & Hallquist 2025PDF DOI ↗

Result: Emotion-related impulsivity lowers both v and boundary a under emotion-word conflict, in a hierarchical DDM fit.

Lesson for our design: HDDM recovers coordinated shifts in v and a even in a modest clinical sample (N=86) — its shrinkage is exactly what makes small-N designs tractable, reinforcing our EZ+HDDM pairing over EZ alone.
PARTIAL (strong) · Soc Cogn Affect Neurosci

Haller et al. 2024 DOI ↗

Result: Sensitivity (v) on a happy-to-angry morph continuum increases with age; anterior-insula response to ambiguity also increases with age.

Lesson for our design: A morph continuum supports stable per-subject sensitivity estimates even at small N=44 by leaning on within-subject repeated measures — repeated measures, not raw headcount, is what buys a stable drift function.
PARTIAL (strong) · Consciousness & Cognition

Schrader, Habel, Jo, Walter & Wagels 2023 DOI ↗

Result: Sad (vs. neutral/happy) faces lower both v and accuracy across three graded exposure durations (8.3/16.7/25 ms); 16.7 ms is flagged as the optimal subconscious-priming duration.

Lesson for our design: Graded brief-exposure duration is the closest existing analog to a graded masking-severity ladder — keep its 3-way response format from leaking into our design; ours must stay strictly 2AFC.
PARTIAL · Front Hum Neurosci

Yang, Brunet-Gouet, Burca, Kalunga & Amorim 2020PDF DOI ↗

Result: Face inversion lowers v; image impoverishment (photo→sketch) raises encoding time t₀; N170 ERP amplitude tracks v.

Lesson for our design: Pairing DDM with an EEG marker (N170) validates that v reflects a genuine perceptual/encoding stage rather than pure decision noise — worth keeping as a future confirmatory extension of our own pipeline.
FAIL (axis) · Cognition & Emotion

Evans et al. 2025 DOI ↗

Result: Conflict (approach-avoidance) trials lower v (noisier accumulation), in a pre-registered replication design.

Lesson for our design: A strong pre-registered replication proves reliability within its own task family, but a wrong discrimination axis (approach/avoidance vs. judging the face itself) can never be borrowed as a design template, however well-powered.
FAIL (axis) · Psychophysiology

Maka, Chrustowicz & Okruszek 2023 DOI ↗

Result: No threat-specific drift bias in a dot-probe task; the lonely group shows generally lower v and higher drift variability — a genuine null, not a non-result.

Lesson for our design: DDM can surface an informative null that mean RT would have missed entirely — don't discard a “boring” null from our own T1/T2 tests; report and interpret it with the same rigor as a positive effect.
PARTIAL · Emotion

Tipples 2023PDF DOI ↗

Result: The angry-vs-happy drift-rate conclusion flips depending on outlier-removal and RT-distribution modeling choices; no single new effect size is defended.

Lesson for our design: Commit to one pre-registered RT-trimming and DDM-fitting procedure before looking at the data — never trust a conclusion that depends on which of several plausible RT-cleaning choices you happened to make.
PARTIAL · QJEP

Alister, McKay, Sewell & Evans 2023 DOI ↗

Result: Gaze cueing loads primarily on non-decision time (attentional orienting) in a Posner gaze-cueing task, with the largest trial count (139,001) of any paper reviewed — v's inclusion probability is only 3–13%, so a clean drift effect is not supported.

Lesson for our design: With enough trials-per-subject, even a small non-drift-dominant effect becomes reliably estimable — our gaze-arm pilot should prioritize trials per subject, not just headcount, in case the true gaze→v effect turns out to be small.
PARTIAL · R Soc Open Sci

Palmer, Caruana & Clifford 2018 DOI ↗

Result: Head and eye angle combine to determine perceived gaze direction, with per-participant discrimination thresholds reported in degrees — no DDM parameter (v/a/t₀) is fit at all.

Lesson for our design: This is the best available task-structure template for a gaze 2AFC — bolt a DDM fit onto this exact graded head/eye-angle stimulus design rather than inventing a new gaze paradigm from scratch.
🎧 Contextual deep-dive

The emotion-DDM literature, walked through

A spoken tour of the emotion-axis papers in this corpus — Sawada, Williams, Nan, Haller, Schrader — and what each does and doesn't establish about masking/morphing and drift rate.

3 · Full structured comparison table

All 19 papers, all columns. “Not reported in KB” means the vault card's Results/Calibration block is empty and the narrative does not state the value — retrieval of the primary-source PDF/OSF data would be needed to fill it in. No number below has been invented.

CitationStimulus setAxisDifficulty/degradationN Trials/conditionDDM softwareParameters movedEffect size on v Reliability/robustness
Sawada et al. 2022, CognitionNot specified (RaFD not confirmed)Emotion (angry & happy vs. anti-expression)Anti-expression control, not graded maskingNot reportedNot reportedNot specifiedv↑, t0↓ for normal vs. anti-expression; a also larger; arousal correlates + with vDirection only — no numeric Δv/d/raw vNot pre-registered per KB
Brennan & Baskin-Sommers 2020, Psych ScienceNot specifiedEmotion (anger, among others)None (individual-difference design)N=90 (incarcerated men)Not reportedNot specified (likely HDDM)v↑ with aggression, specific to anger; no effect on z or aDirection onlyClinical/forensic sample; not pre-registered
Brennan & Baskin-Sommers 2021, Personality DisordersNot specifiedEmotion (anger/happiness/fear blends + threat context)Context/blend, not maskingN=92 (incarcerated men)Not reportedNot specifiedPsychopathy → t0↑; externalizing → t0↓, a↑Not reportedSame forensic cohort as 2020; not pre-registered
White, Ratcliff, Vasey & McKoon 2010PDF, EmotionN/A — threat words, not facesN/A (lexical decision)NoneNot reportedNot reportedRatcliff full-DDM (chi-square/quantile)v↑ for threatening words in high-anxious group; missed by mean RT/accuracyNot reportedFoundational demonstration paper, not a face study
Williams, Haque, Mai & Venkatraman 2023PDF, Sci ReportsRaFD (Study 1, N=228) / RADIATE (Study 2, N=264)Emotion (6-way, incl. anger)Face mask occlusion: upper/lower masked vs. unmasked (binary)Study 1 N=228; Study 2 N=264Not reportedDDM (software not named)v↓ when expressive region occluded; effect intensified later in pandemicb=−0.38→−0.65 collapsed (see Power page)Two pre-registered studies — strongest replication signal in the set
Hartmann et al. 2021PDF (preprint of Williams 2023PDF)Same as Williams 2023PDFSameSameSame designNot reportedNot specifiedSame as Williams 2023PDFNot reportedNo linked public OSF data — duplicate of Williams, excluded from counts (corpus effectively ≤18 distinct studies)
Ozturk et al. 2024, Biol Psychiatry: GOSNot specifiedThreat vs. neutral (binary)Pre-stimulus cue uncertainty, not face maskingN=55Not reportedHierarchical DDM (HDDM-class) + SDTHigh-uncertainty cues → v↑ and z shift toward threat; anxiety → z↑ not vNot reportedNot stated
Nagrodzki et al. 2025, EmotionNot specifiedAngry vs. neutral (task-irrelevant attention)None (passive/incidental viewing)N=134 (population-based fMRI)Not reportedDDM (software not named)v↓ for angry vs. neutral, scales with depression historyNot reportedCorrelates with insula/IFG/parietal activity
Klein & Todd 2024PDF, Psychon Bull RevNot specified (Black/White male faces)N/A (weapon identification; emotion is a prime)Emotion-expression salience as a factorN=546 (two experiments)Not reportedDiffusion decision modelRace → starting-point (z) shift toward "gun"; not a v effectN/A — effect lands on zTwo-experiment internal replication
Nan et al. 2024PDF, PsychoneuroendocrinologyNot specifiedEmotion (anger/fear vs. neutral morphs)Morph continuum (graded)N=120 men (double-blind RCT)Not reportedDDM + SDT + regressionTestosterone → sensitivity/v↓, specific to angerNot reported numericallyPharmacological RCT — causal manipulation, rare strength
Schreiber et al. 2025PDF, Psych MedicineNot specifiedEmotion (facial emotion decoding)Congruent/conflicting emotion word (Stroop-like), not maskingN=86 (adolescents/young adults, incl. BPD)Not reportedHierarchical DDMEmotion-related impulsivity → v↓ under conflict, a↓Not reportedClinical/developmental sample
Haller et al. 2024, Soc Cogn Affect NeurosciNot specifiedEmotion (happy-to-angry morph continuum)Morph continuum (graded ambiguity)N=44 (youth + adults, fMRI)Not reportedDDM (sensitivity/bias decomposition)Age → v (sensitivity)↑; anterior insula response to ambiguity ↑ with ageNot reportedDevelopmental design; small N typical of fMRI-DDM
Schrader et al. 2023, Consciousness & CognitionNot specifiedEmotion (sad/neutral/happy, 3-way)Graded exposure duration (8.3/16.7/25 ms) — masking-likeN=40Not reportedHierarchical DDMSad trials → v↓, accuracy↓; 16.7 ms flagged optimal subconscious-priming durationNot reportedLinks to subjective/objective awareness measures
Yang et al. 2020PDF, Front Hum NeurosciNot specified (photo/sketch versions)Emotion recognitionUpright/inverted × photo/sketch (image impoverishment)Not reportedNot reportedDiffusion decision model + ERPInversion → v↓; impoverishment → t0 (encoding)↑; N170 amplitude tracks vNot reportedEEG+DDM bridge; N170 as neural v-readout
Evans et al. 2025, Cognition & EmotionNot specified (happiness/anger morphs)N/A (approach-avoidance; emotion is graded input)Morph continuum (parametric social reward/threat/conflict)Not reported (pre-registered replication)Not reportedDDMConflict trials → v↓ (noisier accumulation)Not reportedPre-registered replication — genuine reliability strength, wrong axis
Maka et al. 2023, PsychophysiologyNot specifiedN/A (dot-probe; threat faces incidental)None (attention-capture paradigm)26 lonely vs. 26 non-lonely (N=52)Not reportedDDM + N2pc ERPNull for threat-specific drift bias; lonely group → v↓ generally, drift variability ↑N/A (null)Cautionary/null result — DDM more sensitive than raw RT
Tipples 2023PDF, EmotionNot specifiedEmotion (angry vs. happy, 2AFC)None (methods critique)Not reportedNot reportedRecommends DDM/ex-Gaussian/ex-Wald; no new v reportedN/A — conclusions flip with outlier/model choiceN/AMethodological warning, not an effect-size source
Alister, McKay, Sewell & Evans 2023, QJEPNot specified (real face photos)Gaze (cue direction; response to probed target)None reported as graded stimulus manipulationN=171 across 3 datasets139,001 trials totalDiffusion model / LBA, individual + hierarchical Bayesian→ primarily non-decision-time / orienting; v least-included (3–13%)Not reported numerically; v has lowest inclusion probability (3–13%)Largest trial count of any paper reviewed
Palmer, Caruana & Clifford 2018, R Soc Open SciNot specified (own gaze photos, not RaFD)Gaze (left/right, direct 2AFC)Head angle × eye/pupil angle independently varied (natural graded axis)Not reported (schizophrenia vs. control groups)Not reportedNone — threshold-only, no DDMN/AN/A — no drift-rate parameter exists in this paperPer-participant thresholds (degrees); task-structure template only

3.5 · Aggregate tally — how many papers made each design choice

Counted across all 19 target papers (White 2010PDF retained in the denominator despite failing C2, since it is on the named list). Every count is traceable to a specific paper above; nothing is estimated.

(a) Stimulus set

RaFD
1/19
RADIATE
1/19
KDEF / NimStim / FACES
0/19
Not specified in KB
16/19
N/A — not face stimuli
1/19

Modal choice: mostly unspecified. Most corpus papers do not name a validated stimulus database in a form this KB captured. RaFD is not novel to this corpus: our anchor paper, Williams et al. 2023, uses RaFD (Study 1) and RADIATE (Study 2). What is genuinely novel is not the database but the combination — RaFD + graded (non-binary) masking + strict 2AFC + full v/a/t₀/z fit — and the graded-masking severity itself (0/19; see below).

(b) Discrimination axis

Emotion
12/19
Gaze
2/19
Threat (binary, not one basic emotion)
1/19
Axis mismatch (face incidental)
3/19
Non-face
1/19

Modal choice: emotion (12/19, 63%). Emotion outnumbers gaze 6:1 among papers that actually judge the face's own content — the quantitative basis for the V1/V2 power imbalance (see §4.3 below).

(c) Degradation/difficulty manipulation

None / no graded manipulation
6/19
Morph continuum
4/19
Occlusion/mask (binary)
2/19
Non-stimulus contextual manipulation
3/19
Brief/backward-masked exposure duration
1/19
Graded stimulus angle (gaze-specific)
1/19
Image degradation (inversion/sketch)
1/19
N/A (no face stimulus)
1/19
Graded occlusion severity (our design)
0/19
Zero precedent flag. Occlusion/masking — the manipulation our design specifies — has only 2 precedents, and both are binary (masked vs. unmasked). A continuously graded occlusion/masking manipulation has zero precedent in this corpus. This is the single largest methodological gap our design must justify from first principles (drawing on the RDM coherence-calibration tradition — Newsome & Paré 1988 for per-subject coherence/noise-dot calibration — and the b-CFS masking literature), rather than from a face-specific precedent.

(d) DDM software

Not specified in KB
13/19
HDDM (explicitly named)
3/19
Custom hierarchical Bayesian
1/19
Ratcliff-custom (classic quantile)
1/19
No DDM fit at all
1/19
EZ-diffusion (our screening method)
0/19

Zero papers in this corpus report EZ-diffusion, despite it being the estimator this KB's best-practices docs recommend as the per-subject baseline. Of the 6 papers that name a method, HDDM is explicitly named in 3 and a further 1 uses a custom hierarchical-Bayesian fit — so hierarchical/Bayesian methods are the plurality (4 of 6 named cases, HDDM alone the single most common at 3/19). This supports HDDM as the confirmatory half of our EZ+HDDM pairing, even though EZ itself has no face-DDM precedent.

(e) Titration/calibration method

No titration / single fixed condition
13/19
Fixed graded levels (not adaptive)
5/19
Per-participant threshold, method unspecified
1/19
Adaptive staircase / QUEST / Ψ (our design)
0/19
Zero precedent flag. Not one of the 19 target papers uses an adaptive per-subject staircase or Bayesian procedure (QUEST/Ψ) to equate face-task difficulty across participants. Our planned per-subject Ψ/QUEST+ calibration is imported wholesale from the RDM/psychophysics tradition (Watson & Pelli 1983PDF; Kontsevich & Tyler 1999) — a genuine methodological contribution, not a replication of face-DDM practice.

(f) Trials-per-condition and total N ranges

QuantityValue
Papers reporting a numeric total N13/19 (Brennan 2020 N=90; Brennan 2021 N=92; Williams 2023PDF N=228/264; Ozturk 2024 N=55; Nagrodzki 2025 N=134; Klein & Todd 2024PDF N=546; Nan 2024PDF N=120; Schreiber 2025PDF N=86; Haller 2024 N=44; Schrader 2023 N=40; Maka 2023 N=52; Alister 2023 N=171)
N range40 (Schrader 2023) to 546 (Klein & Todd 2024PDF); median ≈ 92 across the 13 reported values
Papers reporting a trials/condition or total-trials figure1/19: Alister 2023 (139,001 trials total, N=171, not broken down per condition)
Papers with no trial-count figure at all18/19

Modal choice: trial counts are essentially unreported (18/19). The ≥100/cell (EZ) and ~20–40/cell (HDDM) benchmarks used in the power analysis come entirely from best-practices docs' independent RDM/psychophysics literature, not from a count taken across these 19 papers.

4 · Golden-standard synthesis

4.1 Ranking by evidential strength (drift-rate effects)

  1. Williams et al. 2023PDF (face-mask occlusion → v↓). Two independent pre-registered studies, consistent direction, real photographic faces — the strongest single anchor in the corpus for “obscuring the face lowers v” — with the caveat that its manipulation is binary garment occlusion, not our graded visual-noise mask; use as a directionally-valid, cross-manipulation anchor triangulated with Kalhan 2022 / Schrader 2023.
  2. Alister, McKay, Sewell & Evans 2023 (gaze → primarily non-decision time / attentional orienting; drift inclusion only 3–13%). By far the largest N and trial count — the best-powered DDM-on-gaze-behavior precedent, though the task is cueing rather than direct gaze-direction discrimination, and it is a cautionary precedent for our gaze arm: the cueing effect does not land cleanly on drift.
  3. Sawada et al. 2022 (anger/happy → v↑, t0↓). The cleanest mechanistic control (energy-matched anti-expressions) for isolating an emotion-specific v effect.
  4. Brennan & Baskin-Sommers 2020 (anger → v↑, no z/a). Clean parameter dissociation, single-cohort forensic sample, no independent replication.
  5. Schrader 2023 / Nan 2024PDF / Haller 2024 — each independently reports emotion-driven v modulation with a genuine graded manipulation, but each is a single N=40–120 study with no cross-paper replication of the same contrast.
  6. Palmer & Clifford 2018 (gaze discrimination, threshold-only). The best task-structure analog on the gaze axis but contributes no DDM parameter at all.
  7. Everything with an axis mismatch (Nagrodzki, Klein & Todd, Maka, Evans) — informative as cautionary/contrast cases, not as positive anchors.

4.2 What should anchor our pipeline

Grounded in the §3.5 counts, not assumption: EZ has zero precedent here and only 3/19 papers explicitly name HDDM, so the recommendation to pair EZ-diffusion (screen) + HDDM (confirm) is not the modal practice in this literature — it is a choice imported from best-practices docs drawing on the RDM literature (Wagenmakers 2007PDF for EZ, Wiecki 2013PDF for HDDM). Trial counts cannot be empirically derived from this corpus at all (only 1/19 reports any figure); the ≥100/cell and ~20–40/cell benchmarks come from independent RDM/psychophysics sources. Total N (13/19 reported, range 40–546, median ≈92) is the only recruitment-relevant number this literature actually supports empirically.

4.3 The gaze (V1) vs. emotion (V2) power imbalance

Emotion axis: 12/19 papers (63%), spanning clinical, developmental, pharmacological, and EEG designs, with two independent pre-registered studies and a mechanistically controlled anchor. Gaze axis: 2/19 papers (11%) — and neither is a full match: Alister 2023 has the best power in the entire corpus but studies gaze-cueing, not direct discrimination; Palmer & Clifford 2018 has the right task structure but no DDM fit whatsoever. There is no paper in this corpus that is both adequately powered and structurally a direct gaze-direction 2AFC with a DDM fit. This is the quantitative basis for calling gaze the weaker-anchored, higher-novelty arm.

4.4 Recommendation

Anchor the pipeline on Williams 2023PDF (masking→v, the strongest pre-registered emotion-side precedent) and Alister 2023 (largest-N gaze-side DDM precedent, cueing caveat) for seeding priors, use EZ+HDDM as the estimator pair, and treat the RaFD + graded-masking + strict-2AFC + full-v/a/t0/z combination as novel rather than a replication — no paper in this corpus has run that exact combination on either axis. Budget substantially more power/pilot data for the gaze axis than the emotion axis given the 6:1 paper-count imbalance.

4.5 Design choices with zero precedent — explicit flags

Per the counts above, the following elements of our planned design are not represented by even a single paper among the 19 reviewed, and must be treated as novel methodological steps requiring first-principles justification, not as replications of face-DDM practice:

  1. RaFD + graded masking + strict 2AFC + full v/a/t0/z as a single bundle — 0/19 (RaFD itself is not novel: Williams 2023 uses it; the novelty is the combination).
  2. Graded (non-binary) occlusion/masking severity — 0/19; only binary masked/unmasked exists.
  3. Adaptive per-subject staircase/QUEST/Ψ calibration of face-task difficulty — 0/19.
  4. EZ-diffusion as the estimator — 0/19; the pairing with HDDM is a best-practices import, not a literature norm.
  5. A single paper combining strict 2AFC + graded masking + full v/a/t0/z fit on the gaze axis — 0/19; the closest is split across two non-overlapping papers.

None of these gaps is disqualifying — they are precisely the gaps this design's best-practices spine (see Design §7) was written to fill using the RDM/psychophysics tradition. But they must be named as novel, not presented as if a precedent existed.