Our design (the benchmark): within-participant 2AFC on RaFD faces, judging gaze direction (left/right) or emotion (angry/happy), with faces graded in difficulty by a masking/obscuring manipulation, analyzed with a full DDM fit (choice + RT → v, a, t₀, z) — scored against 19 published face-2AFC-DDM papers.
A study is scored against eight checklist items. Satisfying all eight is a strict PASS; satisfying the task-family/axis intent but failing one or more structural items is PARTIAL; failing the core task family (wrong response type, wrong stimulus modality, or no DDM fit at all) is FAIL.
Note on C4 vs. morph continua: several papers grade difficulty via emotion-morphing (neutral↔angry blends) rather than occlusion/masking. That is a legitimate difficulty manipulation but is not the masking/obscuring manipulation our design specifies, so it is flagged separately rather than silently counted as a C4 pass.
| Paper | Verdict | Reason |
|---|---|---|
| Sawada, Sato, Nakashima & Kumada 2022 · Cognition · DOI | PARTIAL | Face-in-crowd visual search/detection, not single-face 2AFC; no graded masking (categorical normal-vs-anti-expression); best precedent for the direction of the anger/happy→v effect. |
| Brennan & Baskin-Sommers 2020 · Psych Science · DOI | PARTIAL | Emotion-identification task (likely >2 alternatives), incarcerated men; no masking manipulation; individual-difference (aggression) design, not stimulus-difficulty. |
| Brennan & Baskin-Sommers 2021 · Personality Disorders · DOI | PARTIAL | Three-way anger/happiness/fear blends with contextual-threat manipulation — graded via blend/context not occlusion; >2 response options. |
| White, Ratcliff, Vasey & McKoon 2010PDF · Emotion · DOI | FAIL | Stimuli are threat words (lexical decision), not faces — fails C2 outright. Retained as the methodological template for DDM-reveals-what-mean-RT-misses. |
| Williams, Haque, Mai & Venkatraman 2023PDF · Sci Reports · DOI | PARTIAL | Best masking/occlusion precedent (upper/lower face masked vs. unmasked → lower v) — but six-emotion categorization (fails strict C1) and masking is binary, not graded (partial C4). Two pre-registered studies, N=228/264 — otherwise the strongest structural match. |
| Hartmann et al. 2021PDF (preprint) · DOI | PARTIAL (low confidence) | Same design as Williams 2023PDF; no linked public OSF data found — treat as a duplicate/lower-confidence data point. |
| Ozturk et al. 2024 · Biol Psychiatry: GOS · DOI | PARTIAL | Binary threat/neutral (2AFC-like) but graded manipulation is pre-stimulus cue uncertainty, not face masking; face-stimulus status uncertain; HDDM used (C5 satisfied). |
| Nagrodzki et al. 2025 · Emotion · DOI | FAIL | DDM applied to a task-irrelevant sustained-attention measure — the face is an incidental prime, not the discriminated stimulus (fails C3). |
| Klein & Todd 2024PDF · Psychon Bull Rev · DOI | FAIL | 2AFC dimension is weapon-vs-tool identification; face emotion/race are priming context, not the judged axis (fails C3). |
| Nan et al. 2024PDF · Psychoneuroendocrinology · DOI | PARTIAL (strong) | Anger/fear-vs-neutral morph judgment — binary-like response with graded morph intensity (good C1/C6); grading is morph, not occlusion (C4 mismatch). Pharmacological RCT. |
| Schreiber, Hall, Parr & Hallquist 2025PDF · Psych Medicine · DOI | PARTIAL | HDDM on facial-emotion decoding (C5 satisfied) but difficulty manipulated via congruent/conflicting emotion words, not face masking (fails C4); response set likely >2 emotions. |
| Haller et al. 2024 · Soc Cogn Affect Neurosci · DOI | PARTIAL (strong) | Happy-to-angry morph continuum labeling task — closest emotion-axis analog to our design's intent, but graded via morphing not masking; small N=44 (fMRI). |
| Schrader, Habel, Jo, Walter & Wagels 2023 · Consciousness & Cognition · DOI | PARTIAL (strong) | Sad/neutral/happy faces at graded exposure durations (8.3/16.7/25 ms) — a genuine masking-like manipulation (best C4 match alongside Williams), HDDM, within-subject — but 3-way choice, not strict 2AFC; N=40. |
| Yang, Brunet-Gouet, Burca, Kalunga & Amorim 2020PDF · Front Hum Neurosci · DOI | PARTIAL | Upright/inverted × photo/sketch is a form of image degradation (partial C4), DDM+ERP (C5), but binary factorial conditions, not a graded continuum. |
| Evans et al. 2025 · Cognition & Emotion · DOI | FAIL (axis) | Choice dimension is approach-vs-avoid, not emotion/gaze identification (fails C3), despite graded happy/anger morph intensity and a strong pre-registered-replication design. |
| Maka, Chrustowicz & Okruszek 2023 · Psychophysiology · DOI | FAIL (axis) | Dot-probe task — the 2AFC response is about probe location, not the face itself (fails C3). Valuable as a null/cautionary result, not a design template. |
| Tipples 2023PDF · Emotion · DOI | PARTIAL | Genuine angry-vs-happy 2AFC axis (C1/C3 satisfied) but a methods-critique/re-analysis paper, not a new effect-size report; no masking manipulation. Retained as the mandatory methodological caution. |
| Alister, McKay, Sewell & Evans 2023 · QJEP · DOI | PARTIAL | Best-powered DDM+gaze precedent (N=171, 139,001 trials) — but the response is to a cued target's location/identity (Posner gaze-cueing), not a judgment of the face's own gaze direction (fails strict C3); no masking. |
| Palmer, Caruana & Clifford 2018 · R Soc Open Sci · DOI | PARTIAL | Closest task-structure match on the gaze axis: genuine left/right gaze 2AFC, head+eye angle parametrically varied — but reports thresholds in degrees only, no DDM fit at all (fails C5). A psychophysics paper, not a DDM paper. |
For each of the 19 papers: its key empirical result, and the best-practice lesson it teaches for our own design — argued in one or two sentences, drawn from the verdicts above and the §4 synthesis below. All 19 link to their bibliography entry, and to a local PDF where one exists.
Result: Normal (angry/happy) expressions show larger drift v, shorter non-decision time t₀, and larger boundary a than energy-matched anti-expressions; arousal ratings correlate positively with v.
Result: Aggression predicts a v increase specific to anger, with no effect on starting point z or boundary a.
Result: Psychopathy predicts a slower non-decision time t₀ (especially on angry-ambiguous trials); externalizing predicts a faster t₀ and larger boundary a.
Result: High-anxious participants show a v advantage for threat words that mean RT and accuracy miss entirely.
Result: Masking the expressive (lower) face region lowers v across two pre-registered studies (N=228, N=264); the effect intensified later in the pandemic.
Result: Same masking→v effect as Williams 2023 (same design), reported in a preprint with no independently linked public data.
Result: High pre-stimulus threat-cue uncertainty raises v (efficiency) and shifts starting point z toward threat; anxiety raises z but not v.
v.Result: Angry (vs. neutral) faces lower v during an orthogonal sustained-attention task, an effect that scales with depression history.
Result: Race shifts the starting point z toward “gun” after Black faces; the shift weakens when emotional-expression salience is heightened — the effect lands on z, not v.
z rather than v — never assume a manipulation targets drift; test parameter selectivity explicitly (our T2).Result: Testosterone lowers sensitivity/v specifically for anger (not fear) in a double-blind, placebo-controlled design.
Result: Emotion-related impulsivity lowers both v and boundary a under emotion-word conflict, in a hierarchical DDM fit.
v and a even in a modest clinical sample (N=86) — its shrinkage is exactly what makes small-N designs tractable, reinforcing our EZ+HDDM pairing over EZ alone.Result: Sensitivity (v) on a happy-to-angry morph continuum increases with age; anterior-insula response to ambiguity also increases with age.
Result: Sad (vs. neutral/happy) faces lower both v and accuracy across three graded exposure durations (8.3/16.7/25 ms); 16.7 ms is flagged as the optimal subconscious-priming duration.
Result: Face inversion lowers v; image impoverishment (photo→sketch) raises encoding time t₀; N170 ERP amplitude tracks v.
v reflects a genuine perceptual/encoding stage rather than pure decision noise — worth keeping as a future confirmatory extension of our own pipeline.Result: Conflict (approach-avoidance) trials lower v (noisier accumulation), in a pre-registered replication design.
Result: No threat-specific drift bias in a dot-probe task; the lonely group shows generally lower v and higher drift variability — a genuine null, not a non-result.
Result: The angry-vs-happy drift-rate conclusion flips depending on outlier-removal and RT-distribution modeling choices; no single new effect size is defended.
Result: Gaze cueing loads primarily on non-decision time (attentional orienting) in a Posner gaze-cueing task, with the largest trial count (139,001) of any paper reviewed — v's inclusion probability is only 3–13%, so a clean drift effect is not supported.
Result: Head and eye angle combine to determine perceived gaze direction, with per-participant discrimination thresholds reported in degrees — no DDM parameter (v/a/t₀) is fit at all.
A spoken tour of the emotion-axis papers in this corpus — Sawada, Williams, Nan, Haller, Schrader — and what each does and doesn't establish about masking/morphing and drift rate.
All 19 papers, all columns. “Not reported in KB” means the vault card's Results/Calibration block is empty and the narrative does not state the value — retrieval of the primary-source PDF/OSF data would be needed to fill it in. No number below has been invented.
| Citation | Stimulus set | Axis | Difficulty/degradation | N | Trials/condition | DDM software | Parameters moved | Effect size on v | Reliability/robustness |
|---|---|---|---|---|---|---|---|---|---|
| Sawada et al. 2022, Cognition | Not specified (RaFD not confirmed) | Emotion (angry & happy vs. anti-expression) | Anti-expression control, not graded masking | Not reported | Not reported | Not specified | v↑, t0↓ for normal vs. anti-expression; a also larger; arousal correlates + with v | Direction only — no numeric Δv/d/raw v | Not pre-registered per KB |
| Brennan & Baskin-Sommers 2020, Psych Science | Not specified | Emotion (anger, among others) | None (individual-difference design) | N=90 (incarcerated men) | Not reported | Not specified (likely HDDM) | v↑ with aggression, specific to anger; no effect on z or a | Direction only | Clinical/forensic sample; not pre-registered |
| Brennan & Baskin-Sommers 2021, Personality Disorders | Not specified | Emotion (anger/happiness/fear blends + threat context) | Context/blend, not masking | N=92 (incarcerated men) | Not reported | Not specified | Psychopathy → t0↑; externalizing → t0↓, a↑ | Not reported | Same forensic cohort as 2020; not pre-registered |
| White, Ratcliff, Vasey & McKoon 2010PDF, Emotion | N/A — threat words, not faces | N/A (lexical decision) | None | Not reported | Not reported | Ratcliff full-DDM (chi-square/quantile) | v↑ for threatening words in high-anxious group; missed by mean RT/accuracy | Not reported | Foundational demonstration paper, not a face study |
| Williams, Haque, Mai & Venkatraman 2023PDF, Sci Reports | RaFD (Study 1, N=228) / RADIATE (Study 2, N=264) | Emotion (6-way, incl. anger) | Face mask occlusion: upper/lower masked vs. unmasked (binary) | Study 1 N=228; Study 2 N=264 | Not reported | DDM (software not named) | v↓ when expressive region occluded; effect intensified later in pandemic | b=−0.38→−0.65 collapsed (see Power page) | Two pre-registered studies — strongest replication signal in the set |
| Hartmann et al. 2021PDF (preprint of Williams 2023PDF) | Same as Williams 2023PDF | Same | Same | Same design | Not reported | Not specified | Same as Williams 2023PDF | Not reported | No linked public OSF data — duplicate of Williams, excluded from counts (corpus effectively ≤18 distinct studies) |
| Ozturk et al. 2024, Biol Psychiatry: GOS | Not specified | Threat vs. neutral (binary) | Pre-stimulus cue uncertainty, not face masking | N=55 | Not reported | Hierarchical DDM (HDDM-class) + SDT | High-uncertainty cues → v↑ and z shift toward threat; anxiety → z↑ not v | Not reported | Not stated |
| Nagrodzki et al. 2025, Emotion | Not specified | Angry vs. neutral (task-irrelevant attention) | None (passive/incidental viewing) | N=134 (population-based fMRI) | Not reported | DDM (software not named) | v↓ for angry vs. neutral, scales with depression history | Not reported | Correlates with insula/IFG/parietal activity |
| Klein & Todd 2024PDF, Psychon Bull Rev | Not specified (Black/White male faces) | N/A (weapon identification; emotion is a prime) | Emotion-expression salience as a factor | N=546 (two experiments) | Not reported | Diffusion decision model | Race → starting-point (z) shift toward "gun"; not a v effect | N/A — effect lands on z | Two-experiment internal replication |
| Nan et al. 2024PDF, Psychoneuroendocrinology | Not specified | Emotion (anger/fear vs. neutral morphs) | Morph continuum (graded) | N=120 men (double-blind RCT) | Not reported | DDM + SDT + regression | Testosterone → sensitivity/v↓, specific to anger | Not reported numerically | Pharmacological RCT — causal manipulation, rare strength |
| Schreiber et al. 2025PDF, Psych Medicine | Not specified | Emotion (facial emotion decoding) | Congruent/conflicting emotion word (Stroop-like), not masking | N=86 (adolescents/young adults, incl. BPD) | Not reported | Hierarchical DDM | Emotion-related impulsivity → v↓ under conflict, a↓ | Not reported | Clinical/developmental sample |
| Haller et al. 2024, Soc Cogn Affect Neurosci | Not specified | Emotion (happy-to-angry morph continuum) | Morph continuum (graded ambiguity) | N=44 (youth + adults, fMRI) | Not reported | DDM (sensitivity/bias decomposition) | Age → v (sensitivity)↑; anterior insula response to ambiguity ↑ with age | Not reported | Developmental design; small N typical of fMRI-DDM |
| Schrader et al. 2023, Consciousness & Cognition | Not specified | Emotion (sad/neutral/happy, 3-way) | Graded exposure duration (8.3/16.7/25 ms) — masking-like | N=40 | Not reported | Hierarchical DDM | Sad trials → v↓, accuracy↓; 16.7 ms flagged optimal subconscious-priming duration | Not reported | Links to subjective/objective awareness measures |
| Yang et al. 2020PDF, Front Hum Neurosci | Not specified (photo/sketch versions) | Emotion recognition | Upright/inverted × photo/sketch (image impoverishment) | Not reported | Not reported | Diffusion decision model + ERP | Inversion → v↓; impoverishment → t0 (encoding)↑; N170 amplitude tracks v | Not reported | EEG+DDM bridge; N170 as neural v-readout |
| Evans et al. 2025, Cognition & Emotion | Not specified (happiness/anger morphs) | N/A (approach-avoidance; emotion is graded input) | Morph continuum (parametric social reward/threat/conflict) | Not reported (pre-registered replication) | Not reported | DDM | Conflict trials → v↓ (noisier accumulation) | Not reported | Pre-registered replication — genuine reliability strength, wrong axis |
| Maka et al. 2023, Psychophysiology | Not specified | N/A (dot-probe; threat faces incidental) | None (attention-capture paradigm) | 26 lonely vs. 26 non-lonely (N=52) | Not reported | DDM + N2pc ERP | Null for threat-specific drift bias; lonely group → v↓ generally, drift variability ↑ | N/A (null) | Cautionary/null result — DDM more sensitive than raw RT |
| Tipples 2023PDF, Emotion | Not specified | Emotion (angry vs. happy, 2AFC) | None (methods critique) | Not reported | Not reported | Recommends DDM/ex-Gaussian/ex-Wald; no new v reported | N/A — conclusions flip with outlier/model choice | N/A | Methodological warning, not an effect-size source |
| Alister, McKay, Sewell & Evans 2023, QJEP | Not specified (real face photos) | Gaze (cue direction; response to probed target) | None reported as graded stimulus manipulation | N=171 across 3 datasets | 139,001 trials total | Diffusion model / LBA, individual + hierarchical Bayesian | → primarily non-decision-time / orienting; v least-included (3–13%) | Not reported numerically; v has lowest inclusion probability (3–13%) | Largest trial count of any paper reviewed |
| Palmer, Caruana & Clifford 2018, R Soc Open Sci | Not specified (own gaze photos, not RaFD) | Gaze (left/right, direct 2AFC) | Head angle × eye/pupil angle independently varied (natural graded axis) | Not reported (schizophrenia vs. control groups) | Not reported | None — threshold-only, no DDM | N/A | N/A — no drift-rate parameter exists in this paper | Per-participant thresholds (degrees); task-structure template only |
Counted across all 19 target papers (White 2010PDF retained in the denominator despite failing C2, since it is on the named list). Every count is traceable to a specific paper above; nothing is estimated.
Modal choice: mostly unspecified. Most corpus papers do not name a validated stimulus database in a form this KB captured. RaFD is not novel to this corpus: our anchor paper, Williams et al. 2023, uses RaFD (Study 1) and RADIATE (Study 2). What is genuinely novel is not the database but the combination — RaFD + graded (non-binary) masking + strict 2AFC + full v/a/t₀/z fit — and the graded-masking severity itself (0/19; see below).
Modal choice: emotion (12/19, 63%). Emotion outnumbers gaze 6:1 among papers that actually judge the face's own content — the quantitative basis for the V1/V2 power imbalance (see §4.3 below).
Zero papers in this corpus report EZ-diffusion, despite it being the estimator this KB's best-practices docs recommend as the per-subject baseline. Of the 6 papers that name a method, HDDM is explicitly named in 3 and a further 1 uses a custom hierarchical-Bayesian fit — so hierarchical/Bayesian methods are the plurality (4 of 6 named cases, HDDM alone the single most common at 3/19). This supports HDDM as the confirmatory half of our EZ+HDDM pairing, even though EZ itself has no face-DDM precedent.
| Quantity | Value |
|---|---|
| Papers reporting a numeric total N | 13/19 (Brennan 2020 N=90; Brennan 2021 N=92; Williams 2023PDF N=228/264; Ozturk 2024 N=55; Nagrodzki 2025 N=134; Klein & Todd 2024PDF N=546; Nan 2024PDF N=120; Schreiber 2025PDF N=86; Haller 2024 N=44; Schrader 2023 N=40; Maka 2023 N=52; Alister 2023 N=171) |
| N range | 40 (Schrader 2023) to 546 (Klein & Todd 2024PDF); median ≈ 92 across the 13 reported values |
| Papers reporting a trials/condition or total-trials figure | 1/19: Alister 2023 (139,001 trials total, N=171, not broken down per condition) |
| Papers with no trial-count figure at all | 18/19 |
Modal choice: trial counts are essentially unreported (18/19). The ≥100/cell (EZ) and ~20–40/cell (HDDM) benchmarks used in the power analysis come entirely from best-practices docs' independent RDM/psychophysics literature, not from a count taken across these 19 papers.
Grounded in the §3.5 counts, not assumption: EZ has zero precedent here and only 3/19 papers explicitly name HDDM, so the recommendation to pair EZ-diffusion (screen) + HDDM (confirm) is not the modal practice in this literature — it is a choice imported from best-practices docs drawing on the RDM literature (Wagenmakers 2007PDF for EZ, Wiecki 2013PDF for HDDM). Trial counts cannot be empirically derived from this corpus at all (only 1/19 reports any figure); the ≥100/cell and ~20–40/cell benchmarks come from independent RDM/psychophysics sources. Total N (13/19 reported, range 40–546, median ≈92) is the only recruitment-relevant number this literature actually supports empirically.
Emotion axis: 12/19 papers (63%), spanning clinical, developmental, pharmacological, and EEG designs, with two independent pre-registered studies and a mechanistically controlled anchor. Gaze axis: 2/19 papers (11%) — and neither is a full match: Alister 2023 has the best power in the entire corpus but studies gaze-cueing, not direct discrimination; Palmer & Clifford 2018 has the right task structure but no DDM fit whatsoever. There is no paper in this corpus that is both adequately powered and structurally a direct gaze-direction 2AFC with a DDM fit. This is the quantitative basis for calling gaze the weaker-anchored, higher-novelty arm.
Anchor the pipeline on Williams 2023PDF (masking→v, the strongest pre-registered emotion-side precedent) and Alister 2023 (largest-N gaze-side DDM precedent, cueing caveat) for seeding priors, use EZ+HDDM as the estimator pair, and treat the RaFD + graded-masking + strict-2AFC + full-v/a/t0/z combination as novel rather than a replication — no paper in this corpus has run that exact combination on either axis. Budget substantially more power/pilot data for the gaze axis than the emotion axis given the 6:1 paper-count imbalance.
Per the counts above, the following elements of our planned design are not represented by even a single paper among the 19 reviewed, and must be treated as novel methodological steps requiring first-principles justification, not as replications of face-DDM practice:
None of these gaps is disqualifying — they are precisely the gaps this design's best-practices spine (see Design §7) was written to fill using the RDM/psychophysics tradition. But they must be named as novel, not presented as if a precedent existed.