Methodology — Baseline
From Pyroclast Systems
The science behind Baseline · Contact sys_admin@pyroclastsystems.com
This page lists every scientific framework Baseline implements, with citations to the primary peer-reviewed source for each signal; the frameworks we deliberately do notimplement, with the validation literature explaining why; what we do and don’t claim; and the results of running Baseline against an external academic corpus — including the null result. Baseline is not a lie detector and produces a signal, not a verdict. The product is defined as much by what it excludes as by what it includes.
Your privacy is part of the method: recordings are processed on-device and never uploaded. Read exactly what is and isn’t collected in the Privacy Policy.
What Baseline reads, all on-device:
Voice
Pitch, loudness, pace, pauses, and voice quality (jitter, shimmer, HNR).
Face
Gaze direction, head pose, blink rate, and facial-action geometry.
Language
Word choice, sensory detail, hedging, and narrative structure.
What Baseline does and doesn’t claim
The honest-accuracy disclosure · this is a signal, not a verdict
Claims we make
- Baseline computes signals from your recording (vocal prosody, gaze, linguistic patterns, affect) using methods grounded in published peer-reviewed research.
- Baseline highlights moments where multiple signals deviate from your personal baseline simultaneously, on the theory that multimodal deviation is more informative than any single channel.
- Baseline cites the primary scientific source for every signal it extracts, including the realistic accuracy ceiling (e.g. Bond & DePaulo 2006: 54% human accuracy at deception detection).
Claims we do NOT make
- Baseline is not a lie detector. The flagged moments are discussion prompts, not verdicts.
- Baseline does not predict deception with high accuracy. No method, polygraph included, performs reliably above ~70% in field conditions, and most perform much worse.
- Baseline does not assign meaning to gaze direction, microexpressions as a deception cue, or voice stress. These claims have failed validation in the literature; we explicitly exclude them.
- Baseline output is not appropriate as evidence in legal, hiring, or other adverse-decision contexts. Use it as a self-coaching tool or as the basis for further conversation — never as a determination.
Anchor: realistic accuracy ceiling
The realistic ceiling for any deception-detection method is the Bond & DePaulo (2006) meta-analysis result: 54% human accuracy across 206 documents and 24,483 judges. Methods claiming substantially better than this in field conditions should be treated with skepticism. (Bond & DePaulo, 2006. Accuracy of deception judgments. Personality and Social Psychology Review, 10(3), 214–234.)
Accuracy ceiling for deception detection. Humans average 54% (Bond & DePaulo, 2006); no method reliably exceeds ~70% in field conditions. Any product claiming much more should be treated with skepticism.
Methods we include
8 frameworks · 17 primary citations · each grounded in published research
Reality Monitoring
LanguageTracks the balance of perceptual detail (sights, sounds, locations, times) versus cognitive operations (inference, reasoning, hedging). Recalled real events lean perceptual; constructed narratives lean cognitive.
Reality Monitoring is the foundational psychological framework distinguishing memories of perceived events from memories of imagined or fabricated events. Real-event memories carry richer sensory detail, spatial anchoring, and temporal markers; constructed memories rely more on inferential framing (“I think”, “I figured”, “must have”). Baseline computes per-segment and session-level perceptual and cognitive densities, plus the RM ratio (perceptual / total) — the most-replicated single metric in the deception-cue literature.
Features computed
Honest accuracy note
Sporer (1997) and Masip et al. (2005) report 65–75% lab-condition accuracy for RM-based classifiers. Field generalization is weaker. Use as one input among many — never as a verdict.
Primary sources (3)
- Johnson & Raye (1981). Reality monitoring. Psychological Review, 88(1), 67–85. — Originated the perceptual-vs-cognitive distinction.
- Sporer (1997). The less travelled road to truth: Verbal cues in deception detection. Applied Cognitive Psychology, 11. — RM ratio was the most consistently replicated single feature across deception experiments.
- Masip et al. (2005). The detection of deception with the reality monitoring approach. Psychology, Crime & Law, 11. — RM-based classifiers achieve 65–75% in lab conditions; field generalization is lower.
Criteria-Based Content Analysis (automatable subset)
LanguageImplements the automatable subset of the 19 CBCA criteria — quantity of details, contextual embedding, reproduction of conversation, and memory-gap admissions. Designed for forensic statement assessment.
CBCA is a 19-criterion framework developed for evaluating the credibility of child-witness testimony in legal contexts. Several of the criteria can be approximated automatically: detail counts, verifiable entities, spatial-temporal anchoring (contextual embedding), reported-speech rate, and admissions of memory gaps. We compute these as composites and surface them as discussion prompts. Other CBCA criteria — logical structure, unstructured production, unusual details, accurately reported peripheral details — require expert human judgment and are not automated. We are explicit that this is the automated subset, not a CBCA verdict.
Features computed
Honest accuracy note
Vrij’s 2005 review of 37 CBCA studies reports ~73% accuracy with trained human coders following the full 19-criterion protocol plus Validity Checklist. Automated subset accuracy is not separately validated and should be treated as informational only — never as a legal or hiring determination.
Primary sources (2)
- Steller & Köhnken (1989). Criteria-Based Content Analysis. In Raskin (ed.), Psychological Methods in Criminal Investigation. — Original 19-criterion framework; detail criteria most predictive.
- Vrij (2005). Criteria-Based Content Analysis: A qualitative review. Psychology, Public Policy, and Law, 11. — 73% accuracy with trained coders; partial implementations score lower; CBCA alone is insufficient for legal determination.
Linguistic deception cues (Newman / Pennebaker)
LanguageTracks first-person pronoun usage, self-references, hedge density, and tense distribution — features that correlate with deceptive communication in published LIWC-based studies.
The Newman/Pennebaker line of work used LIWC (Linguistic Inquiry and Word Count) to identify language patterns that distinguished deceptive from truthful written and spoken statements. Notable findings: liars use fewer first-person singular pronouns (distancing), more negative-emotion words under stress, more third-person pronouns, and present a less complex self-narrative. Baseline computes the relevant per-100-word densities and the first-person self-reference ratio.
Features computed
Honest accuracy note
Newman et al. (2003) classifier reached ~67% accuracy in their sample. DePaulo et al.’s (2003) meta-analysis: average effect sizes of individual linguistic cues are small (Cohen’s d ≈ 0.1–0.3).
Primary sources (2)
- Newman, Pennebaker, Berry & Richards (2003). Lying words: Predicting deception from linguistic styles. Personality and Social Psychology Bulletin, 29. — 67% classification accuracy from pronoun + emotion-word + complexity features.
- DePaulo et al. (2003). Cues to deception. Psychological Bulletin, 129(1), 74–118. — Meta-analysis of 158 cues across 120 studies; average effect sizes small but reliable.
Vocal prosody (Banse-Scherer affect + load proxies)
VoiceTracks fundamental frequency (pitch), pitch variability, intensity, speaking rate, and pause structure. Used both for affect inference and as a cognitive-load proxy.
Banse & Scherer (1996) established that vocal prosody encodes emotional state in patterns predictable from arousal × valence. Baseline tracks the canonical features (F0 mean, F0 SD, intensity, speaking rate) and matches their session profile against the Banse-Scherer prototype for each of 14 emotional states. Independently, abnormally high pitch and reduced pitch variability have been associated with cognitive load — the work of telling a lie or recalling under stress (Vrij 2008, Ekman 2009).
Features computed
Honest accuracy note
Banse-Scherer emotion classification: 50–60% accuracy across 14 emotional states (random baseline 7%). Vocal cues to deception specifically have weaker effect sizes than affect inference.
Primary sources (2)
- Banse & Scherer (1996). Acoustic profiles in vocal emotion expression. Journal of Personality and Social Psychology, 70. — Each of 14 emotional states produces a distinct, reliably-classified acoustic profile.
- Scherer (2003). Vocal communication of emotion: A review. Speech Communication, 40. — Vocal prosody encodes affect more reliably than facial expression for some states.
Visual signals (gaze stability, head pose, facial action)
FaceTracks gaze direction, head yaw/pitch/roll, blink rate, and basic facial action units. Reports per-segment deviations from the personal baseline — never assigns meaning to gaze direction.
Baseline tracks visual signals via MediaPipe Face Mesh: gaze x/y, head pose, blink rate, mouth-open ratio. We track deviation from the speaker’s own baseline and surface that as a discussion prompt. Crucially, we do NOT assign deception meaning to gaze direction — the popular claim that “looking up-and-right means lying” (NLP-derived eye-accessing-cues theory) was directly tested by Wiseman et al. (2012) and failed validation.
Features computed
Honest accuracy note
Visual deviation from personal baseline is informative as one input among many. Single-feature visual deception classifiers perform poorly. Bond & DePaulo (2006) meta-analysis: humans average 54% accuracy at deception detection from any cue type.
Primary sources (2)
- Wiseman, Watt, Ten Brinke, Porter, Couper & Rankin (2012). The eyes don’t have it: Lie detection and Neuro-Linguistic Programming. PLOS ONE, 7(7). — Direct test of the NLP eye-accessing-cues claim. No relationship between gaze direction and deception. We cite this to make explicit what we DON’T claim.
- Bond & DePaulo (2006). Accuracy of deception judgments. Personality and Social Psychology Review, 10. — Across 206 documents and 24,483 judges: average human accuracy at deception detection is 54%. Sets the realistic ceiling.
Personal-baseline deviation scoring
ScoringAll flagged moments are computed as deviations from the speaker’s own running statistics, not from population norms. The signal that matters is “this is unusual for you”.
Population norms for behavioral signals are noisy because of huge between-speaker variation. A given person’s baseline pitch, speaking rate, hedge density, and gaze stability are far more stable across sessions than any cross-population norm. Baseline accumulates running mean and standard deviation per feature per profile, scores each new session relative to that baseline, and flags moments where multiple features simultaneously exceed thresholds. This “multimodal cross-channel deviation” approach is well-grounded in the broader signal-detection literature even though its application to deception is still novel.
Honest accuracy note
The personal-baseline approach addresses one of the largest sources of error in single-shot lie detection (between-person variability) but does not solve within-person variability — stress, fatigue, topic, and recording conditions still introduce noise. Always interpret deviations as discussion prompts.
Primary sources (1)
- DePaulo et al. (2003). Cues to deception. Psychological Bulletin, 129. — Within-person variability in cue presentation is large; cross-population norms are weak predictors.
Cognitive Load practice protocols (Vrij)
PracticeA practice-mode prompt category that exercises Vrij’s cognitive-load techniques: reverse recall, dual-task narration, and unanticipated questioning.
Vrij’s cognitive-load approach is one of the few interview-based deception-detection protocols with replicated effect sizes meaningfully above chance. The intuition: lying takes more cognitive effort than truth-telling, and you can amplify the difference by imposing additional cognitive load — asking subjects to narrate in reverse chronological order, perform a secondary task while speaking, or answer unanticipated angles on the same event. Baseline implements this as a practice-mode prompt category so users can self-administer the protocol and observe their own signal under load.
Honest accuracy note
Vrij et al. (2008) reverse-recall protocol: 71% accuracy in a lab study, vs. 56% under standard interviewing — meaningful but still imperfect. Effect sizes hold across replications; field generalization remains an open question.
Primary sources (2)
- Vrij, Mann, Fisher, Leal, Milne & Bull (2008). Increasing cognitive load to facilitate lie detection: The benefit of recalling an event in reverse order. Law and Human Behavior, 32. — Reverse-order recall amplified deception cues: 71% lie-detection accuracy vs. 56% in the standard-interview condition.
- Vrij & Fisher (2016). Which lie detection tools are ready for use in the criminal justice system? Journal of Applied Research in Memory and Cognition, 5. — Cognitive-load and Verifiability Approach techniques have the strongest empirical support among interview-based methods, though none reach polygraph-claimed accuracy levels.
Cross-session narrative consistency
LanguageCompares a session against earlier sessions answering the same Practice prompt. Surfaces recurring named entities, behavioral feature drift across takes, and structural changes in narrative shape. Designed as self-coaching feedback, not a contradiction detector.
Vrij and Granhag’s work on interrogation-based deception detection established that cross-session statement consistency is among the more robust verbal cues — when a narrator can produce stable central facts while showing the natural surface variation expected of authentic memory, that pattern differs from rehearsed-and-rigid recall and from inconsistent-and-improvised recall. Baseline links sessions through their Practice prompt_id: when a user records the same prompt 2+ times, we compute named-entity overlap and per-feature drift across takes (RM ratio, hedge density, specificity, etc.). The framing is neutral and self-coaching-oriented: “here is what recurred” and “here is how your signals shifted”, never “CONTRADICTION DETECTED”. Adversarial framing belongs to interrogation contexts that Baseline is explicitly not.
Honest accuracy note
Cross-session consistency is informational, not a verdict. Granhag & Hartwig (2008) showed strategic-evidence-disclosure interview techniques that exploit between-session inconsistency achieve 60–70% lie-detection accuracy in trained-interrogator lab studies — but those involve adversarial questioning, not self-coaching. The self-coaching application implemented here has no published accuracy benchmark because it’s a different use case. Treat the output as a tool for noticing your own narrative patterns, not as evidence about anyone else.
Primary sources (3)
- Vrij & Granhag (2007). Interviewing to detect deception. In Granhag (Ed.), Forensic Psychology in Context. Willan. — Cross-session consistency, when paired with appropriate interview design, distinguishes truthful from fabricated accounts more reliably than single-session cue analysis.
- Granhag & Hartwig (2008). A new theoretical perspective on deception detection: On the psychology of instrumental mind-reading. Psychology, Crime & Law, 14. — Strategic Use of Evidence (SUE) interview technique exploiting between-session inconsistency: 60–70% lie-detection accuracy in trained-coder lab studies.
- Granhag, Strömwall & Jonsson (2003). Partners in crime: How liars in collusion betray themselves. Journal of Applied Social Psychology, 33. — Inter-witness consistency analysis: even rehearsed liars show characteristic surface-detail divergence on unanticipated peripheral questions.
Methods we deliberately exclude
6 frameworks · failed-validation literature cited · a trust signal
Scientific Content Analysis (SCAN)
not usedA pronoun-and-verb-tense analysis system marketed for credibility assessment in law-enforcement contexts.
Why we exclude it
SCAN has been repeatedly tested under controlled conditions and fails to discriminate truth from deception above chance. Despite decades of commercial marketing and use by some agencies, the validation literature is consistently negative. Adding SCAN would compromise Baseline’s commitment to citing primary peer-reviewed sources for every signal.
Failed-validation literature (2)
- Vrij (2008). Detecting Lies and Deceit (2nd ed.), Wiley — SCAN review chapter. — Comprehensive review found no controlled study in which SCAN reliably distinguished truth from deception above chance.
- Bogaard, Meijer, Vrij & Merckelbach (2016). Scientific Content Analysis (SCAN) cannot distinguish between truthful and fabricated accounts. PLOS ONE, 11(1). — Direct test with 234 statements: SCAN-trained coders performed at chance (49.6%).
Voice Stress Analysis (CVSA, LVA, etc.)
not usedCommercial “microtremor” detection systems claiming to identify deception from voice stress signatures.
Why we exclude it
Multiple government-funded validation studies — including National Institute of Justice and Department of Defense reviews — have found voice-stress products to perform at or near chance for deception detection. The underlying premise that “stress” uniquely indicates deception conflates arousal with deception and is not supported. Including voice stress would put Baseline in the same category as discredited polygraph alternatives.
Failed-validation literature (3)
- Damphousse (2008). Voice Stress Analysis: Only 15 percent of lies about drug use detected in field test. NIJ Journal, 259. — Field test of 319 arrestees: VSA detected lies at 15%, far below polygraph and below chance for the alternative class.
- Hollien, Harnsberger, Martin & Hollien (2008). Evaluation of the NITV CVSA. Journal of Forensic Sciences, 53. — Controlled lab evaluation: CVSA performance not significantly different from chance.
- Eriksson & Lacerda (2007). Charlatanry in forensic speech science: A problem to be taken seriously. International Journal of Speech, Language and the Law, 14. — Critical review of voice-stress and lie-detection commercial products; widespread methodological flaws in vendor-supplied validation.
Microexpression-based deception detection
not usedThe popularized claim — associated with Paul Ekman’s work and the TV show “Lie to Me” — that 1/25-second facial expressions reliably reveal concealed emotion and, by extension, deception.
Why we exclude it
Microexpressions exist as a phenomenon, but the specific claim that they reliably indicate deception (as opposed to general emotional leakage) does not replicate well under controlled study. Effect sizes for microexpression-based deception classification are far smaller than the popular framing implies, and human judges trained on microexpressions do not reliably outperform untrained controls at detecting lies. We do not implement microexpression-based deception inference. We do track facial action units as one input into multimodal baseline-deviation scoring, which is a different and more defensible claim.
Failed-validation literature (2)
- Burgoon (2018). Microexpressions are not the best way to catch a liar. Frontiers in Psychology, 9. — Direct review of microexpression deception research: classification accuracy not reliably better than chance; trained-judge advantage does not replicate.
- Jordan et al. (2019). A test of the micro-expressions training tool: Does it improve lie detection? Journal of Investigative Psychology and Offender Profiling, 16. — Participants trained with the standard microexpression training tool did not detect lies better than untrained controls.
Polygraph-style autonomic-arousal deception detection
not usedMulti-channel autonomic measurement (heart rate, GSR, respiration) interpreted via Comparison Question Technique or Concealed Information Test to infer deception.
Why we exclude it
Polygraph-based deception inference rests on the contested premise that arousal under specific question types reliably indicates deception. The 2003 National Research Council report found controlled-condition accuracy substantially below claims, and polygraph evidence is inadmissible in most US court systems. Baseline does not infer deception from autonomic measures. If biometric integration is added in a future phase, it will be framed strictly as an arousal/deviation signal contributing to multimodal baseline scoring — not as a lie indicator.
Failed-validation literature (2)
- National Research Council (2003). The Polygraph and Lie Detection. National Academies Press. — Controlled-condition accuracy substantially below claims; polygraph not recommended for security screening.
- Iacono & Lykken (1997). The validity of the lie detector: Two surveys of scientific opinion. Journal of Applied Psychology, 82. — Surveys of psychophysiologists found majority view that CQT polygraph testing has “no scientific basis” for the strong claims made by practitioners.
Training on adjudicated interrogation video
not usedAn attractive shortcut: train a classifier on confessed-and-convicted interrogation footage to learn “real” deception cues.
Why we exclude it
Three independent reasons not to do this. (1) Ground-truth labels are contaminated — confession and conviction both include known false-positive rates (the Innocence Project documents ~25% of DNA-exoneration cases involved confessed-then-exonerated defendants). (2) Severe selection bias — interrogation-video subjects are stressed under custodial conditions; a classifier trained on them learns custodial-stress signatures, not deception. (3) Existing attempts (Pérez-Rosas Real-Life Trial corpus, DARPA CROAD) achieve 60–75% in-sample accuracy with poor out-of-sample generalization. The marketing appeal is strong; the science is weak; the brand cost of being wrong is high.
Failed-validation literature (3)
- Innocence Project (2024). False confessions or admissions resource page (innocenceproject.org). — Approximately 25% of DNA-exoneration cases involved a false confession or self-incriminating statement. Conviction is not a reliable ground-truth label for deception.
- Pérez-Rosas, Abouelenien, Mihalcea & Burzo (2015). Deception detection using real-life trial data. ICMI ’15. — Most-published academic dataset of adjudicated trial video. Best classifiers reach 60–75% in-sample; cross-corpus generalization is poor.
- Kassin & Gudjonsson (2004). The psychology of confessions: A review of the literature and issues. Psychological Science in the Public Interest, 5. — Comprehensive review of false-confession mechanisms; documents systematic bias toward confession from sleep-deprived, mentally ill, and coerced suspects.
Neuro-Linguistic Programming eye-accessing cues
not usedThe folk claim that gaze direction (e.g. “looking up and to the right”) reveals deception.
Why we exclude it
Tested directly by Wiseman et al. (2012) and found to have no relationship with deception. Baseline does track gaze stability as a baseline-deviation signal but does NOT assign deception meaning to gaze direction itself. We cite the Wiseman study in our visual-feature documentation specifically to make the non-claim explicit.
Failed-validation literature (1)
- Wiseman, Watt, Ten Brinke, Porter, Couper & Rankin (2012). The eyes don’t have it: Lie detection and Neuro-Linguistic Programming. PLOS ONE, 7(7). — Direct empirical test of NLP eye-accessing-cues theory. No relationship between gaze direction and deception in two pre-registered experiments.
External validation
Results against academic corpora — including null and negative results
Empirical results from running Baseline against academic deception corpora — including null and negative results, published as a structural commitment to honesty.
Pérez-Rosas Real-Life Trial corpus (RLT)
null resultPérez-Rosas, Abouelenien, Mihalcea & Burzo (2015). Deception detection using real-life trial data. Proceedings of ICMI ’15. · Run 2026-05-05
Sample: 121 clips (61 lie, 60 truth) across 56 unique speakers. Metrics: calibration r = −0.024 (95% CI [−0.226, 0.133], p = 0.797); balanced accuracy = 0.554; ROC AUC = 0.498.
Headline: On a one-shot stranger-evaluation corpus with no per-speaker baseline, Baseline does not produce above-chance deception classification.
What this did NOT test:Baseline is designed for intra-personal deviation analysis — how a single individual’s signals deviate from their own established baseline across multiple sessions. The RLT corpus has no prior baseline data and typically only 1–3 clips per speaker, so the personal-baseline engine had nothing to score against and fell back to population-level scoring. This validation tested the wrong job for the tool.
What we learned: Several signals correlated in the opposite direction the literature predicts for personal-baseline deviation — filler density (r = −0.17), CBCA detail markers (r = −0.16), perceptual specificity (r = −0.13), and sensory density (r = −0.14) all skewed toward the lie class. This is consistent with the rehearsal-effect literature: high-stakes courtroom liars are typically coached by counsel, producing more polished accounts than genuinely-distressed truthful witnesses. It underlines why personal-baseline anchoring matters.
Why we publish it: Hiding a null result would be the same methodological dishonesty we explicitly call out in the polygraph literature — and the result is itself an argument for why personal baselining is the right design. Cross-corpus validation of the personal-baseline use case requires longitudinal data with multiple sessions per speaker, which RLT does not provide.

