Pyroclast Systems
Back
Baseline

Methodology — Baseline

From Pyroclast Systems

The science behind Baseline · Contact sys_admin@pyroclastsystems.com

This page lists every scientific framework Baseline implements, with citations to the primary peer-reviewed source for each signal; the frameworks we deliberately do notimplement, with the validation literature explaining why; what we do and don’t claim; and the results of running Baseline against an external academic corpus — including the null result. Baseline is not a lie detector and produces a signal, not a verdict. The product is defined as much by what it excludes as by what it includes.

Your privacy is part of the method: recordings are processed on-device and never uploaded. Read exactly what is and isn’t collected in the Privacy Policy.

What Baseline reads, all on-device:

Voice

Pitch, loudness, pace, pauses, and voice quality (jitter, shimmer, HNR).

Face

Gaze direction, head pose, blink rate, and facial-action geometry.

Language

Word choice, sensory detail, hedging, and narrative structure.

What Baseline does and doesn’t claim

The honest-accuracy disclosure · this is a signal, not a verdict

Claims we make

  • Baseline computes signals from your recording (vocal prosody, gaze, linguistic patterns, affect) using methods grounded in published peer-reviewed research.
  • Baseline highlights moments where multiple signals deviate from your personal baseline simultaneously, on the theory that multimodal deviation is more informative than any single channel.
  • Baseline cites the primary scientific source for every signal it extracts, including the realistic accuracy ceiling (e.g. Bond & DePaulo 2006: 54% human accuracy at deception detection).

Claims we do NOT make

  • Baseline is not a lie detector. The flagged moments are discussion prompts, not verdicts.
  • Baseline does not predict deception with high accuracy. No method, polygraph included, performs reliably above ~70% in field conditions, and most perform much worse.
  • Baseline does not assign meaning to gaze direction, microexpressions as a deception cue, or voice stress. These claims have failed validation in the literature; we explicitly exclude them.
  • Baseline output is not appropriate as evidence in legal, hiring, or other adverse-decision contexts. Use it as a self-coaching tool or as the basis for further conversation — never as a determination.

Anchor: realistic accuracy ceiling

The realistic ceiling for any deception-detection method is the Bond & DePaulo (2006) meta-analysis result: 54% human accuracy across 206 documents and 24,483 judges. Methods claiming substantially better than this in field conditions should be treated with skepticism. (Bond & DePaulo, 2006. Accuracy of deception judgments. Personality and Social Psychology Review, 10(3), 214–234.)

Accuracy ceiling for deception detection. Humans average 54% (Bond & DePaulo, 2006); no method reliably exceeds ~70% in field conditions. Any product claiming much more should be treated with skepticism.

Methods we include

8 frameworks · 17 primary citations · each grounded in published research

Reality Monitoring

Language

Tracks the balance of perceptual detail (sights, sounds, locations, times) versus cognitive operations (inference, reasoning, hedging). Recalled real events lean perceptual; constructed narratives lean cognitive.

Reality Monitoring is the foundational psychological framework distinguishing memories of perceived events from memories of imagined or fabricated events. Real-event memories carry richer sensory detail, spatial anchoring, and temporal markers; constructed memories rely more on inferential framing (“I think”, “I figured”, “must have”). Baseline computes per-segment and session-level perceptual and cognitive densities, plus the RM ratio (perceptual / total) — the most-replicated single metric in the deception-cue literature.

Features computed

sensory_densityspatial_densitytemporal_densitycognitive_operation_densityrm_perceptual_scorerm_cognitive_scorerm_ratio

Honest accuracy note

Sporer (1997) and Masip et al. (2005) report 65–75% lab-condition accuracy for RM-based classifiers. Field generalization is weaker. Use as one input among many — never as a verdict.

Primary sources (3)

Criteria-Based Content Analysis (automatable subset)

Language

Implements the automatable subset of the 19 CBCA criteria — quantity of details, contextual embedding, reproduction of conversation, and memory-gap admissions. Designed for forensic statement assessment.

CBCA is a 19-criterion framework developed for evaluating the credibility of child-witness testimony in legal contexts. Several of the criteria can be approximated automatically: detail counts, verifiable entities, spatial-temporal anchoring (contextual embedding), reported-speech rate, and admissions of memory gaps. We compute these as composites and surface them as discussion prompts. Other CBCA criteria — logical structure, unstructured production, unusual details, accurately reported peripheral details — require expert human judgment and are not automated. We are explicit that this is the automated subset, not a CBCA verdict.

Features computed

specificity_densityverifiability_densitycbca_details_scorecbca_admissions_scorecbca_conversation_score

Honest accuracy note

Vrij’s 2005 review of 37 CBCA studies reports ~73% accuracy with trained human coders following the full 19-criterion protocol plus Validity Checklist. Automated subset accuracy is not separately validated and should be treated as informational only — never as a legal or hiring determination.

Primary sources (2)

Linguistic deception cues (Newman / Pennebaker)

Language

Tracks first-person pronoun usage, self-references, hedge density, and tense distribution — features that correlate with deceptive communication in published LIWC-based studies.

The Newman/Pennebaker line of work used LIWC (Linguistic Inquiry and Word Count) to identify language patterns that distinguished deceptive from truthful written and spoken statements. Notable findings: liars use fewer first-person singular pronouns (distancing), more negative-emotion words under stress, more third-person pronouns, and present a less complex self-narrative. Baseline computes the relevant per-100-word densities and the first-person self-reference ratio.

Features computed

self_reference_ratiopronoun_first_singularpronoun_thirdhedge_density

Honest accuracy note

Newman et al. (2003) classifier reached ~67% accuracy in their sample. DePaulo et al.’s (2003) meta-analysis: average effect sizes of individual linguistic cues are small (Cohen’s d ≈ 0.1–0.3).

Primary sources (2)

Vocal prosody (Banse-Scherer affect + load proxies)

Voice

Tracks fundamental frequency (pitch), pitch variability, intensity, speaking rate, and pause structure. Used both for affect inference and as a cognitive-load proxy.

Banse & Scherer (1996) established that vocal prosody encodes emotional state in patterns predictable from arousal × valence. Baseline tracks the canonical features (F0 mean, F0 SD, intensity, speaking rate) and matches their session profile against the Banse-Scherer prototype for each of 14 emotional states. Independently, abnormally high pitch and reduced pitch variability have been associated with cognitive load — the work of telling a lie or recalling under stress (Vrij 2008, Ekman 2009).

Features computed

f0_hz_meanf0_hz_stdintensity_db_meanspeaking_ratepause_density

Honest accuracy note

Banse-Scherer emotion classification: 50–60% accuracy across 14 emotional states (random baseline 7%). Vocal cues to deception specifically have weaker effect sizes than affect inference.

Primary sources (2)

Visual signals (gaze stability, head pose, facial action)

Face

Tracks gaze direction, head yaw/pitch/roll, blink rate, and basic facial action units. Reports per-segment deviations from the personal baseline — never assigns meaning to gaze direction.

Baseline tracks visual signals via MediaPipe Face Mesh: gaze x/y, head pose, blink rate, mouth-open ratio. We track deviation from the speaker’s own baseline and surface that as a discussion prompt. Crucially, we do NOT assign deception meaning to gaze direction — the popular claim that “looking up-and-right means lying” (NLP-derived eye-accessing-cues theory) was directly tested by Wiseman et al. (2012) and failed validation.

Features computed

gaze_x_stdgaze_y_stdhead_yaw_stdhead_pitch_stdblink_rate

Honest accuracy note

Visual deviation from personal baseline is informative as one input among many. Single-feature visual deception classifiers perform poorly. Bond & DePaulo (2006) meta-analysis: humans average 54% accuracy at deception detection from any cue type.

Primary sources (2)

Personal-baseline deviation scoring

Scoring

All flagged moments are computed as deviations from the speaker’s own running statistics, not from population norms. The signal that matters is “this is unusual for you”.

Population norms for behavioral signals are noisy because of huge between-speaker variation. A given person’s baseline pitch, speaking rate, hedge density, and gaze stability are far more stable across sessions than any cross-population norm. Baseline accumulates running mean and standard deviation per feature per profile, scores each new session relative to that baseline, and flags moments where multiple features simultaneously exceed thresholds. This “multimodal cross-channel deviation” approach is well-grounded in the broader signal-detection literature even though its application to deception is still novel.

Honest accuracy note

The personal-baseline approach addresses one of the largest sources of error in single-shot lie detection (between-person variability) but does not solve within-person variability — stress, fatigue, topic, and recording conditions still introduce noise. Always interpret deviations as discussion prompts.

Primary sources (1)

Cognitive Load practice protocols (Vrij)

Practice

A practice-mode prompt category that exercises Vrij’s cognitive-load techniques: reverse recall, dual-task narration, and unanticipated questioning.

Vrij’s cognitive-load approach is one of the few interview-based deception-detection protocols with replicated effect sizes meaningfully above chance. The intuition: lying takes more cognitive effort than truth-telling, and you can amplify the difference by imposing additional cognitive load — asking subjects to narrate in reverse chronological order, perform a secondary task while speaking, or answer unanticipated angles on the same event. Baseline implements this as a practice-mode prompt category so users can self-administer the protocol and observe their own signal under load.

Honest accuracy note

Vrij et al. (2008) reverse-recall protocol: 71% accuracy in a lab study, vs. 56% under standard interviewing — meaningful but still imperfect. Effect sizes hold across replications; field generalization remains an open question.

Primary sources (2)

Cross-session narrative consistency

Language

Compares a session against earlier sessions answering the same Practice prompt. Surfaces recurring named entities, behavioral feature drift across takes, and structural changes in narrative shape. Designed as self-coaching feedback, not a contradiction detector.

Vrij and Granhag’s work on interrogation-based deception detection established that cross-session statement consistency is among the more robust verbal cues — when a narrator can produce stable central facts while showing the natural surface variation expected of authentic memory, that pattern differs from rehearsed-and-rigid recall and from inconsistent-and-improvised recall. Baseline links sessions through their Practice prompt_id: when a user records the same prompt 2+ times, we compute named-entity overlap and per-feature drift across takes (RM ratio, hedge density, specificity, etc.). The framing is neutral and self-coaching-oriented: “here is what recurred” and “here is how your signals shifted”, never “CONTRADICTION DETECTED”. Adversarial framing belongs to interrogation contexts that Baseline is explicitly not.

Honest accuracy note

Cross-session consistency is informational, not a verdict. Granhag & Hartwig (2008) showed strategic-evidence-disclosure interview techniques that exploit between-session inconsistency achieve 60–70% lie-detection accuracy in trained-interrogator lab studies — but those involve adversarial questioning, not self-coaching. The self-coaching application implemented here has no published accuracy benchmark because it’s a different use case. Treat the output as a tool for noticing your own narrative patterns, not as evidence about anyone else.

Primary sources (3)

Methods we deliberately exclude

6 frameworks · failed-validation literature cited · a trust signal

Several widely-marketed deception-detection methods have failed empirical validation. Including them would compromise Baseline’s commitment to citing primary peer-reviewed sources for every signal. We document them here so you know what we’re not doing and why.

Scientific Content Analysis (SCAN)

not used

A pronoun-and-verb-tense analysis system marketed for credibility assessment in law-enforcement contexts.

Why we exclude it

SCAN has been repeatedly tested under controlled conditions and fails to discriminate truth from deception above chance. Despite decades of commercial marketing and use by some agencies, the validation literature is consistently negative. Adding SCAN would compromise Baseline’s commitment to citing primary peer-reviewed sources for every signal.

Failed-validation literature (2)

Voice Stress Analysis (CVSA, LVA, etc.)

not used

Commercial “microtremor” detection systems claiming to identify deception from voice stress signatures.

Why we exclude it

Multiple government-funded validation studies — including National Institute of Justice and Department of Defense reviews — have found voice-stress products to perform at or near chance for deception detection. The underlying premise that “stress” uniquely indicates deception conflates arousal with deception and is not supported. Including voice stress would put Baseline in the same category as discredited polygraph alternatives.

Failed-validation literature (3)

Microexpression-based deception detection

not used

The popularized claim — associated with Paul Ekman’s work and the TV show “Lie to Me” — that 1/25-second facial expressions reliably reveal concealed emotion and, by extension, deception.

Why we exclude it

Microexpressions exist as a phenomenon, but the specific claim that they reliably indicate deception (as opposed to general emotional leakage) does not replicate well under controlled study. Effect sizes for microexpression-based deception classification are far smaller than the popular framing implies, and human judges trained on microexpressions do not reliably outperform untrained controls at detecting lies. We do not implement microexpression-based deception inference. We do track facial action units as one input into multimodal baseline-deviation scoring, which is a different and more defensible claim.

Failed-validation literature (2)

Polygraph-style autonomic-arousal deception detection

not used

Multi-channel autonomic measurement (heart rate, GSR, respiration) interpreted via Comparison Question Technique or Concealed Information Test to infer deception.

Why we exclude it

Polygraph-based deception inference rests on the contested premise that arousal under specific question types reliably indicates deception. The 2003 National Research Council report found controlled-condition accuracy substantially below claims, and polygraph evidence is inadmissible in most US court systems. Baseline does not infer deception from autonomic measures. If biometric integration is added in a future phase, it will be framed strictly as an arousal/deviation signal contributing to multimodal baseline scoring — not as a lie indicator.

Failed-validation literature (2)

Training on adjudicated interrogation video

not used

An attractive shortcut: train a classifier on confessed-and-convicted interrogation footage to learn “real” deception cues.

Why we exclude it

Three independent reasons not to do this. (1) Ground-truth labels are contaminated — confession and conviction both include known false-positive rates (the Innocence Project documents ~25% of DNA-exoneration cases involved confessed-then-exonerated defendants). (2) Severe selection bias — interrogation-video subjects are stressed under custodial conditions; a classifier trained on them learns custodial-stress signatures, not deception. (3) Existing attempts (Pérez-Rosas Real-Life Trial corpus, DARPA CROAD) achieve 60–75% in-sample accuracy with poor out-of-sample generalization. The marketing appeal is strong; the science is weak; the brand cost of being wrong is high.

Failed-validation literature (3)

Neuro-Linguistic Programming eye-accessing cues

not used

The folk claim that gaze direction (e.g. “looking up and to the right”) reveals deception.

Why we exclude it

Tested directly by Wiseman et al. (2012) and found to have no relationship with deception. Baseline does track gaze stability as a baseline-deviation signal but does NOT assign deception meaning to gaze direction itself. We cite the Wiseman study in our visual-feature documentation specifically to make the non-claim explicit.

Failed-validation literature (1)

External validation

Results against academic corpora — including null and negative results

Empirical results from running Baseline against academic deception corpora — including null and negative results, published as a structural commitment to honesty.

Pérez-Rosas Real-Life Trial corpus (RLT)

null result

Pérez-Rosas, Abouelenien, Mihalcea & Burzo (2015). Deception detection using real-life trial data. Proceedings of ICMI ’15. · Run 2026-05-05

−0.024calibration r
0.554balanced acc
0.498ROC AUC

Sample: 121 clips (61 lie, 60 truth) across 56 unique speakers. Metrics: calibration r = −0.024 (95% CI [−0.226, 0.133], p = 0.797); balanced accuracy = 0.554; ROC AUC = 0.498.

Headline: On a one-shot stranger-evaluation corpus with no per-speaker baseline, Baseline does not produce above-chance deception classification.

What this did NOT test:Baseline is designed for intra-personal deviation analysis — how a single individual’s signals deviate from their own established baseline across multiple sessions. The RLT corpus has no prior baseline data and typically only 1–3 clips per speaker, so the personal-baseline engine had nothing to score against and fell back to population-level scoring. This validation tested the wrong job for the tool.

What we learned: Several signals correlated in the opposite direction the literature predicts for personal-baseline deviation — filler density (r = −0.17), CBCA detail markers (r = −0.16), perceptual specificity (r = −0.13), and sensory density (r = −0.14) all skewed toward the lie class. This is consistent with the rehearsal-effect literature: high-stakes courtroom liars are typically coached by counsel, producing more polished accounts than genuinely-distressed truthful witnesses. It underlines why personal-baseline anchoring matters.

Why we publish it: Hiding a null result would be the same methodological dishonesty we explicitly call out in the polygraph literature — and the result is itself an argument for why personal baselining is the right design. Cross-corpus validation of the personal-baseline use case requires longitudinal data with multiple sessions per speaker, which RLT does not provide.