🚦 English documentation
日本語 →

🚦Verdict Logic — The Big Picture

L1 deterministic gates → the scoring world → thresholds → the 3-way branch, the asymmetric termination valve, and the mainline/observation split — in three diagrams.

Three diagrams show how Aegis decides whether to hang up, ring a person, or let the AI answer. Every number (0.8 / 0.3 / 0.6 / 0.95) was verified against the code on 2026-07-21. These are structure diagrams only — no performance figures or operational guarantees (How to Read This Spec).

Diagram 1: The full verdict flow

Stage 1 — at ring time (pre-call): route-decision (never looks at a score) Incoming call L1 deterministic gates (checked in order) No caller ID (if enabled) → block Blocklist (confirmed) → block Whitelist → human Protection OFF → human ⚑ independent lane — no score exists here block → announce, hang up [company part] → [body A–D], played in sequence human → straight to a person ai → AI receptionist (default) Stage 2 — during the call (ai only): the rules layer (Scorer via verdict port) settled WITHOUT a score (short-circuit on hit) blocklist → blocked whitelist → human hours → after_hours auth KW → human (0.0) the scoring world KW score high 0.3 / med 0.15 / low 0.05 opener floor score = max(score, 0.6) category hit → settled before thresholds Threshold verdict block ≥ 0.8 grey 0.3 < s < 0.8 → human pass ≤ 0.3 → human ⚑ fraud enum: when classify_call accepts "fraud", spam_score = 0.95 (constant). 0.95 ≥ valve 0.8, so termination derives — the only LLM-sourced score injection.
Diagram: the full verdict flow in two stages. Stage 1 fires at ring time: route-decision checks the L1 deterministic gates in order (no caller ID if enabled / blocklist / whitelist / protection flag) and branches block / human / ai — no score exists in this lane. Announcements ([company part] → [body A–D]) ride only on block. Stage 2 fires during the call, for ai calls only: the Scorer first runs its own score-free gates (blocklist → whitelist → business hours → auth keywords; a hit settles the action immediately), then enters the scoring world: keyword score → opener floor (raises to 0.6, below the 0.8 valve) → category (a hit settles before thresholds) → threshold verdict (block ≥ 0.8 / grey 0.3 < s < 0.8 to a human / pass ≤ 0.3 to a human). The fraud enum is the only LLM-sourced score: 0.95, above the valve. Structure only — verified against code on 2026-07-21.

Diagram 2: The asymmetry of the safety valve

The road to hangup (narrow) ① action is terminate-class form_guided ② spam_score ≥ 0.8 (valve) AEGIS_TERMINATE_MIN_SCORE AND should_terminate = true sole exception: blocked / after_hours (L1 gates) terminate unconditionally — the valve does not apply The road to a person (wide, fail-open) empty transcript → pending internal error → error score < 0.8 (grey / pass) transfer-number lookup fails → always to a person “When in doubt, don’t hang up” hanging up = two conditions ANDed / doubt & failure = always a person (guaranteed by structure)
Diagram: the asymmetry of the termination valve. Left (narrow road): should_terminate becomes true only when BOTH conditions hold — the action is terminate-class (form_guided) AND spam_score ≥ 0.8 (the valve, AEGIS_TERMINATE_MIN_SCORE). The only exception is the L1 deterministic gates (blocked / after_hours), which terminate unconditionally and sit outside the valve. Right (wide road): an empty transcript (pending), an internal error (error), any score under 0.8, or a failed transfer-number lookup — all of them fall open to a person. Verified against verdict_port.py on 2026-07-21.

Diagram 3: Mainline vs. observation

Mainline (decides) call verdict port (champion = rules) action / should_terminate respond: hang up or to a person solid arrows = affects the verdict Observation (records only) at the settle point (transcript-final), only when AEGIS_VERDICT_MODE=shadow (default rules = no recording) rules re-run alongside LLM shadow_* arc score (Peak + Accumulation) arc_* verification-layer trigger watch verification_* 3 independent try/except — one failure never drags down another, nor the call record call_logs (recording columns) × ⚑ observation NEVER touches the verdict — no wiring into verdict / should_terminate / route-decision
Diagram: the mainline and observation are structurally separated. Left: the mainline decides — call → verdict port (champion = rules) → action / should_terminate → respond. Right: at the settle point (transcript-final), and only when AEGIS_VERDICT_MODE=shadow (the default is rules, meaning no recording), three collectors write records only: the rules re-run recorded beside the LLM value (shadow_* columns), the conversation arc score Peak + Accumulation (arc_* columns), and the verification-layer trigger watch (verification_* columns). Each sits in its own try/except so one failure never drags down another or the call record itself. The crossed-out arrow is the point of the diagram: observation has no wiring into verdict, should_terminate, or route-decision. Verified against realtime_session.py / arc_score.py / verification_trigger.py on 2026-07-21.

The thinking behind the shape

  • L1 never looks at a score — obvious callers (listed numbers, after-hours calls) need no computation. Deterministic gates are fast, free, and their reasons are unambiguous. Only callers the lists cannot settle enter the scoring world.
  • The cascade principle — easy calls are settled instantly by rules; only the grey zone goes to the AI (LLM). Money and time are spent only on the calls worth doubting (same principle as Diagrams 3–4 in Getting Started).
  • The shadow principle — every new measuring stick (LLM verdicts, the arc score, the verification-layer watch) starts as observation before it is allowed to influence decisions. Growing the recording columns now, at zero real calls, means data accumulates from day one on real devices.
  • FPR discipline — measurements are never stated as point estimates; they carry Wilson 95% confidence intervals. The numbers live in Detection Logic & Accuracy.
ℹ️Source code (canonical)

Verified against aegis-platform backend/engine/scorer.py (steps 1–7), verdict_port.py (should_terminate derivation, fail-open), opener_floor.py (opener floor 0.6), backend/api/route_decision.py (3-way branch, announcement resolution), backend/api/realtime_session.py (the three shadow recorders at the settle point), and backend/services/tools_core.py (fraud enum 0.95). Settled values change in the SSoT first.