🎯 English documentation
日本語 →

🎯Single Source of Truth (SSoT)

Confirmed cross-repo facts (authoritative): target OS / Asterisk version / ports / verdict path. All repos reference this page.

🕒 Last updated: 2026-07-14 (JST)

This page is the Single Source of Truth for the whole Aegis project (aegis-platform / aegis-sip-bridge / voice-edge / aegis-docs). To keep the target OS, Asterisk version, and similar facts from drifting apart across repos (cross-repo drift), every confirmed fact is consolidated here.

ℹ️Operating rules
  • This page is authoritative. Each repo’s README / docs / code comments reference this page instead of hard-coding values.
  • When a value changes, update this page first, then propagate to the repos (never the other way around).
  • Items marked “being verified on device” are not treated as confirmed (see the ⚠️ note below).

Confirmed facts (authoritative)

ItemConfirmed valueNotes / background
Target OSRaspberry Pi OS Bookworm Lite 64bit (aarch64)Chris’s choice; unified on aarch64
Asterisk22 LTS (source build)Built from source, not the packaged version
Python (Pi host)3.11 (Bookworm default)Uses the OS default as-is
Python (bridge)3.12 planned ⚠️Being verified on device (see note below)
Primary voice routeRoute C (OpenAI Realtime GA)The main line for AI responses. voice-edge A22 is its telephony foundation (below)
Realtime default modelgpt-realtime-2.1Decided by Okada on 2026-09-02 (bridge src/bridge.py:173 AEGIS_REALTIME_MODEL_DEFAULT). Chosen to match the default voice marin. Overridable via env AEGIS_REALTIME_MODEL (2 stages: (1) env, (2) default; there is no file stage). ⚑ We do NOT claim that plain gpt-realtime and gpt-realtime-2.1 are the same model under different names (all that was measured is that their default voices differ; plain defaults to alloy)
Primary telephony foundationvoice-edge A22The main line. See the comparison table below
AudioSocket port9092Asterisk ↔ bridge.py
Post-call data recordingPOST /api/realtime/transcript-final (once, after the call ends)Saves the full conversation + LLM verdict to call_logs. The main line for training data
Verdict source for spam_scorecall_logs.score_source = rules / llm / llm_conversationMultiple verdicts can land in the same score column; the column tells them apart (guards against verdict duplication) + precedence (routing = live truth / recording & training = full-conversation verdict truth; below)
Verdict termination (mid-call)No automatic hangup on a conversation verdict (slated for removal; final design 2026-09-16)The bridge’s AMI Hangup (D-1) code remains but conversation_may_terminate() is always False, so it never fires (br=a557a5b1d41d). The AstDB DBPut path was rejected (ADR-0001)
Termination safety valveAEGIS_TERMINATE_MIN_SCORE=0.8 (bridge side: remnant of an unreachable path; slated for removal)The should_terminate derived by the platform verdict port stays for recording and the browser demo; it is not used to cut phone calls
route-decision (pre-call)curl (separate layer)Route selection at call arrival. Distinct from the verdict path
Entry point to the classifier brainverdict port (single contract)Every path calls the port. should_terminate is a derived value from action + the safety valve (see “verdict port contract” below)
⚠️Being verified on device (not confirmed)

Python 3.12 for the bridge has not finished on-device verification. Until it does, treat it as “3.12 targeted, being verified on device” and do not state 3.12 as settled in any repo. (Python 3.11 on the Pi host IS confirmed, since it is the Bookworm default.)

The main line (A22) and the frozen design

How A22 and Route C relate: voice-edge A22 is the telephony foundation (lines, Asterisk, AudioSocket), while Route C (OpenAI Realtime GA, default model gpt-realtime-2.1) is the main line for AI responses. They are separate layers: Route C runs on top of A22.

voice-edge A22 (main line, active) v2 standalone (frozen)
Status Main line, under active development Frozen (reference only)
OS Raspberry Pi OS Bookworm 64bit (aarch64) Ubuntu 24.04
Asterisk 22 LTS (source build) 20
Treatment Implement against this as the truth Not used for new references or new implementation
📌What “frozen” means

v2 standalone (Ubuntu 24.04 / Asterisk 20) is frozen. It remains as historical material, but is not used as the basis for new implementation, documentation, or estimates. Going forward, the only truth is voice-edge A22.

The verdict flow (don’t mix up the paths)

The path that returns the verdict and the pre-call route decision (route-decision) are separate layers. Do not conflate them.

  • Pre-call — route-decision: decides which route handles an incoming call. Queried via curl (HTTP). The response’s decision has 4 values (block / human / ai / take_message). ★Not the same thing as action (the verdict vocabulary) — different words, different layer. ⚑ Besides the values there are 2 other cases (not values): empty string = no decision could be made → the dialplan falls to the static fallback and logs a WARNING (voice-edge 20-from-hikari.conf:161-162, ve=e94704495635) / an unknown value outside the 4 → hand to a human (same file :167-168, ve=e94704495635). ⚑ In the final design (2026-09-16), empty / invalid / 401 / timeout ring the fixed phone once instead of the static fallback (stage 1 E01, not implemented). ★Unless these 2 cases are written next to the 4 values, the site reads it as “only 4 values ever arrive”. ⚑ Line numbers move as versions change (as the voice-edge note says). The ve= above is the voice-edge version (sha) at the time this line was written — check the real file at that sha first.
  • Mid-call — no termination on the verdict (corrected 2026-09-16): the bridge still contains the D-1 code (classify_call with should_terminate=true → after the closing line, look up AEGIS_UUID via AMI Status → Hangup(cause=21)), but conversation_may_terminate() is always False, so it never fires on any call (br=a557a5b1d41d, src/bridge.py). should_terminate only remains in the record as material that steers the AI toward taking a message. The final design (Plan A) removes this path together with terminate_pending, the threshold check and the channel scan. The old option A (AMI DBPut → AstDB → dialplan DB()) was also rejected (ADR-0001); a DB_DELETE(aegis_verdict/…) remains in the dialplan h extension with no writer.
  • Termination safety valve (AEGIS_TERMINATE_MIN_SCORE, default 0.8): since the path above never fires, it has no effect on the phone path. The should_terminate derived by the platform verdict port is used by the browser demo and for records; it is not an instruction to cut a call.
  • Opener-floor guardrail (opener floor / L1 rules): “authority frame + just-a-confirmation”-style opener impersonation (sales / scams) keeps spam_score low in turns 1–2, before any decisive utterance arrives (0.1–0.2 measured), and gets misjudged as legitimate and passed straight through to a human. To close this hole, an opener floor of 0.6 is OR-merged (max merge) into the champion’s spam_score on its way through the verdict port.
    • Detection (backend/engine/opener_floor.py): after NFKC normalization, ① opener-template substrings (還付金 / 保険年金課 / 無料点検 / モニター価格 / 内容のご確認だけ, etc.) and ② a compound match of “authority words (市の / 行政 / 役所 / 自治体 …) ∩ delegation words (委託 / 委任 / 受託)” (independent of word order and particles; neither side fires alone).
    • The 0.6 floor is below the 0.8 safety valve, so it never terminates alone (it only lifts the call into the grey band = toward a human). should_terminate is derived from the post-floor spam_score by the existing rule (no duplicated terminate logic = the verdict port contract).
    • score_source stays rules (it is a guardrail in the L1 rules layer). Even when an LLM becomes the champion in the future, the same function OR-merges into the LLM score.
    • Evidence: adversarial A/B (56 attack turns, 48 legitimate turns) measured breakthroughs (score<0.5) 3→0 and 0 legitimate calls caught (= zero false-positive cost). Details: aegis-platform decisions/ab_replay_2026-07-06.md.

Post-call data recording (the main line for training data)

Separate from mid-call termination, there is a layer that records the full conversation and the verdict once the call ends. This is the main line that feeds the moat (data asset) and classifier improvement (eval).

  • Path: the bridge (Route C) sends to POST /api/realtime/transcript-final exactly once at call end. The payload is the full conversation (ordered caller: / ai: lines) plus the latest mid-call classify_call LLM verdict (spam_score / detected_category).
  • fail-safe: a failed send never blocks call-end processing. Even with zero transcript lines, the verdict is sent if present (so the verdict is never lost).
  • Destination: upserted into call_logs (key = call UUID). The rows accumulated here feed the human review → training → eval loop — that is the main line.
  • Independent of the in-call verdict. The point is to record every call, including the ones handed to a human and the ones where a message was taken.
  • Full-conversation verdict (llm_conversation): on this call-end path, the full conversation is run through a conversation judge (one LLM pass over the whole transcript), and that verdict becomes the truth for the record (precedence below). The aim is to override noisy per-turn verdicts (sinking in the opening turns, spiking on sensitive words) with a verdict that sees the whole conversation. Design: aegis-platform decisions/aegis_scorer_v3_design_2026-07-07.md (wiring not yet implemented — per our rule, this section settles the contract ahead of the implementation).

Usage metadata (seconds the AI actually used, and dropped-record accounting)

The post-call POST body (above) also carries usage that maps to OpenAI’s billed cost. Send seconds the AI was actively generating a response, not the whole call’s duration (rationale: cost is not incurred while the AI is idle).

FieldTypeNotes
ai_active_secondsfloat | nullSeconds the AI actually used (the response.created→response.done span, summed per the counting policy below). Not the whole call’s elapsed time. Distinct from the existing call_seconds (the usage_events table’s separate billing figure) — do not conflate them (there is a recorded incident where call_seconds was inflated to the whole-call cap: aegis-sip-bridge docs/audits/2026-08-06_implementation_quality_audit.md)
tokens_in / tokens_outint | nullToken counts aggregated from response.done’s usage object. The exact field names still need verification against the real OpenAI Realtime API (the bridge-side design doc records this as unconfirmed; whoever confirms it with a real key must append the date and repo version)
dropped_after_snapshotint | nullCount of response.done events that arrived after usage was finalized (at the top of finally, the same instant as the call-end timestamp). Late events are not waited for, but the count is always sent (never silently dropped)
occurred_atstr (ISO 8601, UTC, Z suffix)When the call ended (same instant as call_ended_at, captured at the top of finally). Example: 2026-09-10T15:04:05Z. Send it in UTC — never convert to JST before sending (rationale: keep one absolute-time representation and let the receiving side (pf) own the calendar-month rounding; converting at the sender would split the UTC⇄JST conversion rule across two places)

Aggregation unit (confirmed 2026-09-10 by Okada — this also constrains bridge): usage is aggregated per customer, per month, where month means the JST calendar month — not a UTC month boundary, and not a trailing 30-day window. Converting occurred_at (UTC) to JST and rounding to a month boundary is pf’s aggregation responsibility. If this is left ambiguous, a call that ends near JST month-end can appear to fall in the previous or next day’s month under UTC and get miscounted into the wrong month. Bridge performs no conversion or rounding — it only sends the raw UTC instant (keeping the conversion rule in exactly one place, per the no-duplication principle). The remaining four aggregation dimensions (call count / tokens in and out kept separate / JPY conversion rate / turn count) are all settings or computations that live on the pf side — bridge sends no new fields for them (tokens in/out are already sent above; turn count reuses the existing caller_turns). Caps may be expressed against any of several of these measures, and an unset cap means no limit (pf’s design).

The null contract (both repos must honor it — a promise made by only one side is not a promise):

  • null (including an omitted field) means “could not be measured.” 0 means “genuinely used none.” Do not conflate the two (e.g. a field-name mismatch must yield null, never a substituted 0).
  • The receiving side (pf) must not treat null as “used nothing ($0).” Exclude it from cost_usd calculation and cap enforcement, and do not write a 0 row into usage_events (keep it as something a human reviews). For cost caps specifically, “when unsure, let it through (fail-open)” is the one place that is wrong — leave it visibly unresolved instead.
  • When dropped_after_snapshot is 1 or more, treat that call’s usage as an incomplete tally (details follow pf’s own design).

Source: aegis-sip-bridge docs/design/DESIGN_D0_F14_USAGE_METADATA_20260910.md (PR #160 — what the bridge side captures and sends; T2’s second review round closed it out). The aegis-platform receiving side is PR #280 (to be wired up after this page is confirmed).

score_source (which brain produced the score)

Because three kinds of verdict sources can land in call_logs.spam_score, they are always distinguished by the score_source column (guards against verdict duplication).

score_sourceVerdict sourcePath
rulesRule-based scorer (keyword weights, categories)written by transcript-tick
llmThe LLM’s classify_call (gpt-realtime) = per-turn signal-so-farwritten by transcript-final
llm_conversationConversation judge over the full transcript (gpt-realtime etc.)written by transcript-final (call-end). The truth for recording & training (precedence below)
nullNo score assigned / legacy data—
  • Which score_source is the more accurate classifier (= the eval question) is undecided. Once real data accumulates, it will be decided by eval (scripts/eval_scorer.py --source rules|llm|llm_conversation) against the same gold set.
  • On a separate axis, “which score_source serves which decision (routing vs recording)” IS settled — see the precedence below (independent of classifier quality; do not conflate the two).
  • Early observations: the rule layer is weak against paraphrase and STT errors and misses a lot / the LLM layer (per-turn) catches them but costs money and latency, and is noisy — sinking in turns 1–2, spiking on sensitive words / the full-conversation verdict smooths that noise over the whole arc but is only available after the call. Because everything carries a name tag, we can measure fairly later — that much is confirmed today.

The verdict port contract (the single entry point to the classifier brain)

Every path that asks for a verdict (transcript-tick / transcript-final / the bridge’s classify / the future Rhodium webhook) calls the classifier brain through a single contract: the verdict port. Whether the champion is rules or llm, whether a shadow runs alongside, and how guardrails apply are internal concerns of the port, invisible to callers (detailed design: 00_Daimyo_Brain/decisions/aegis_L3_design_2026-07-04.md L3-D4).

Input (VerdictRequest):

FieldTypeNotes
call_refstrCall UUID (same as the call_logs upsert key)
client_idstrTenant
caller_numberstr | nullE.164. May be null on unwired paths
transcriptstrAccumulated conversation text
turn_indexint | nullTurn number for mid-call verdicts (null for post-call verdicts)
channelsip / realtime / telephony_webhookName tag of the calling path

Output (Verdict):

FieldTypeNotes
actionstrThe verdict truth. 8 values (direct_transfer / after_hours / blocked / transferred / allowed_sales / form_guided / grey_transferred plus pending / error). blocked = L1 blocklist hit (explicitly registered blocked number)
spam_scorefloat | null0.0–1.0. null on short-circuit
detected_categorystr | nullThe 8-way classification is the primary output
confidencefloat | nullOnly on L3 (llm) verdicts
score_sourcerules / llm / llm_conversation / nullWhich brain produced the score (the same name tag as the section above). llm_conversation = the post-call full-conversation verdict (precedence below)
should_terminateboolA derived value, not an independent channel. When action is a terminating one: L1 short-circuit origins (after_hours / blocked) are unconditionally true (deterministic gates are outside the safety valve), score-based ones (form_guided) are true only when spam_score ≥ AEGIS_TERMINATE_MIN_SCORE (the safety valve)
should_transfer / transfer_numberbool / str | nullDerived values from transfer-type actions

Contract invariants (fail-open baked into the contract):

  1. Never compute should_terminate outside the port (duplication is forbidden). The truth is always action; should_terminate is its projection.
  2. The port never leaks exceptions to callers. Internal errors and timeouts return a Verdict with action="error" (grey side = to a human).
  3. L1 deterministic gates (whitelist / hours / auth / blocklist) always rank above L3 (llm) verdicts.

Verdict precedence (when multiple score_sources coexist)

A single call can carry multiple score_sources (rules / llm / llm_conversation). Which one is “the truth” depends on the decision being asked (a separate axis from classifier quality — the eval question in the score_source section above).

DecisionAuthoritative score_sourceInformationPath (verdict port)
Routing decision (transfer / cut)Live signal-so-far (rules / llm, per-turn)What was available mid-callturn_index non-null (mid-call)
Ground-truth label for recording & trainingPost-call full-conversation verdict (llm_conversation)The full conversationturn_index = null (call-end, post-hoc)
  1. Routing (AMI Hangup / transfer) is decided mid-call, so in principle only signal-so-far can be used. The post-call llm_conversation never overturns routing after the fact (the call is already over).
  2. The “record truth” in call_logs and the eval training label prefer llm_conversation. The verdict that saw the whole conversation — not the noisy per-turn ones — is the truth for the moat (data asset). Live per-turn verdicts are kept, never discarded, with their score_source name tags (so everything can be measured fairly later).
  3. Why this does not violate the no-duplication rule: contract invariant 1 forbids computing should_terminate outside of action — it does not forbid multiple name-tagged score_sources coexisting on the same call (the score_source section was designed for coexistence from the start). Inside each verdict, should_terminate remains a unique derivation from action.
  4. fail-open invariant (post-call verdicts never cut live calls): llm_conversation is for recording & training only and never fires a real-channel termination (AMI Hangup). However high the post-call score, the call has already ended and is never cut retroactively. The live termination safety valve (spam_score ≥ AEGIS_TERMINATE_MIN_SCORE) continues to apply to signal-so-far only. (A future variant that runs the full-conversation judge mid-call and lets it participate in termination — semi-live — is a separate decision, gated on meeting the per-tenant false-positive bar. As of this section, llm_conversation plays no part in termination.)
  5. llm_conversation is observed, not gold (per D6): the full-conversation verdict’s output is a model prediction record (llm_conversation_observed, same rank as llm_observed) and not ground truth (gold_is_spam / gold_category). Promotion to gold always goes through a human label (keeping call_log_id for training isolation = D6). Never promote the full-conversation verdict’s predictions to gold unverified (no contaminating the eval foundation).

Source-sharing policy

Shared withScopeNotes
Chris (Namzak Labs)The entire repositoryFoundation choices such as the OS are Chris’s call
ChatVoiceconfigs and scripts onlyLimited to what installation and operations require

The phone-call hand-off contract (fixed 2026-09-16, not implemented)

The hand-off for the final design (Plan A) is fixed in one sheet: aegis-docs docs/PHONE_CALL_CONTRACT.md. The sections above describe the current code; the contract items are not implemented yet. Okada-san’s final answers (not to be re-decided):

ItemFixed value
At the cap (500 unsent messages / 30 days)AI intake stops. The fixed phone still rings (fixed phone takes priority over AI-intake stop)
After a human answersNo time limit (no deadline announcement, no automatic cut)
ClientsOne company for now (one device = one client; routing by called number is outside the contract)
RecordingsDeleted from the Pi after the NAS confirms them; untransferred ones kept within half of the SD’s free space, oldest first. Messages are never deleted
ClockThe AI intake path is managed at 300 s from ring. The 90 s re-intake is a guide, not a separate hard cap

Changelog

  • 2026-09-16: Corrected descriptions that differed from the current code (stage 1, D01). The in-call automatic hangup (AMI Hangup, D-1) never fires because conversation_may_terminate() is always False, so it is now “not performed; slated for removal”; the safety valve is noted as having no effect on the phone path. Added the “phone-call hand-off contract” section registering docs/PHONE_CALL_CONTRACT.md (D00) and the five final answers.

  • 2026-09-10: Added the “Usage metadata (seconds the AI actually used, and dropped-record accounting)” section. Registered ai_active_seconds / tokens_in / tokens_out / dropped_after_snapshot / occurred_at as fixed values, and codified the contract distinguishing null (could not be measured) from 0 (genuinely used none). occurred_at is sent as a fixed UTC ISO 8601 instant, with the calendar-month (JST) rounding left entirely to pf’s aggregation (Okada confirmed: usage is aggregated per customer per JST calendar month; converting at the sender would split the conversion rule across two places). This registration precedes the aegis-sip-bridge implementation (PR #160, T2’s second review round closed it out), per the no-reverse-order rule. Also noted this is distinct from the existing call_seconds (usage_events). The aegis-platform receiving side (PR #280) will land in a follow-up once this page is confirmed.

  • 2026-09-05: Added the Realtime default model as a fixed value (gpt-realtime-2.1). Okada’s 2026-09-02 decision had landed in bridge (src/bridge.py:166,173 / deploy/bridge.env.example:141,174, measured at br=5a8a221) but had never reached this SSoT (gpt-realtime-2.1 appeared 0 times here). Fixed both the fixed-value table and the prose, in ja and en together. ⚑ aegis-platform has not followed (the value that actually takes effect is backend/config.py:160 openai_realtime_model: str = "gpt-realtime"; backend/adapters/realtime_adapter.py:131 reads it first and falls back to _DEFAULT_OPENAI_REALTIME_MODEL = "gpt-realtime" at :38 — both are plain, measured at pf=8e3be70). Not changed in this repo (another repo’s implementation belongs in another PR).

  • 2026-07-07: Added the verdict-precedence section (v3 design, aegis-platform decisions/aegis_scorer_v3_design_2026-07-07.md). Added llm_conversation (post-call full-conversation verdict) to score_source and settled the precedence: routing is owned by live signal-so-far (rules/llm); recording/training is owned by the full-conversation verdict (llm_conversation). Clarified that the no-duplication rule forbids external computation of should_terminate, not the coexistence of score_sources. Post-call verdicts are record-only and never cut live calls (fail-open) / observed, not gold (human promotion required, D6). Also added llm_conversation to the Verdict output table’s score_source in the “verdict port contract” section (consistent with the precedence section). Wiring not yet implemented (per our rule, the SSoT is settled ahead of the implementation).

  • 2026-07-07: Added the opener-floor guardrail (opener floor / L1 rules — a 0.6 floor OR-merged into the champion score) to the verdict flow. Adversarial A/B (56 attack / 48 legitimate) measured breakthroughs 3→0 with 0 legitimate calls caught (aegis-platform decisions/ab_replay_2026-07-06.md). Below the 0.8 safety valve, so it never terminates alone; should_terminate is derived from the post-floor score by the existing rule (no duplication).

  • 2026-07-04: Added the “verdict port contract” section (an update ahead of implementation, following the approval of L3 design L3-D4). Consolidated the classifier brain’s entry point into a single contract and defined should_terminate as a derived value from action + the safety valve (independent channels forbidden). Also added blocked to action (now 8 values) — a vocabulary addition accompanying the L1 blocklist’s scorer wiring (resolving the L3-D2 drift).

  • 2026-07-04: Updated the verdict path to implementation fact (active termination via AMI Hangup (D-1); the AstDB DBPut path rejected/removed) and added the termination safety valve AEGIS_TERMINATE_MIN_SCORE=0.8. Consistent with ADR-0001 (option B adopted).

  • 2026-07-04: Added the “post-call data recording (transcript-final — the main line for training data)” section and the score_source (rules/llm distinction) section. Positioned Route C (gpt-realtime) as the primary voice route (A22 = telephony foundation).

  • 2026-06-18: First version. Recorded the day’s decisions (OS / Asterisk 22 / port 9092 / verdict path / A22 as main line & standalone frozen / source-sharing policy) as confirmed.