🎯Single Source of Truth (SSoT)
Confirmed cross-repo facts (authoritative): target OS / Asterisk version / ports / verdict path. All repos reference this page.
🕒 Last updated: 2026-07-14 (JST)
This page is the Single Source of Truth for the whole Aegis project (aegis-platform / aegis-sip-bridge / voice-edge / aegis-docs). To keep the target OS, Asterisk version, and similar facts from drifting apart across repos (cross-repo drift), every confirmed fact is consolidated here.
- This page is authoritative. Each repo’s README / docs / code comments reference this page instead of hard-coding values.
- When a value changes, update this page first, then propagate to the repos (never the other way around).
- Items marked “being verified on device” are not treated as confirmed (see the ⚠️ note below).
Confirmed facts (authoritative)
| Item | Confirmed value | Notes / background |
|---|---|---|
| Target OS | Raspberry Pi OS Bookworm Lite 64bit (aarch64) | Chris’s choice; unified on aarch64 |
| Asterisk | 22 LTS (source build) | Built from source, not the packaged version |
| Python (Pi host) | 3.11 (Bookworm default) | Uses the OS default as-is |
| Python (bridge) | 3.12 planned ⚠️ | Being verified on device (see note below) |
| Primary voice route | Route C (OpenAI Realtime GA) | The main line for AI responses. voice-edge A22 is its telephony foundation (below) |
| Realtime default model | gpt-realtime-2.1 | Decided by Okada on 2026-09-02 (bridge src/bridge.py:173 AEGIS_REALTIME_MODEL_DEFAULT). Chosen to match the default voice marin. Overridable via env AEGIS_REALTIME_MODEL (2 stages: (1) env, (2) default; there is no file stage). ⚑ We do NOT claim that plain gpt-realtime and gpt-realtime-2.1 are the same model under different names (all that was measured is that their default voices differ; plain defaults to alloy) |
| Primary telephony foundation | voice-edge A22 | The main line. See the comparison table below |
| AudioSocket port | 9092 | Asterisk ↔ bridge.py |
| Post-call data recording | POST /api/realtime/transcript-final (once, after the call ends) | Saves the full conversation + LLM verdict to call_logs. The main line for training data |
| Verdict source for spam_score | call_logs.score_source = rules / llm / llm_conversation | Multiple verdicts can land in the same score column; the column tells them apart (guards against verdict duplication) + precedence (routing = live truth / recording & training = full-conversation verdict truth; below) |
| Verdict termination (mid-call) | No automatic hangup on a conversation verdict (slated for removal; final design 2026-09-16) | The bridge’s AMI Hangup (D-1) code remains but conversation_may_terminate() is always False, so it never fires (br=a557a5b1d41d). The AstDB DBPut path was rejected (ADR-0001) |
| Termination safety valve | AEGIS_TERMINATE_MIN_SCORE=0.8 (bridge side: remnant of an unreachable path; slated for removal) | The should_terminate derived by the platform verdict port stays for recording and the browser demo; it is not used to cut phone calls |
| route-decision (pre-call) | curl (separate layer) | Route selection at call arrival. Distinct from the verdict path |
| Entry point to the classifier brain | verdict port (single contract) | Every path calls the port. should_terminate is a derived value from action + the safety valve (see “verdict port contract” below) |
Python 3.12 for the bridge has not finished on-device verification. Until it does, treat it as “3.12 targeted, being verified on device” and do not state 3.12 as settled in any repo. (Python 3.11 on the Pi host IS confirmed, since it is the Bookworm default.)
The main line (A22) and the frozen design
How A22 and Route C relate: voice-edge A22 is the telephony foundation (lines, Asterisk, AudioSocket),
while Route C (OpenAI Realtime GA, default model gpt-realtime-2.1) is the main line for AI responses. They are separate
layers: Route C runs on top of A22.
v2 standalone (Ubuntu 24.04 / Asterisk 20) is frozen. It remains as historical material, but is not used as the basis for new implementation, documentation, or estimates. Going forward, the only truth is voice-edge A22.
The verdict flow (don’t mix up the paths)
The path that returns the verdict and the pre-call route decision (route-decision) are separate layers. Do not conflate them.
- Pre-call — route-decision: decides which route handles an incoming call. Queried via curl (HTTP).
The response’s
decisionhas 4 values (block/human/ai/take_message). ★Not the same thing asaction(the verdict vocabulary) — different words, different layer. ⚑ Besides the values there are 2 other cases (not values): empty string = no decision could be made → the dialplan falls to the static fallback and logs a WARNING (voice-edge20-from-hikari.conf:161-162, ve=e94704495635) / an unknown value outside the 4 → hand to a human (same file:167-168, ve=e94704495635). ⚑ In the final design (2026-09-16), empty / invalid / 401 / timeout ring the fixed phone once instead of the static fallback (stage 1 E01, not implemented). ★Unless these 2 cases are written next to the 4 values, the site reads it as “only 4 values ever arrive”. ⚑ Line numbers move as versions change (as the voice-edge note says). Theve=above is the voice-edge version (sha) at the time this line was written — check the real file at that sha first. - Mid-call — no termination on the verdict (corrected 2026-09-16): the bridge still contains the D-1 code
(
classify_callwithshould_terminate=true→ after the closing line, look upAEGIS_UUIDvia AMIStatus→Hangup(cause=21)), butconversation_may_terminate()is always False, so it never fires on any call (br=a557a5b1d41d,src/bridge.py).should_terminateonly remains in the record as material that steers the AI toward taking a message. The final design (Plan A) removes this path together withterminate_pending, the threshold check and the channel scan. The old option A (AMIDBPut→ AstDB → dialplanDB()) was also rejected (ADR-0001); aDB_DELETE(aegis_verdict/…)remains in the dialplanhextension with no writer. - Termination safety valve (
AEGIS_TERMINATE_MIN_SCORE, default 0.8): since the path above never fires, it has no effect on the phone path. Theshould_terminatederived by the platform verdict port is used by the browser demo and for records; it is not an instruction to cut a call. - Opener-floor guardrail (opener floor / L1 rules): “authority frame + just-a-confirmation”-style
opener impersonation (sales / scams) keeps spam_score low in turns 1–2, before any decisive utterance
arrives (0.1–0.2 measured), and gets misjudged as legitimate and passed straight through to a human.
To close this hole, an opener floor of 0.6 is OR-merged (
maxmerge) into the champion’sspam_scoreon its way through the verdict port.- Detection (
backend/engine/opener_floor.py): after NFKC normalization, ① opener-template substrings (還付金/保険年金課/無料点検/モニター価格/内容のご確認だけ, etc.) and ② a compound match of “authority words (市の/行政/役所/自治体…) ∩ delegation words (委託/委任/受託)” (independent of word order and particles; neither side fires alone). - The 0.6 floor is below the 0.8 safety valve, so it never terminates alone (it only lifts the
call into the grey band = toward a human).
should_terminateis derived from the post-floorspam_scoreby the existing rule (no duplicated terminate logic = the verdict port contract). score_sourcestaysrules(it is a guardrail in the L1 rules layer). Even when an LLM becomes the champion in the future, the same function OR-merges into the LLM score.- Evidence: adversarial A/B (56 attack turns, 48 legitimate turns) measured breakthroughs (score<0.5) 3→0
and 0 legitimate calls caught (= zero false-positive cost). Details: aegis-platform
decisions/ab_replay_2026-07-06.md.
- Detection (
Post-call data recording (the main line for training data)
Separate from mid-call termination, there is a layer that records the full conversation and the verdict once the call ends. This is the main line that feeds the moat (data asset) and classifier improvement (eval).
- Path: the bridge (Route C) sends to
POST /api/realtime/transcript-finalexactly once at call end. The payload is the full conversation (orderedcaller:/ai:lines) plus the latest mid-callclassify_callLLM verdict (spam_score/detected_category). - fail-safe: a failed send never blocks call-end processing. Even with zero transcript lines, the verdict is sent if present (so the verdict is never lost).
- Destination: upserted into
call_logs(key = call UUID). The rows accumulated here feed the human review → training → eval loop — that is the main line. - Independent of the in-call verdict. The point is to record every call, including the ones handed to a human and the ones where a message was taken.
- Full-conversation verdict (
llm_conversation): on this call-end path, the full conversation is run through a conversation judge (one LLM pass over the whole transcript), and that verdict becomes the truth for the record (precedence below). The aim is to override noisy per-turn verdicts (sinking in the opening turns, spiking on sensitive words) with a verdict that sees the whole conversation. Design: aegis-platformdecisions/aegis_scorer_v3_design_2026-07-07.md(wiring not yet implemented — per our rule, this section settles the contract ahead of the implementation).
Usage metadata (seconds the AI actually used, and dropped-record accounting)
The post-call POST body (above) also carries usage that maps to OpenAI’s billed cost. Send seconds the AI was actively generating a response, not the whole call’s duration (rationale: cost is not incurred while the AI is idle).
| Field | Type | Notes |
|---|---|---|
ai_active_seconds | float | null | Seconds the AI actually used (the response.created→response.done span, summed per the counting policy below). Not the whole call’s elapsed time. Distinct from the existing call_seconds (the usage_events table’s separate billing figure) — do not conflate them (there is a recorded incident where call_seconds was inflated to the whole-call cap: aegis-sip-bridge docs/audits/2026-08-06_implementation_quality_audit.md) |
tokens_in / tokens_out | int | null | Token counts aggregated from response.done’s usage object. The exact field names still need verification against the real OpenAI Realtime API (the bridge-side design doc records this as unconfirmed; whoever confirms it with a real key must append the date and repo version) |
dropped_after_snapshot | int | null | Count of response.done events that arrived after usage was finalized (at the top of finally, the same instant as the call-end timestamp). Late events are not waited for, but the count is always sent (never silently dropped) |
occurred_at | str (ISO 8601, UTC, Z suffix) | When the call ended (same instant as call_ended_at, captured at the top of finally). Example: 2026-09-10T15:04:05Z. Send it in UTC — never convert to JST before sending (rationale: keep one absolute-time representation and let the receiving side (pf) own the calendar-month rounding; converting at the sender would split the UTC⇄JST conversion rule across two places) |
Aggregation unit (confirmed 2026-09-10 by Okada — this also constrains bridge): usage is aggregated
per customer, per month, where month means the JST calendar month — not a UTC month boundary,
and not a trailing 30-day window. Converting occurred_at (UTC) to JST and rounding to a month boundary
is pf’s aggregation responsibility. If this is left ambiguous, a call that ends near JST month-end can
appear to fall in the previous or next day’s month under UTC and get miscounted into the wrong month.
Bridge performs no conversion or rounding — it only sends the raw UTC instant (keeping the conversion
rule in exactly one place, per the no-duplication principle). The remaining four aggregation
dimensions (call count / tokens in and out kept separate / JPY conversion rate / turn count) are all
settings or computations that live on the pf side — bridge sends no new fields for them (tokens
in/out are already sent above; turn count reuses the existing caller_turns). Caps may be expressed
against any of several of these measures, and an unset cap means no limit (pf’s design).
The null contract (both repos must honor it — a promise made by only one side is not a promise):
null(including an omitted field) means “could not be measured.”0means “genuinely used none.” Do not conflate the two (e.g. a field-name mismatch must yieldnull, never a substituted0).- The receiving side (
pf) must not treatnullas “used nothing ($0).” Exclude it from cost_usd calculation and cap enforcement, and do not write a0row intousage_events(keep it as something a human reviews). For cost caps specifically, “when unsure, let it through (fail-open)” is the one place that is wrong — leave it visibly unresolved instead. - When
dropped_after_snapshotis 1 or more, treat that call’s usage as an incomplete tally (details followpf’s own design).
Source: aegis-sip-bridge docs/design/DESIGN_D0_F14_USAGE_METADATA_20260910.md (PR #160 —
what the bridge side captures and sends; T2’s second review round closed it out). The
aegis-platform receiving side is PR #280 (to be wired up after this page is confirmed).
score_source (which brain produced the score)
Because three kinds of verdict sources can land in call_logs.spam_score, they are always
distinguished by the score_source column (guards against verdict duplication).
score_source | Verdict source | Path |
|---|---|---|
rules | Rule-based scorer (keyword weights, categories) | written by transcript-tick |
llm | The LLM’s classify_call (gpt-realtime) = per-turn signal-so-far | written by transcript-final |
llm_conversation | Conversation judge over the full transcript (gpt-realtime etc.) | written by transcript-final (call-end). The truth for recording & training (precedence below) |
null | No score assigned / legacy data | — |
- Which score_source is the more accurate classifier (= the eval question) is undecided. Once real
data accumulates, it will be decided by eval (
scripts/eval_scorer.py --source rules|llm|llm_conversation) against the same gold set. - On a separate axis, “which score_source serves which decision (routing vs recording)” IS settled — see the precedence below (independent of classifier quality; do not conflate the two).
- Early observations: the rule layer is weak against paraphrase and STT errors and misses a lot / the LLM layer (per-turn) catches them but costs money and latency, and is noisy — sinking in turns 1–2, spiking on sensitive words / the full-conversation verdict smooths that noise over the whole arc but is only available after the call. Because everything carries a name tag, we can measure fairly later — that much is confirmed today.
The verdict port contract (the single entry point to the classifier brain)
Every path that asks for a verdict (transcript-tick / transcript-final / the bridge’s classify / the
future Rhodium webhook) calls the classifier brain through a single contract: the verdict port.
Whether the champion is rules or llm, whether a shadow runs alongside, and how guardrails apply are
internal concerns of the port, invisible to callers (detailed design:
00_Daimyo_Brain/decisions/aegis_L3_design_2026-07-04.md L3-D4).
Input (VerdictRequest):
| Field | Type | Notes |
|---|---|---|
call_ref | str | Call UUID (same as the call_logs upsert key) |
client_id | str | Tenant |
caller_number | str | null | E.164. May be null on unwired paths |
transcript | str | Accumulated conversation text |
turn_index | int | null | Turn number for mid-call verdicts (null for post-call verdicts) |
channel | sip / realtime / telephony_webhook | Name tag of the calling path |
Output (Verdict):
| Field | Type | Notes |
|---|---|---|
action | str | The verdict truth. 8 values (direct_transfer / after_hours / blocked / transferred / allowed_sales / form_guided / grey_transferred plus pending / error). blocked = L1 blocklist hit (explicitly registered blocked number) |
spam_score | float | null | 0.0–1.0. null on short-circuit |
detected_category | str | null | The 8-way classification is the primary output |
confidence | float | null | Only on L3 (llm) verdicts |
score_source | rules / llm / llm_conversation / null | Which brain produced the score (the same name tag as the section above). llm_conversation = the post-call full-conversation verdict (precedence below) |
should_terminate | bool | A derived value, not an independent channel. When action is a terminating one: L1 short-circuit origins (after_hours / blocked) are unconditionally true (deterministic gates are outside the safety valve), score-based ones (form_guided) are true only when spam_score ≥ AEGIS_TERMINATE_MIN_SCORE (the safety valve) |
should_transfer / transfer_number | bool / str | null | Derived values from transfer-type actions |
Contract invariants (fail-open baked into the contract):
- Never compute
should_terminateoutside the port (duplication is forbidden). The truth is alwaysaction;should_terminateis its projection. - The port never leaks exceptions to callers. Internal errors and timeouts return a Verdict with
action="error"(grey side = to a human). - L1 deterministic gates (whitelist / hours / auth / blocklist) always rank above L3 (llm) verdicts.
Verdict precedence (when multiple score_sources coexist)
A single call can carry multiple score_sources (rules / llm / llm_conversation).
Which one is “the truth” depends on the decision being asked (a separate axis from classifier
quality — the eval question in the score_source section above).
| Decision | Authoritative score_source | Information | Path (verdict port) |
|---|---|---|---|
| Routing decision (transfer / cut) | Live signal-so-far (rules / llm, per-turn) | What was available mid-call | turn_index non-null (mid-call) |
| Ground-truth label for recording & training | Post-call full-conversation verdict (llm_conversation) | The full conversation | turn_index = null (call-end, post-hoc) |
- Routing (AMI Hangup / transfer) is decided mid-call, so in principle only signal-so-far can be used.
The post-call
llm_conversationnever overturns routing after the fact (the call is already over). - The “record truth” in
call_logsand the eval training label preferllm_conversation. The verdict that saw the whole conversation — not the noisy per-turn ones — is the truth for the moat (data asset). Live per-turn verdicts are kept, never discarded, with theirscore_sourcename tags (so everything can be measured fairly later). - Why this does not violate the no-duplication rule: contract invariant 1 forbids computing
should_terminateoutside ofaction— it does not forbid multiple name-taggedscore_sources coexisting on the same call (thescore_sourcesection was designed for coexistence from the start). Inside each verdict,should_terminateremains a unique derivation fromaction. - fail-open invariant (post-call verdicts never cut live calls):
llm_conversationis for recording & training only and never fires a real-channel termination (AMI Hangup). However high the post-call score, the call has already ended and is never cut retroactively. The live termination safety valve (spam_score ≥ AEGIS_TERMINATE_MIN_SCORE) continues to apply to signal-so-far only. (A future variant that runs the full-conversation judge mid-call and lets it participate in termination — semi-live — is a separate decision, gated on meeting the per-tenant false-positive bar. As of this section,llm_conversationplays no part in termination.) llm_conversationis observed, not gold (per D6): the full-conversation verdict’s output is a model prediction record (llm_conversation_observed, same rank asllm_observed) and not ground truth (gold_is_spam/gold_category). Promotion to gold always goes through a human label (keepingcall_log_idfor training isolation = D6). Never promote the full-conversation verdict’s predictions to gold unverified (no contaminating the eval foundation).
Source-sharing policy
| Shared with | Scope | Notes |
|---|---|---|
| Chris (Namzak Labs) | The entire repository | Foundation choices such as the OS are Chris’s call |
| ChatVoice | configs and scripts only | Limited to what installation and operations require |
The phone-call hand-off contract (fixed 2026-09-16, not implemented)
The hand-off for the final design (Plan A) is fixed in one sheet: aegis-docs docs/PHONE_CALL_CONTRACT.md. The sections above describe the
current code; the contract items are not implemented yet. Okada-san’s final answers (not to be re-decided):
| Item | Fixed value |
|---|---|
| At the cap (500 unsent messages / 30 days) | AI intake stops. The fixed phone still rings (fixed phone takes priority over AI-intake stop) |
| After a human answers | No time limit (no deadline announcement, no automatic cut) |
| Clients | One company for now (one device = one client; routing by called number is outside the contract) |
| Recordings | Deleted from the Pi after the NAS confirms them; untransferred ones kept within half of the SD’s free space, oldest first. Messages are never deleted |
| Clock | The AI intake path is managed at 300 s from ring. The 90 s re-intake is a guide, not a separate hard cap |
Changelog
-
2026-09-16: Corrected descriptions that differed from the current code (stage 1, D01). The in-call automatic hangup (AMI Hangup, D-1) never fires because
conversation_may_terminate()is always False, so it is now “not performed; slated for removal”; the safety valve is noted as having no effect on the phone path. Added the “phone-call hand-off contract” section registeringdocs/PHONE_CALL_CONTRACT.md(D00) and the five final answers. -
2026-09-10: Added the “Usage metadata (seconds the AI actually used, and dropped-record accounting)” section. Registered
ai_active_seconds/tokens_in/tokens_out/dropped_after_snapshot/occurred_atas fixed values, and codified the contract distinguishingnull(could not be measured) from0(genuinely used none).occurred_atis sent as a fixed UTC ISO 8601 instant, with the calendar-month (JST) rounding left entirely topf’s aggregation (Okada confirmed: usage is aggregated per customer per JST calendar month; converting at the sender would split the conversion rule across two places). This registration precedes the aegis-sip-bridge implementation (PR #160, T2’s second review round closed it out), per the no-reverse-order rule. Also noted this is distinct from the existingcall_seconds(usage_events). The aegis-platform receiving side (PR #280) will land in a follow-up once this page is confirmed. -
2026-09-05: Added the Realtime default model as a fixed value (
gpt-realtime-2.1). Okada’s 2026-09-02 decision had landed in bridge (src/bridge.py:166,173/deploy/bridge.env.example:141,174, measured at br=5a8a221) but had never reached this SSoT (gpt-realtime-2.1appeared 0 times here). Fixed both the fixed-value table and the prose, in ja and en together. ⚑ aegis-platform has not followed (the value that actually takes effect isbackend/config.py:160openai_realtime_model: str = "gpt-realtime";backend/adapters/realtime_adapter.py:131reads it first and falls back to_DEFAULT_OPENAI_REALTIME_MODEL = "gpt-realtime"at:38— both are plain, measured at pf=8e3be70). Not changed in this repo (another repo’s implementation belongs in another PR). -
2026-07-07: Added the verdict-precedence section (v3 design, aegis-platform
decisions/aegis_scorer_v3_design_2026-07-07.md). Addedllm_conversation(post-call full-conversation verdict) toscore_sourceand settled the precedence: routing is owned by live signal-so-far (rules/llm); recording/training is owned by the full-conversation verdict (llm_conversation). Clarified that the no-duplication rule forbids external computation ofshould_terminate, not the coexistence of score_sources. Post-call verdicts are record-only and never cut live calls (fail-open) / observed, not gold (human promotion required, D6). Also addedllm_conversationto the Verdict output table’sscore_sourcein the “verdict port contract” section (consistent with the precedence section). Wiring not yet implemented (per our rule, the SSoT is settled ahead of the implementation). -
2026-07-07: Added the opener-floor guardrail (opener floor / L1 rules — a 0.6 floor OR-merged into the champion score) to the verdict flow. Adversarial A/B (56 attack / 48 legitimate) measured breakthroughs 3→0 with 0 legitimate calls caught (aegis-platform
decisions/ab_replay_2026-07-06.md). Below the 0.8 safety valve, so it never terminates alone;should_terminateis derived from the post-floor score by the existing rule (no duplication). -
2026-07-04: Added the “verdict port contract” section (an update ahead of implementation, following the approval of L3 design L3-D4). Consolidated the classifier brain’s entry point into a single contract and defined
should_terminateas a derived value fromaction+ the safety valve (independent channels forbidden). Also addedblockedtoaction(now 8 values) — a vocabulary addition accompanying the L1 blocklist’s scorer wiring (resolving the L3-D2 drift). -
2026-07-04: Updated the verdict path to implementation fact (active termination via AMI Hangup (D-1); the AstDB DBPut path rejected/removed) and added the termination safety valve
AEGIS_TERMINATE_MIN_SCORE=0.8. Consistent with ADR-0001 (option B adopted). -
2026-07-04: Added the “post-call data recording (transcript-final — the main line for training data)” section and the
score_source(rules/llm distinction) section. Positioned Route C (gpt-realtime) as the primary voice route (A22 = telephony foundation). -
2026-06-18: First version. Recorded the day’s decisions (OS / Asterisk 22 / port 9092 / verdict path / A22 as main line & standalone frozen / source-sharing policy) as confirmed.