Skip to content

Latest commit

 

History

History
651 lines (519 loc) · 37.8 KB

File metadata and controls

651 lines (519 loc) · 37.8 KB
title Locus Signal — High Level Design
version 2.5
status implemented
owner Founder's Office assignment
tags
architecture
hld
locus
weekly-brief
updated 2026-08-22

Locus Signal — High Level Design

[!abstract] The job Every week, tell the CRO what happened to Locus customers that changes whether they are about to churn or about to expand — split into risk and opportunity, readable in four minutes, forwardable one item at a time.

This is a precision product, not a coverage product. The value is in what it throws away. Four items out of four hundred is the specification; a system that surfaces forty has failed even if all forty are true.


1. What changed from v1

The v1 diagram had the right spine and several good instincts. This revision keeps them and fixes four structural problems.

[!success] Kept from v1 — these were right

  • Accounts as a data record at the top of the flow, not a list buried in the engine. This is what makes the account list editable.
  • A Status field rather than deletion. Churned customers stay as history.
  • Scheduling drawn as a layer above the engine, with Pause / Resume named explicitly.
  • A cheap deterministic filter before the expensive AI stage. Correct and load-bearing.
  • Event clustering placed before AI analysis. The strongest single call on the page — one event gets judged once instead of nine near-identical articles being judged nine times.
  • Multiple independent sources rather than one vendor.
  • Prioritisation as a stage distinct from classification.
  • Field choices: Aliases, Monitoring Keywords, Locus Use Case, Account Owner, Tier. Better than what is currently built.

[!danger] Structural problems in v1 1. The three certainties had no home. Account list, recipients and cadence were drawn as annotations on data stores. Nothing showed who edits them or how. The assignment's hardest requirement — "none of that should require you, or anybody who can code" — was invisible in the design. → Now a first-class Control Plane.

2. Cadence lived inside the cron job. If the cron expression encodes the schedule, moving Monday to Thursday is a code change and a redeploy. → The schedule now lives in the database; the cron is a dumb hourly heartbeat that asks "is anything due?".

3. The engine was one long synchronous job. The margin note — "usually going to take the time depending on the number of accounts" — is the sharpest observation on the page and it identifies a fatal flaw: on serverless it gets killed mid-run; on a long-lived host a restart loses the week. → The engine is now a resumable state machine with checkpoints.

4. Feedback was a terminal box with no outbound arrows. Drawn as the last thing the pipeline does rather than a system that observes the pipeline. → Now an Observability Plane that reads from every stage and closes the loop back to configuration.

Plus two pipeline-ordering corrections and one source removed — see [[#5. The pipeline]] and [[#12. Decision records|ADR-006]].


2. Design principles

# Principle Consequence
P1 Precision over recall Willing to ship two items, or zero. An empty section beats a manufactured one.
P2 Config is data, never code Every operational lever lives in Postgres and is edited in a UI. No redeploy to change behaviour.
P3 Deterministic where possible, model where necessary The model makes judgement calls. Code makes every rule that can be written down — cheaper, inspectable, fails predictably.
P4 Every rejection is recorded "Why didn't I hear about X?" must be answerable. The filter is auditable, not trusted.
P5 Assume nobody is watching Silent failure is the primary risk of a weekly job. Absence of a brief must be distinguishable from a quiet week.
P6 The human keeps the judgement The system triggers an intervention; it never performs one.

3. System context

flowchart TB
    subgraph People["People"]
        CRO["CRO and other recipients"]
        OPS["Brief owner - commercial ops, non-technical"]
    end

    subgraph Control["Control Plane - auth-gated web app"]
        UIACC["Accounts - add, edit, mute, aliases, tier, context"]
        UIREC["Recipients - add, pause, remove"]
        UISCH["Schedule - cadence, day, hour, timezone, pause, thresholds"]
        UIRUN["Run now - preview or send"]
        UIHEALTH["Is it working - funnel, drop reasons, blind spots, cost"]
    end

    subgraph Store["Durable state - Postgres"]
        DB[("Postgres - accounts, recipients, settings, runs, candidates, briefs, brief_items, feedback, audit_log")]
    end

    subgraph Sched["Scheduler"]
        TICK["Hourly cron tick - knows nothing but 'wake up'"]
        DUE{"isDue - reads the schedule from the database"}
    end

    subgraph Engine["Run engine - resumable state machine"]
        RUN["startRun - stepRun - finalizeRun - sendBrief"]
    end

    subgraph Ext["External services"]
        SRC["News sources - paid chain NewsAPI.ai then NewsAPI.org, baseline Google News RSS plus GDELT"]
        LLM["Model providers - OpenRouter then Azure OpenAI, 5-key pool each"]
        MAIL["Resend - per-recipient send"]
    end

    OPS --> UIACC
    OPS --> UIREC
    OPS --> UISCH
    OPS --> UIRUN
    OPS --> UIHEALTH

    UIACC --> DB
    UIREC --> DB
    UISCH --> DB
    UIRUN --> RUN
    DB --> UIHEALTH

    TICK --> DUE
    DUE -->|"due now and not already sent"| RUN
    DUE -.->|"not due - no-op, 167 ticks a week"| TICK

    DB <--> RUN
    RUN <--> SRC
    RUN <--> LLM
    RUN --> MAIL
    MAIL --> CRO
    CRO -->|"one-click verdict"| DB
    CRO -->|"forwards one item to the account owner"| OPS
    DB -.->|"blind spots and quality signals"| OPS
Loading

[!note] Why four planes and not one pipeline The v1 diagram was a single flow because that is how the data moves. But the assignment is graded on the parts that are not the data flow: who can change the account list, what happens when it breaks, and how you prove it works. Those need to be visible as subsystems, not implied.


4. The four planes

4.1 Control Plane — the part the assignment actually grades

Everything a commercial user must be able to change, without a terminal.

Surface Handles Status
Accounts Add / edit / mute / remove. Aliases, tier, segment, region, customer-vs-prospect, free-text context ✅ Built — src/app/accounts/page.tsx
Test aliases Shows what news the current aliases find right now, pre-AI. Distinguishes a quiet account from a misconfigured one ✅ Built — src/app/accounts/TestAliases.tsx
Recipients Add / pause / remove. Pause ≠ delete — history is preserved ✅ Built — src/app/recipients/page.tsx
Schedule Cadence, day, hour, timezone, pause-until, and selectivity (max items, min severity, lookback) ✅ Built — src/app/schedule/page.tsx
Run now Preview or send on demand, with live progress ✅ Built — src/app/RunButton.tsx
Access control Password gate; cron endpoint guarded by shared secret ⚠️ Partial — fails open when unset (dev convenience)
Account owner + renewal date + ARR Owner name/email, renewal date and ARR per account. Renewal proximity feeds the ranking score ✅ Built — actions.ts, accounts/page.tsx
Monitoring keywords Per-account watch terms. Become their own search query and raise entity-resolution confidence ✅ Built — sources.ts
Exclusion terms Per-account veto. How a commercial user fixes "wrong Landmark" without a developer ✅ Built — resolveEntity in sources.ts
Markets / countries, Locus use case Recorded on the account and fed to the triage prompt ✅ Built
Alert resolution Acknowledge an open alert from the health page ✅ Built — resolveAlert

[!tip] The design move that matters The schedule is a row in the database, not a cron expression. A fixed hourly cron wakes the app; the app reads cadence, send_day, send_hour, timezone and paused_until and decides. That single indirection is what turns "change the send day" from a deploy into a dropdown.

4.2 Ingestion & Analysis Plane

The engine. Covered in [[#5. The pipeline]].

4.3 Delivery Plane

Rendering and delivery are separate concerns — v1 merged them into one box. Delivery fails independently, and it fails per recipient.

Concern Status
Brief rendered as a durable web object, archived at a permalink ✅ Built — src/lib/email.ts → renderBriefHtml
Per-recipient email so feedback links are attributable to a person ✅ Built — loop in sendBrief
Plain-text alternative ✅ Built — renderBriefText
Per-recipient delivery state and retry ✅ Built — deliveries table; "Retry the ones that failed" on the brief page

4.4 Observability Plane

Covered in [[#9. Observability]].


5. The pipeline

flowchart TB
    A["Account dequeued from runs.pending_ids - batch of 4"]:::code
    B["Collect - paid chain first success wins, Google News per locale always, GDELT only when thin"]:::code
    C["Normalise and exact dedupe - canonical URL, strips hash, utm, oc, hl, gl, ceid, ref"]:::code
    D["Cheap pre-filter - junk domains, outside lookback window"]:::code
    E["Entity resolution - alias must appear in title or snippet"]:::code
    F["Event clustering - Jaccard 0.6 on title tokens, within one account"]:::code
    G["Cap at 25 newest articles for this account"]:::code
    H{{"MODEL CALL 1 - triageAccount, cheap model, one call per account"}}:::ai
    I["Drop ladder - name collision, not revenue-relevant, neutral, below severity floor, confidence under 0.5"]:::code
    J[("candidates - every judged article persisted, kept or dropped with a reason")]:::store
    K["Editor pool - dropped_reason is null, order by severity then confidence, limit 60"]:::code
    L{{"MODEL CALL 2 - selectItems, strong model, once per run, may legitimately return zero"}}:::ai
    M["Assemble - match back to candidate by source_url, write briefs and brief_items"]:::code
    N["Render - risk and opportunity halves, one forwardable card each"]:::code
    O["Deliver - one email per active recipient, signed feedback links"]:::code

    A --> B --> C --> D --> E --> F --> G --> H --> I --> J --> K --> L --> M --> N --> O

    classDef code fill:#e8eef7,stroke:#1f3a5f,color:#10233d
    classDef ai fill:#f3e8f7,stroke:#6b2d7a,color:#331239
    classDef store fill:#eef7ee,stroke:#1e7a4f,color:#0f3a26
Loading

Legend — rectangle = deterministic code · hexagon = model call · cylinder = persisted state

5.1 Ordering corrections against v1

[!warning] Correction 1 — split deduplication, move half of it upstream v1 ran ENTITY RESOLUTION → DEDUPLICATION. That pays resolution cost on nine copies of the same wire story and then throws eight away. Exact/canonical dedupe is the cheapest stage in the pipeline and belongs at the front. Near-duplicate grouping is not a separate stage at all — it is event clustering.

[!warning] Correction 2 — collapse RELEVANCE + AI ANALYSIS and RISK / OPPORTUNITY v1 drew these as sequential stages. Relevance, direction, severity and why-it-matters are one judgement about one event. Two passes either duplicate the model call for no gain, or produce two verdicts with nothing to arbitrate between them. They are now a single structured triage call. (This is already how the implementation works — the correction is to the diagram.)

[!warning] Correction 3 — prioritisation cannot be one global step This is the most consequential flaw in v1. A flat top-4 over all candidates lets one noisy account take all four slots — which is a news digest about one company, precisely the failure the CRO named. It must be two stages: reduce each account to its strongest signal, then rank across accounts. Status: ✅ built. shortlist() reduces each account to its strongest signal, scores across accounts by tier weight, renewal proximity and escalation, and hands the editor a pool of ~4x the target. One-item-per-account is then re-enforced in code after the editor returns, so it is a guarantee rather than a request.

5.2 Stage reference

Stage Kind Implementation Status
Collect code sources.ts — premiumNews chain (retry + circuit breaker) over newsApiAi then newsApiOrg, plus googleNews per locale and gdelt as gap-filler ✅
Canonical dedupe code canonicalUrl ✅
Pre-filter code junk domains, lookback cutoff ✅
Entity resolution code resolveEntity — scored: alias phrase 0.6, partial 0.3, domain corroboration +0.25, keyword +0.2, region +0.1; operator veto terms; floor 0.4 ✅
Event clustering code similar(), Jaccard ≥ 0.6 ⚠️ near-dup only; true multi-article event clustering not built
Triage AI triageAccount via OpenRouter, OPENROUTER_TRIAGE_MODEL (default google/gemini-2.5-flash) ✅
Drop ladder code severity floor, confidence floor ✅
Per-account winner code shortlist() — one item per account guaranteed; a second only at severity 5 with a different signal type ✅
Novelty check code sent_signals over a 60-day window, by URL key and per-account title key ✅
Cross-account rank code tier weight x severity + confidence + renewal proximity + escalation ✅
Editorial cut AI selectItems via OpenRouter, OPENROUTER_EDITOR_MODEL (default anthropic/claude-sonnet-5), URL-validated against the candidate set ✅
Render + deliver code email.ts ✅

6. Run lifecycle — durability

This is the answer to the v1 margin note about runtime scaling with account count.

stateDiagram-v2
    state "stage=collect" as collect
    state "stage=select" as select
    state "stage=send" as send
    state "stage=done status=ok" as done_ok
    state "stage=done status=failed" as done_failed

    [*] --> collect : startRun writes pending_ids, period_key, should_send
    collect --> collect : stepRun takes a batch of 4, writes candidates, trims pending_ids, commits
    collect --> collect : time budget hit - handler self-re-invokes with the run id
    collect --> select : pending_ids empty
    select --> send : finalizeRun writes the brief - idempotent on re-entry
    send --> done_ok : sendBrief succeeds, last_sent_week set
    collect --> done_failed : failRun
    select --> done_failed : failRun
    send --> done_failed : failRun
    done_ok --> [*]
    done_failed --> [*]
Loading

[!info] Why chunked ~35 accounts means ~35 network fan-outs plus ~35 model calls. That does not fit in one serverless invocation. State lives in the runs row (stage, pending_ids), so any step can be retried or resumed after a crash without losing completed work.

[!success] Recovery sweepStaleRuns() runs at the start of every fresh tick and fails any run whose last_progress_at is over 30 minutes old, raising a stale_run alert. Keying on progress rather than run age matters: a 200-account run spread across many continuations is legitimately long, but it always makes progress. A continuation tick deliberately skips the sweep so it cannot kill the run it is continuing.


7. Scheduling model

flowchart TB
    T["Hourly cron tick"] --> AUTH{"Authorised - bearer or query secret"}
    AUTH -->|"no"| STOP1["401 - no run"]
    AUTH -->|"yes"| S["Load schedule from settings, compute wall clock in the configured timezone"]

    S --> P{"paused_until set and still in the past window"}
    P -->|"yes"| STOP2["Skip - paused, UI shows the resume date"]
    P -->|"no"| HR{"Current hour matches send_hour"}

    HR -->|"no"| STOP3["Skip - not the send hour"]
    HR -->|"yes"| DAY{"Cadence daily, or weekday matches send_day"}

    DAY -->|"no"| STOP4["Skip - not the send day"]
    DAY -->|"yes"| PAR{"Biweekly and ISO week number is odd"}

    PAR -->|"yes"| STOP5["Skip - off week"]
    PAR -->|"no"| SENT{"last_sent_week equals this period key"}

    SENT -->|"yes"| STOP6["Skip - already sent, the at-most-once guard"]
    SENT -->|"no"| GO["Run - collect, select, send, then write last_sent_week"]
Loading

[!success] Catch-up window The hour check is "at or after send_hour, within 6 hours", clamped so it never crosses local midnight. One weekly attempt becomes several, which also absorbs the DST spring-forward skip. last_sent_week keeps repeated attempts idempotent, so extra attempts cannot produce a second email. missedPeriod() raises a critical alert if the window closes with nothing sent.


8. Data model

erDiagram
    ACCOUNTS ||--o{ CANDIDATES : "judged for"
    ACCOUNTS ||--o{ BRIEF_ITEMS : "subject of"
    RUNS ||--o{ CANDIDATES : "produced"
    RUNS ||--o| BRIEFS : "produced"
    BRIEFS ||--o{ BRIEF_ITEMS : "contains"
    BRIEF_ITEMS ||--o{ FEEDBACK : "rated by"

    ACCOUNTS {
        int id PK
        text name UK
        text_array aliases
        text domain
        text region
        text segment
        text tier
        text kind
        text status
        text context
    }
    RUNS {
        int id PK
        text stage
        text status
        int_array pending_ids
        text period_key
        int articles_found
        int candidates
        numeric cost_usd
        jsonb log
    }
    CANDIDATES {
        int id PK
        int run_id FK
        text title
        text url
        text category
        int severity
        text dropped_reason
        boolean selected
    }
    BRIEFS {
        int id PK
        date week_of
        text headline
        text status
        int sent_to
    }
    BRIEF_ITEMS {
        int id PK
        int brief_id FK
        text category
        int severity
        text headline
        text why_it_matters
        text suggested_action
        text source_url
    }
    FEEDBACK {
        int id PK
        int brief_item_id FK
        text verdict
        text reader_email
    }
Loading

[!note] The candidates table is the product It stores every article the model judged, kept or dropped, with the reason. That single table is what makes the filter auditable rather than a black box, and it is what answers "why didn't I hear about X?".

Account fields still to add (all from v1, all good): markets_countries, locus_use_case, monitoring_keywords, account_owner, plus renewal_date and arr — the two things a CRO actually prioritises on.


9. Observability

flowchart TB
    ITEM["brief_items row shipped"] --> MAIL["Per-recipient email with HMAC-signed verdict links"]
    MAIL --> READ["Reader clicks one of three verdicts"]
    READ --> V1["useful - right item"]
    READ --> V2["not_useful - precision problem"]
    READ --> V3["already_knew - latency problem"]

    V1 --> FB[("feedback")]
    V2 --> FB
    V3 --> FB

    RUNS[("runs and candidates - funnel, drop reasons, cost")] --> HEALTH
    BR[("briefs - delivery status")] --> HEALTH
    FB --> HEALTH["Is it working surface"]

    HEALTH --> Q1["Precision - share rated useful"]
    HEALTH --> Q2["Latency - share rated already knew"]
    HEALTH --> Q3["Blind spots - active accounts never surfaced"]

    Q1 --> HUMAN{"Human tuning decision"}
    Q2 --> HUMAN
    Q3 --> ALIAS["Test aliases - what would have matched"]
    ALIAS --> HUMAN

    HUMAN -->|"adjust min_severity, max_items, lookback"| SET[("settings")]
    HUMAN -->|"fix aliases, sharpen context, retier"| ACC[("accounts")]
    SET --> NEXT["Next scheduled run"]
    ACC --> NEXT
    NEXT --> ITEM
Loading

Three questions that are usually collapsed into one:

Question Measure
Does it run? Run success rate, last run, cost per brief
Does it land? Briefs emailed, send failures, per-recipient activity
Is the judgement good? Share rated useful; already_knew tracked separately from not_useful — the first means too slow, the second means wrong. Different fixes.

Plus two things easy to miss: blind spots (active accounts that have never produced a signal — usually wrong aliases, not a quiet customer) and the per-run audit trail.

[!warning] The loop is deliberately open No rating is read back into any prompt, threshold or ranking automatically. The return path is a person making a change in the UI. Auto-tuning a precision filter on a handful of weekly clicks would overfit to noise within a month.

[!success] Push alerting lib/alerts.ts raises deduplicated alerts for run failure, abandoned runs, missed periods, partial sends, analysis failures, hallucinated source URLs and zero-item briefs, then emails them to ALERT_EMAIL (falling back to the first active recipients). Every tick flushes pending alerts, so 167 of the 168 weekly ticks exist purely to notice breakage. Open alerts also surface at the top of /health with a Resolve action.


10. Failure modes

Identified by an SRE review of the implementation, then fixed. Severity is impact on an unattended six-month run.

# Severity Failure Status
D1 🔴 Critical No alerting — /health was pull-only, so nobody would learn the brief had stopped ✅ lib/alerts.ts, deduplicated + emailed, flushed every tick
D2 🔴 Critical Duplicate email — any replay of the resume URL re-sent the whole brief ✅ deliveries gates per recipient; uniq_scheduled_run_per_period blocks a second scheduled sending run per period (manual runs deliberately exempt)
D3 🔴 Critical Model parse failure looked like a quiet week — parsed_output ?? [] turned an API failure into an empty success ✅ AnalysisError thrown on null output or refusal; recorded against the account and alerted
D4 🔴 Critical Partial send never retried — failed recipients were recorded then dropped ✅ Per-recipient state + "Retry the ones that failed"
D5 🔴 Critical Missed send had no catch-up — exactly one attempt per week ✅ 6-hour catch-up window; missedPeriod() alerts if it closes empty
D6 🟠 Major Run never marked failed on the first invocation — failRun(resume,…) where resume was 0 ✅ runId hoisted out of the try; sweepStaleRuns() catches orphans via last_progress_at
D7 🟠 Major Editor could invent a source_url ✅ Every returned item validated against the candidate set; unmatched items dropped and alerted
D8 🟠 Major Candidate pool truncated to 60 ignoring tier ✅ Replaced by shortlist(): tier weight × severity + confidence + renewal proximity + escalation
D9 🟠 Major No cross-run novelty — last week's story could return ✅ sent_signals, 60-day window, URL key + per-account title key; recorded at send, not at build, so previews do not poison it
D10 🟠 Major Source 429 was silent — an account contributed nothing and looked quiet ✅ source_health per run, surfaced on /health
D11 🟠 Major Region captured but unused — non-English news invisible ✅ REGION_LOCALES fans the broad query across locales (es-419/MX, ar/EG, id/ID, en-AE, en-SG, en-IN) alongside en-US
D12 🟡 Minor DST spring-forward skipped the send hour ✅ Absorbed by the D5 catch-up window
D13 🟡 Minor paused_until evaluated in server time ✅ Compared as a zoned calendar date
D14 🔴 Critical send_hour = 0 never fired — hour12:false renders midnight as "24", so 24 !== 0. Reachable from the UI ✅ % 24 in zonedParts

All fourteen are covered by npm test (14 assertions over the scheduling, catch-up, pause and missed-period logic).

[!example] Why D14 mattered more than it looked It is the exact shape of failure this assignment tests for. A user picks midnight from a dropdown; the brief silently never sends; no error, no failed run, nothing to notice. Found by an SRE review pass, then confirmed empirically before being believed.

Still open, deliberately

Item Why it is acceptable
Event clustering is near-duplicate only Jaccard over titles collapses wire copies of one story. True multi-article event clustering (same event, different framing, different headline) is not built. The triage pass absorbs most of the cost, and severity ranking usually surfaces the strongest telling.
Auth fails open when ADMIN_PASSWORD is unset Deliberate so a fresh clone runs with zero configuration. Documented in the README as a deploy requirement.
Unbounded runs.log JSONB Grows per batch. Needs a retention policy well before it matters at this account count.
Biweekly parity uses ISO week number Can skip an extra week at a year boundary. Cosmetic at fortnightly cadence.

11. Scale

Numbers from constants in the code — batch size 4, TIME_BUDGET_MS 220s, maxDuration 300s, 25-article cap, LIMIT 60.

Accounts Model cost / run Wall clock Verdict
35 (today) ~$0.12/account ≈ $4 ~9 batches, minutes Comfortable
200 ~$24 ~50 batches of self-re-invocation Needs a real queue; the continuation chain becomes the weak link
1000 ~$120 — Needs worker infrastructure, per-source quota management, tiered scan frequency

12. Decision records

[!question]- ADR-001 — Schedule in the database, not the cron expression Alternatives: cron expression in vercel.json; a scheduling vendor. Chosen: fixed hourly tick; isDue() reads the schedule from settings. Why: the assignment requires cadence changes and a two-week pause without a coder. A cron expression makes both a redeploy. Tradeoff: 167 no-op invocations a week, and the at-most-once guarantee moves into application logic (last_sent_week).

[!question]- ADR-002 — Resumable chunked runs, not one long job Alternatives: one long invocation; an external worker. Chosen: state machine in the runs row, batches of 4, self-re-invocation on time budget. Why: ~35 accounts cannot finish inside a serverless invocation. v1's own margin note identified this. Tradeoff: more moving parts, and the continuation fetch is fire-and-forget — a dropped continuation strands the run (D6).

[!question]- ADR-003 — Two model passes, on two different models Alternatives: one pass over all articles; pure scoring function; one model for both. Chosen: cheap fast model for per-account triage → strong model for the cross-account editorial cut, both through OpenRouter so the split is a config change. Why: "which four of these forty legitimate signals" is genuinely editorial and needs global context; "is this article about the right company" does not, and is 100× more frequent. Tradeoff: two failure surfaces; the editorial pass sees compressed summaries, not full articles.

[!question]- ADR-004 — Keyless baseline, NewsAPI.ai as the upgrade Alternatives: a single paid vendor as primary; newsapi.org. Chosen: Google News RSS + GDELT as the always-available baseline; NewsAPI.ai layered on when NEWSAPI_AI_KEY is set. Why: the tool must work for whoever clones it on day one with no vendor signup. But NewsAPI.ai is a genuine upgrade rather than more of the same — it returns full article bodies instead of a 200-character RSS teaser (which measurably improves triage), does its own syndication de-duplication, and supports per-language querying so LATAM/MENA/SEA accounts stop looking permanently quiet. newsapi.org was rejected: its free tier is development-only, ~24h delayed and capped. Tradeoff: two code paths for collection, and quota errors that arrive as HTTP 200 with an error body — handled explicitly and surfaced on source health.

[!question]- ADR-007 — A paid chain, not a paid source Alternatives: one paid vendor; no paid vendor; query every vendor every time and merge. Chosen: an ordered chain — NewsAPI.ai, then NewsAPI.org — where the first success wins, wrapped in retry and a circuit breaker, over a keyless baseline that always runs. Why: a keyed vendor is a single point of silent failure. Quota exhaustion and rate limiting return an empty result set, which is indistinguishable from a customer having a quiet fortnight — so the brief gets thinner and nobody learns why. Two vendors with independent failure surfaces, plus a baseline that needs no key at all, means the system degrades in steps instead of going dark. Querying everything every time was rejected: it doubles cost for heavily overlapping coverage. Tradeoff: more moving parts in collection, and the breaker is process-local — on serverless it is per-invocation, so it bounds waste within a run rather than across the week. Recorded here because it is a real limit, not a hidden one.

[!question]- ADR-005 — Persist every rejected candidate Chosen: write all judged articles to candidates with a dropped_reason. Why: the product is the filter. An unauditable filter cannot be trusted or tuned, and "why didn't I hear about X?" is unanswerable without it. Tradeoff: table growth; needs retention policy at scale.

[!question]- ADR-006 — Drop COMPANY WEBSITE scraping Chosen: removed from the design. Why: ~35 different layouts, JS-rendered newsrooms, bot walls, robots.txt questions — and it breaks silently, in exactly the way that matters: a scraper returning zero is indistinguishable from a customer having a quiet week. Unattended, it degrades to nothing within weeks and nobody notices. The signal it would carry (press releases, leadership changes) is already better covered by aggregators. Revisit if: a small set of strategic accounts justifies per-account maintained adapters with explicit health checks.


13. Deliberately not automated

Not automated Why
Routing to account owners No CRM access, so any mapping is a guess. Wrong routing burns trust once and permanently.
Acting on a signal It suggests a next step; it never drafts outreach or opens a task. A wrong automated action is worse than no action.
Padding a thin week The editor may return zero. Three real items beat four with a filler.
Auto-tuning thresholds from feedback Would overfit to a handful of clicks. A human moves the dial.
Human review before send Considered and rejected: a review queue nobody staffs is worse than none, and it defeats "assume you are not around". Mitigated instead by traceable sources and hallucination guardrails.

14. Requirements traceability

Req Requirement Where
R1 Weekly, unattended Scheduler + cron tick
R2 Risk / Opportunity split brief_items.category
R3 Four-minute read max_items, editorial cut
R4 Independently forwardable One card per item, self-contained
R5 Selects, does not aggregate Drop ladder + editor may return zero
R6 The five signal archetypes SIGNAL_TYPES taxonomy + triage prompt
R7 Tied to a tracked account brief_items.account_id
R8 Live source URL + date ✅ built, ⚠️ URL not validated (D7)
R9 The "so what" why_it_matters, suggested_action
R10 No repeats across editions ⬜ Not built (D9)
R11 Wire copies collapse to one similar() clustering
R12 Entity resolution controlled ⚠️ alias match only
R13 Defined zero-signal edition Empty-section copy
R14 Runs independent of the author Vercel + Postgres
R15–R19 Everything editable in a UI Control Plane
R20 On-demand run / re-send Run now, Send again
R21 Push to inbox Resend
R22 Access controlled ⚠️ fails open when unset
R23 Failures visible, non-silent ⚠️ pull-only; no alerting (D1)
R25 Bounded cost Per-run cost_usd
R26 Durable config Postgres
R27 Inspectable run record runs + candidates
R28 Answers "is it working" Observability Plane
R29 One-click per-item feedback HMAC verdict links
R30 Real list from locus.sh 35 seeded accounts
R37 Recipient timezone settings.timezone

15. Open questions for Locus

  1. Which of the 35 are live contracts? A logo wall is not a book of business. Tier is currently a public-information guess.
  2. Renewal dates and ARR. The two strongest prioritisation inputs, and neither is public.
  3. Where does an item go after the CRO forwards it? If it should land in a CRM or a Slack channel, that changes the Delivery Plane.
  4. Non-English coverage for LATAM / MENA / SEA accounts — worth the added noise, or accept the blind spot? (D11)

16. Next actions

  • D1 push alerting for failed runs, missed periods, partial sends
  • D2 duplicate-send guard, per-period run lock
  • D3 model failures no longer masquerade as quiet weeks
  • D4 per-recipient delivery state and retry
  • D5 catch-up window, D12 DST tolerance, D13 zoned pause, D14 midnight
  • D6 stale-run sweep and correct failure marking
  • D7 source-URL validation against the candidate set
  • D8/D9 deterministic shortlist with tier, renewal, escalation and novelty
  • D10 per-source health, D11 locale fan-out
  • account_owner, monitoring_keywords, exclude_terms, renewal_date, arr in the control plane
  • Real multi-article event clustering
  • Retention policy for runs.log and candidates
  • Route each item to its account owner once a CRM mapping exists

17. Provider chain and key pools

Both providers read a pool of up to five keys (OPENROUTER_API_KEY_1 … _5, NEWSAPI_AI_KEY_1 … _5), implemented once in src/lib/keypool.ts.

[!question]- ADR-009 — A provider chain, not a provider Alternatives: one provider; abstract behind a gateway product. Chosen: an ordered, individually-toggleable chain — OpenRouter, then Azure OpenAI — each with its own key pool, model mapping and error classifier. Why: key-level and provider-level failures are different problems with different fixes. Five keys do not help when the provider itself is down, the deployment was deprecated, or billing is suspended. Both providers speak the OpenAI wire format, so the second one costs an adapter rather than an integration. Tradeoff: two model configurations to keep straight, and quality changes when the chain falls back — so usage.servedBy and usage.degraded are recorded per call, and falling back raises an alert rather than passing silently. Azure also does not report per-call cost, so spend is marked unmeasured rather than estimated.

[!question]- ADR-008 — Key pools, not a key Alternatives: one key per provider; a paid plan large enough not to need it. Chosen: an ordered pool per provider with per-key health, failover, and classification of failures into transient vs fatal. Why: the two most common ways an API call dies — rate limit and exhausted credit — are both fixed by using a different key, and quota is per key. On a limited tier one key covers the first dozen accounts and then the brief silently thins out. Getting the transient/fatal split right is the crux: marking a 429 fatal burns a healthy key, and marking a 402 transient retries into the same wall on all 35 accounts. Tradeoff: state is process-local, so on serverless it bounds waste within a run rather than across the week. Dead keys are pushed out as alerts so they survive the invocation.

Behaviour Detail
Resolution Bare name plus _1…_5 and 1…5, de-duplicated
Ordering Lowest slot first, preferring keys with fewer recent failures
Transient failure 20s cooldown, next key tried immediately
Fatal failure 6h cooldown, source_down alert raised once per key
Single-key case Waits up to ~25s for a cooldown to lift instead of failing the account
All keys spent AllKeysExhaustedError, surfaced as a run failure with the upstream message

Verified by npm run check:keys — 11 assertions covering resolution, de-duplication, failover order, the transient/fatal split, once-only dead-key reporting, and exhaustion.


18. Provider notes

Model access goes through OpenRouter (src/lib/llm.ts), one key for any model.

That indirection buys something specific: the pipeline makes two very different kinds of call. Triage runs once per account — roughly 35 times a week — over a large, repetitive payload where the job is mostly "is this the right company and is this a corporate event". The editorial cut runs once and is the hardest judgement in the product. Routing them to different tiers is most of the cost story, and with OpenRouter it is two environment variables rather than two integrations.

Call Env var Default Volume
Triage OPENROUTER_TRIAGE_MODEL google/gemini-2.5-flash ~35 per run
Editorial cut OPENROUTER_EDITOR_MODEL anthropic/claude-sonnet-5 1 per run

Free OpenRouter models work but are a real downgrade, and were measured rather than assumed: nvidia/nemotron-nano-9b-v2:free returns valid structured output at zero cost, openai/gpt-oss-20b:free returned JSON that failed schema validation, and z-ai/glm-5.2:free was rate limited. The schema guard caught the bad one, which is the point of validating rather than trusting.

Two implementation details worth recording:

  • Strict schemas, validated twice. Zod schemas are converted to JSON Schema with every property required and additionalProperties: false, sent as response_format: json_schema, and then the response is re-validated against the Zod schema. A model that ignores the schema fails loudly instead of returning a plausible-looking empty brief — that is D3, and it is the failure mode most likely to go unnoticed.
  • Real cost, not an estimate. OpenRouter reports actual spend per call, so runs.cost_usd is money rather than a hardcoded price table that goes stale the next time a model is repriced.