| title | Locus Signal — High Level Design | ||||
|---|---|---|---|---|---|
| version | 2.5 | ||||
| status | implemented | ||||
| owner | Founder's Office assignment | ||||
| tags |
|
||||
| updated | 2026-08-22 |
[!abstract] The job Every week, tell the CRO what happened to Locus customers that changes whether they are about to churn or about to expand — split into risk and opportunity, readable in four minutes, forwardable one item at a time.
This is a precision product, not a coverage product. The value is in what it throws away. Four items out of four hundred is the specification; a system that surfaces forty has failed even if all forty are true.
The v1 diagram had the right spine and several good instincts. This revision keeps them and fixes four structural problems.
[!success] Kept from v1 — these were right
- Accounts as a data record at the top of the flow, not a list buried in the engine. This is what makes the account list editable.
- A
Statusfield rather than deletion. Churned customers stay as history.- Scheduling drawn as a layer above the engine, with
Pause / Resumenamed explicitly.- A cheap deterministic filter before the expensive AI stage. Correct and load-bearing.
- Event clustering placed before AI analysis. The strongest single call on the page — one event gets judged once instead of nine near-identical articles being judged nine times.
- Multiple independent sources rather than one vendor.
- Prioritisation as a stage distinct from classification.
- Field choices:
Aliases,Monitoring Keywords,Locus Use Case,Account Owner,Tier. Better than what is currently built.
[!danger] Structural problems in v1 1. The three certainties had no home. Account list, recipients and cadence were drawn as annotations on data stores. Nothing showed who edits them or how. The assignment's hardest requirement — "none of that should require you, or anybody who can code" — was invisible in the design. → Now a first-class Control Plane.
2. Cadence lived inside the cron job. If the cron expression encodes the schedule, moving Monday to Thursday is a code change and a redeploy. → The schedule now lives in the database; the cron is a dumb hourly heartbeat that asks "is anything due?".
3. The engine was one long synchronous job. The margin note — "usually going to take the time depending on the number of accounts" — is the sharpest observation on the page and it identifies a fatal flaw: on serverless it gets killed mid-run; on a long-lived host a restart loses the week. → The engine is now a resumable state machine with checkpoints.
4. Feedback was a terminal box with no outbound arrows. Drawn as the last thing the pipeline does rather than a system that observes the pipeline. → Now an Observability Plane that reads from every stage and closes the loop back to configuration.
Plus two pipeline-ordering corrections and one source removed — see [[#5. The pipeline]] and [[#12. Decision records|ADR-006]].
| # | Principle | Consequence |
|---|---|---|
| P1 | Precision over recall | Willing to ship two items, or zero. An empty section beats a manufactured one. |
| P2 | Config is data, never code | Every operational lever lives in Postgres and is edited in a UI. No redeploy to change behaviour. |
| P3 | Deterministic where possible, model where necessary | The model makes judgement calls. Code makes every rule that can be written down — cheaper, inspectable, fails predictably. |
| P4 | Every rejection is recorded | "Why didn't I hear about X?" must be answerable. The filter is auditable, not trusted. |
| P5 | Assume nobody is watching | Silent failure is the primary risk of a weekly job. Absence of a brief must be distinguishable from a quiet week. |
| P6 | The human keeps the judgement | The system triggers an intervention; it never performs one. |
flowchart TB
subgraph People["People"]
CRO["CRO and other recipients"]
OPS["Brief owner - commercial ops, non-technical"]
end
subgraph Control["Control Plane - auth-gated web app"]
UIACC["Accounts - add, edit, mute, aliases, tier, context"]
UIREC["Recipients - add, pause, remove"]
UISCH["Schedule - cadence, day, hour, timezone, pause, thresholds"]
UIRUN["Run now - preview or send"]
UIHEALTH["Is it working - funnel, drop reasons, blind spots, cost"]
end
subgraph Store["Durable state - Postgres"]
DB[("Postgres - accounts, recipients, settings, runs, candidates, briefs, brief_items, feedback, audit_log")]
end
subgraph Sched["Scheduler"]
TICK["Hourly cron tick - knows nothing but 'wake up'"]
DUE{"isDue - reads the schedule from the database"}
end
subgraph Engine["Run engine - resumable state machine"]
RUN["startRun - stepRun - finalizeRun - sendBrief"]
end
subgraph Ext["External services"]
SRC["News sources - paid chain NewsAPI.ai then NewsAPI.org, baseline Google News RSS plus GDELT"]
LLM["Model providers - OpenRouter then Azure OpenAI, 5-key pool each"]
MAIL["Resend - per-recipient send"]
end
OPS --> UIACC
OPS --> UIREC
OPS --> UISCH
OPS --> UIRUN
OPS --> UIHEALTH
UIACC --> DB
UIREC --> DB
UISCH --> DB
UIRUN --> RUN
DB --> UIHEALTH
TICK --> DUE
DUE -->|"due now and not already sent"| RUN
DUE -.->|"not due - no-op, 167 ticks a week"| TICK
DB <--> RUN
RUN <--> SRC
RUN <--> LLM
RUN --> MAIL
MAIL --> CRO
CRO -->|"one-click verdict"| DB
CRO -->|"forwards one item to the account owner"| OPS
DB -.->|"blind spots and quality signals"| OPS
[!note] Why four planes and not one pipeline The v1 diagram was a single flow because that is how the data moves. But the assignment is graded on the parts that are not the data flow: who can change the account list, what happens when it breaks, and how you prove it works. Those need to be visible as subsystems, not implied.
Everything a commercial user must be able to change, without a terminal.
| Surface | Handles | Status |
|---|---|---|
| Accounts | Add / edit / mute / remove. Aliases, tier, segment, region, customer-vs-prospect, free-text context | ✅ Built — src/app/accounts/page.tsx |
| Test aliases | Shows what news the current aliases find right now, pre-AI. Distinguishes a quiet account from a misconfigured one | ✅ Built — src/app/accounts/TestAliases.tsx |
| Recipients | Add / pause / remove. Pause ≠ delete — history is preserved | ✅ Built — src/app/recipients/page.tsx |
| Schedule | Cadence, day, hour, timezone, pause-until, and selectivity (max items, min severity, lookback) | ✅ Built — src/app/schedule/page.tsx |
| Run now | Preview or send on demand, with live progress | ✅ Built — src/app/RunButton.tsx |
| Access control | Password gate; cron endpoint guarded by shared secret | |
| Account owner + renewal date + ARR | Owner name/email, renewal date and ARR per account. Renewal proximity feeds the ranking score | ✅ Built — actions.ts, accounts/page.tsx |
| Monitoring keywords | Per-account watch terms. Become their own search query and raise entity-resolution confidence | ✅ Built — sources.ts |
| Exclusion terms | Per-account veto. How a commercial user fixes "wrong Landmark" without a developer | ✅ Built — resolveEntity in sources.ts |
| Markets / countries, Locus use case | Recorded on the account and fed to the triage prompt | ✅ Built |
| Alert resolution | Acknowledge an open alert from the health page | ✅ Built — resolveAlert |
[!tip] The design move that matters The schedule is a row in the database, not a cron expression. A fixed hourly cron wakes the app; the app reads
cadence,send_day,send_hour,timezoneandpaused_untiland decides. That single indirection is what turns "change the send day" from a deploy into a dropdown.
The engine. Covered in [[#5. The pipeline]].
Rendering and delivery are separate concerns — v1 merged them into one box. Delivery fails independently, and it fails per recipient.
| Concern | Status |
|---|---|
| Brief rendered as a durable web object, archived at a permalink | ✅ Built — src/lib/email.ts → renderBriefHtml |
| Per-recipient email so feedback links are attributable to a person | ✅ Built — loop in sendBrief |
| Plain-text alternative | ✅ Built — renderBriefText |
| Per-recipient delivery state and retry | ✅ Built — deliveries table; "Retry the ones that failed" on the brief page |
Covered in [[#9. Observability]].
flowchart TB
A["Account dequeued from runs.pending_ids - batch of 4"]:::code
B["Collect - paid chain first success wins, Google News per locale always, GDELT only when thin"]:::code
C["Normalise and exact dedupe - canonical URL, strips hash, utm, oc, hl, gl, ceid, ref"]:::code
D["Cheap pre-filter - junk domains, outside lookback window"]:::code
E["Entity resolution - alias must appear in title or snippet"]:::code
F["Event clustering - Jaccard 0.6 on title tokens, within one account"]:::code
G["Cap at 25 newest articles for this account"]:::code
H{{"MODEL CALL 1 - triageAccount, cheap model, one call per account"}}:::ai
I["Drop ladder - name collision, not revenue-relevant, neutral, below severity floor, confidence under 0.5"]:::code
J[("candidates - every judged article persisted, kept or dropped with a reason")]:::store
K["Editor pool - dropped_reason is null, order by severity then confidence, limit 60"]:::code
L{{"MODEL CALL 2 - selectItems, strong model, once per run, may legitimately return zero"}}:::ai
M["Assemble - match back to candidate by source_url, write briefs and brief_items"]:::code
N["Render - risk and opportunity halves, one forwardable card each"]:::code
O["Deliver - one email per active recipient, signed feedback links"]:::code
A --> B --> C --> D --> E --> F --> G --> H --> I --> J --> K --> L --> M --> N --> O
classDef code fill:#e8eef7,stroke:#1f3a5f,color:#10233d
classDef ai fill:#f3e8f7,stroke:#6b2d7a,color:#331239
classDef store fill:#eef7ee,stroke:#1e7a4f,color:#0f3a26
Legend — rectangle = deterministic code · hexagon = model call · cylinder = persisted state
[!warning] Correction 1 — split deduplication, move half of it upstream v1 ran
ENTITY RESOLUTION → DEDUPLICATION. That pays resolution cost on nine copies of the same wire story and then throws eight away. Exact/canonical dedupe is the cheapest stage in the pipeline and belongs at the front. Near-duplicate grouping is not a separate stage at all — it is event clustering.
[!warning] Correction 2 — collapse
RELEVANCE + AI ANALYSISandRISK / OPPORTUNITYv1 drew these as sequential stages. Relevance, direction, severity and why-it-matters are one judgement about one event. Two passes either duplicate the model call for no gain, or produce two verdicts with nothing to arbitrate between them. They are now a single structured triage call. (This is already how the implementation works — the correction is to the diagram.)
[!warning] Correction 3 — prioritisation cannot be one global step This is the most consequential flaw in v1. A flat top-4 over all candidates lets one noisy account take all four slots — which is a news digest about one company, precisely the failure the CRO named. It must be two stages: reduce each account to its strongest signal, then rank across accounts. Status: ✅ built.
shortlist()reduces each account to its strongest signal, scores across accounts by tier weight, renewal proximity and escalation, and hands the editor a pool of ~4x the target. One-item-per-account is then re-enforced in code after the editor returns, so it is a guarantee rather than a request.
| Stage | Kind | Implementation | Status |
|---|---|---|---|
| Collect | code | sources.ts — premiumNews chain (retry + circuit breaker) over newsApiAi then newsApiOrg, plus googleNews per locale and gdelt as gap-filler |
✅ |
| Canonical dedupe | code | canonicalUrl |
✅ |
| Pre-filter | code | junk domains, lookback cutoff | ✅ |
| Entity resolution | code | resolveEntity — scored: alias phrase 0.6, partial 0.3, domain corroboration +0.25, keyword +0.2, region +0.1; operator veto terms; floor 0.4 |
✅ |
| Event clustering | code | similar(), Jaccard ≥ 0.6 |
|
| Triage | AI | triageAccount via OpenRouter, OPENROUTER_TRIAGE_MODEL (default google/gemini-2.5-flash) |
✅ |
| Drop ladder | code | severity floor, confidence floor | ✅ |
| Per-account winner | code | shortlist() — one item per account guaranteed; a second only at severity 5 with a different signal type |
✅ |
| Novelty check | code | sent_signals over a 60-day window, by URL key and per-account title key |
✅ |
| Cross-account rank | code | tier weight x severity + confidence + renewal proximity + escalation |
✅ |
| Editorial cut | AI | selectItems via OpenRouter, OPENROUTER_EDITOR_MODEL (default anthropic/claude-sonnet-5), URL-validated against the candidate set |
✅ |
| Render + deliver | code | email.ts |
✅ |
This is the answer to the v1 margin note about runtime scaling with account count.
stateDiagram-v2
state "stage=collect" as collect
state "stage=select" as select
state "stage=send" as send
state "stage=done status=ok" as done_ok
state "stage=done status=failed" as done_failed
[*] --> collect : startRun writes pending_ids, period_key, should_send
collect --> collect : stepRun takes a batch of 4, writes candidates, trims pending_ids, commits
collect --> collect : time budget hit - handler self-re-invokes with the run id
collect --> select : pending_ids empty
select --> send : finalizeRun writes the brief - idempotent on re-entry
send --> done_ok : sendBrief succeeds, last_sent_week set
collect --> done_failed : failRun
select --> done_failed : failRun
send --> done_failed : failRun
done_ok --> [*]
done_failed --> [*]
[!info] Why chunked ~35 accounts means ~35 network fan-outs plus ~35 model calls. That does not fit in one serverless invocation. State lives in the
runsrow (stage,pending_ids), so any step can be retried or resumed after a crash without losing completed work.
[!success] Recovery
sweepStaleRuns()runs at the start of every fresh tick and fails any run whoselast_progress_atis over 30 minutes old, raising astale_runalert. Keying on progress rather than run age matters: a 200-account run spread across many continuations is legitimately long, but it always makes progress. A continuation tick deliberately skips the sweep so it cannot kill the run it is continuing.
flowchart TB
T["Hourly cron tick"] --> AUTH{"Authorised - bearer or query secret"}
AUTH -->|"no"| STOP1["401 - no run"]
AUTH -->|"yes"| S["Load schedule from settings, compute wall clock in the configured timezone"]
S --> P{"paused_until set and still in the past window"}
P -->|"yes"| STOP2["Skip - paused, UI shows the resume date"]
P -->|"no"| HR{"Current hour matches send_hour"}
HR -->|"no"| STOP3["Skip - not the send hour"]
HR -->|"yes"| DAY{"Cadence daily, or weekday matches send_day"}
DAY -->|"no"| STOP4["Skip - not the send day"]
DAY -->|"yes"| PAR{"Biweekly and ISO week number is odd"}
PAR -->|"yes"| STOP5["Skip - off week"]
PAR -->|"no"| SENT{"last_sent_week equals this period key"}
SENT -->|"yes"| STOP6["Skip - already sent, the at-most-once guard"]
SENT -->|"no"| GO["Run - collect, select, send, then write last_sent_week"]
[!success] Catch-up window The hour check is "at or after
send_hour, within 6 hours", clamped so it never crosses local midnight. One weekly attempt becomes several, which also absorbs the DST spring-forward skip.last_sent_weekkeeps repeated attempts idempotent, so extra attempts cannot produce a second email.missedPeriod()raises a critical alert if the window closes with nothing sent.
erDiagram
ACCOUNTS ||--o{ CANDIDATES : "judged for"
ACCOUNTS ||--o{ BRIEF_ITEMS : "subject of"
RUNS ||--o{ CANDIDATES : "produced"
RUNS ||--o| BRIEFS : "produced"
BRIEFS ||--o{ BRIEF_ITEMS : "contains"
BRIEF_ITEMS ||--o{ FEEDBACK : "rated by"
ACCOUNTS {
int id PK
text name UK
text_array aliases
text domain
text region
text segment
text tier
text kind
text status
text context
}
RUNS {
int id PK
text stage
text status
int_array pending_ids
text period_key
int articles_found
int candidates
numeric cost_usd
jsonb log
}
CANDIDATES {
int id PK
int run_id FK
text title
text url
text category
int severity
text dropped_reason
boolean selected
}
BRIEFS {
int id PK
date week_of
text headline
text status
int sent_to
}
BRIEF_ITEMS {
int id PK
int brief_id FK
text category
int severity
text headline
text why_it_matters
text suggested_action
text source_url
}
FEEDBACK {
int id PK
int brief_item_id FK
text verdict
text reader_email
}
[!note] The
candidatestable is the product It stores every article the model judged, kept or dropped, with the reason. That single table is what makes the filter auditable rather than a black box, and it is what answers "why didn't I hear about X?".
Account fields still to add (all from v1, all good): markets_countries, locus_use_case, monitoring_keywords, account_owner, plus renewal_date and arr — the two things a CRO actually prioritises on.
flowchart TB
ITEM["brief_items row shipped"] --> MAIL["Per-recipient email with HMAC-signed verdict links"]
MAIL --> READ["Reader clicks one of three verdicts"]
READ --> V1["useful - right item"]
READ --> V2["not_useful - precision problem"]
READ --> V3["already_knew - latency problem"]
V1 --> FB[("feedback")]
V2 --> FB
V3 --> FB
RUNS[("runs and candidates - funnel, drop reasons, cost")] --> HEALTH
BR[("briefs - delivery status")] --> HEALTH
FB --> HEALTH["Is it working surface"]
HEALTH --> Q1["Precision - share rated useful"]
HEALTH --> Q2["Latency - share rated already knew"]
HEALTH --> Q3["Blind spots - active accounts never surfaced"]
Q1 --> HUMAN{"Human tuning decision"}
Q2 --> HUMAN
Q3 --> ALIAS["Test aliases - what would have matched"]
ALIAS --> HUMAN
HUMAN -->|"adjust min_severity, max_items, lookback"| SET[("settings")]
HUMAN -->|"fix aliases, sharpen context, retier"| ACC[("accounts")]
SET --> NEXT["Next scheduled run"]
ACC --> NEXT
NEXT --> ITEM
Three questions that are usually collapsed into one:
| Question | Measure |
|---|---|
| Does it run? | Run success rate, last run, cost per brief |
| Does it land? | Briefs emailed, send failures, per-recipient activity |
| Is the judgement good? | Share rated useful; already_knew tracked separately from not_useful — the first means too slow, the second means wrong. Different fixes. |
Plus two things easy to miss: blind spots (active accounts that have never produced a signal — usually wrong aliases, not a quiet customer) and the per-run audit trail.
[!warning] The loop is deliberately open No rating is read back into any prompt, threshold or ranking automatically. The return path is a person making a change in the UI. Auto-tuning a precision filter on a handful of weekly clicks would overfit to noise within a month.
[!success] Push alerting
lib/alerts.tsraises deduplicated alerts for run failure, abandoned runs, missed periods, partial sends, analysis failures, hallucinated source URLs and zero-item briefs, then emails them toALERT_EMAIL(falling back to the first active recipients). Every tick flushes pending alerts, so 167 of the 168 weekly ticks exist purely to notice breakage. Open alerts also surface at the top of/healthwith a Resolve action.
Identified by an SRE review of the implementation, then fixed. Severity is impact on an unattended six-month run.
| # | Severity | Failure | Status |
|---|---|---|---|
| D1 | 🔴 Critical | No alerting — /health was pull-only, so nobody would learn the brief had stopped |
✅ lib/alerts.ts, deduplicated + emailed, flushed every tick |
| D2 | 🔴 Critical | Duplicate email — any replay of the resume URL re-sent the whole brief | ✅ deliveries gates per recipient; uniq_scheduled_run_per_period blocks a second scheduled sending run per period (manual runs deliberately exempt) |
| D3 | 🔴 Critical | Model parse failure looked like a quiet week — parsed_output ?? [] turned an API failure into an empty success |
✅ AnalysisError thrown on null output or refusal; recorded against the account and alerted |
| D4 | 🔴 Critical | Partial send never retried — failed recipients were recorded then dropped | ✅ Per-recipient state + "Retry the ones that failed" |
| D5 | 🔴 Critical | Missed send had no catch-up — exactly one attempt per week | ✅ 6-hour catch-up window; missedPeriod() alerts if it closes empty |
| D6 | 🟠 Major | Run never marked failed on the first invocation — failRun(resume,…) where resume was 0 |
✅ runId hoisted out of the try; sweepStaleRuns() catches orphans via last_progress_at |
| D7 | 🟠 Major | Editor could invent a source_url |
✅ Every returned item validated against the candidate set; unmatched items dropped and alerted |
| D8 | 🟠 Major | Candidate pool truncated to 60 ignoring tier | ✅ Replaced by shortlist(): tier weight × severity + confidence + renewal proximity + escalation |
| D9 | 🟠 Major | No cross-run novelty — last week's story could return | ✅ sent_signals, 60-day window, URL key + per-account title key; recorded at send, not at build, so previews do not poison it |
| D10 | 🟠 Major | Source 429 was silent — an account contributed nothing and looked quiet | ✅ source_health per run, surfaced on /health |
| D11 | 🟠 Major | Region captured but unused — non-English news invisible | ✅ REGION_LOCALES fans the broad query across locales (es-419/MX, ar/EG, id/ID, en-AE, en-SG, en-IN) alongside en-US |
| D12 | 🟡 Minor | DST spring-forward skipped the send hour | ✅ Absorbed by the D5 catch-up window |
| D13 | 🟡 Minor | paused_until evaluated in server time |
✅ Compared as a zoned calendar date |
| D14 | 🔴 Critical | send_hour = 0 never fired — hour12:false renders midnight as "24", so 24 !== 0. Reachable from the UI |
✅ % 24 in zonedParts |
All fourteen are covered by npm test (14 assertions over the scheduling, catch-up, pause and missed-period logic).
[!example] Why D14 mattered more than it looked It is the exact shape of failure this assignment tests for. A user picks midnight from a dropdown; the brief silently never sends; no error, no failed run, nothing to notice. Found by an SRE review pass, then confirmed empirically before being believed.
| Item | Why it is acceptable |
|---|---|
| Event clustering is near-duplicate only | Jaccard over titles collapses wire copies of one story. True multi-article event clustering (same event, different framing, different headline) is not built. The triage pass absorbs most of the cost, and severity ranking usually surfaces the strongest telling. |
Auth fails open when ADMIN_PASSWORD is unset |
Deliberate so a fresh clone runs with zero configuration. Documented in the README as a deploy requirement. |
Unbounded runs.log JSONB |
Grows per batch. Needs a retention policy well before it matters at this account count. |
| Biweekly parity uses ISO week number | Can skip an extra week at a year boundary. Cosmetic at fortnightly cadence. |
Numbers from constants in the code — batch size 4, TIME_BUDGET_MS 220s, maxDuration 300s, 25-article cap, LIMIT 60.
| Accounts | Model cost / run | Wall clock | Verdict |
|---|---|---|---|
| 35 (today) | ~$0.12/account ≈ $4 | ~9 batches, minutes | Comfortable |
| 200 | ~$24 | ~50 batches of self-re-invocation | Needs a real queue; the continuation chain becomes the weak link |
| 1000 | ~$120 | — | Needs worker infrastructure, per-source quota management, tiered scan frequency |
[!question]- ADR-001 — Schedule in the database, not the cron expression Alternatives: cron expression in
vercel.json; a scheduling vendor. Chosen: fixed hourly tick;isDue()reads the schedule fromsettings. Why: the assignment requires cadence changes and a two-week pause without a coder. A cron expression makes both a redeploy. Tradeoff: 167 no-op invocations a week, and the at-most-once guarantee moves into application logic (last_sent_week).
[!question]- ADR-002 — Resumable chunked runs, not one long job Alternatives: one long invocation; an external worker. Chosen: state machine in the
runsrow, batches of 4, self-re-invocation on time budget. Why: ~35 accounts cannot finish inside a serverless invocation. v1's own margin note identified this. Tradeoff: more moving parts, and the continuationfetchis fire-and-forget — a dropped continuation strands the run (D6).
[!question]- ADR-003 — Two model passes, on two different models Alternatives: one pass over all articles; pure scoring function; one model for both. Chosen: cheap fast model for per-account triage → strong model for the cross-account editorial cut, both through OpenRouter so the split is a config change. Why: "which four of these forty legitimate signals" is genuinely editorial and needs global context; "is this article about the right company" does not, and is 100× more frequent. Tradeoff: two failure surfaces; the editorial pass sees compressed summaries, not full articles.
[!question]- ADR-004 — Keyless baseline, NewsAPI.ai as the upgrade Alternatives: a single paid vendor as primary; newsapi.org. Chosen: Google News RSS + GDELT as the always-available baseline; NewsAPI.ai layered on when
NEWSAPI_AI_KEYis set. Why: the tool must work for whoever clones it on day one with no vendor signup. But NewsAPI.ai is a genuine upgrade rather than more of the same — it returns full article bodies instead of a 200-character RSS teaser (which measurably improves triage), does its own syndication de-duplication, and supports per-language querying so LATAM/MENA/SEA accounts stop looking permanently quiet. newsapi.org was rejected: its free tier is development-only, ~24h delayed and capped. Tradeoff: two code paths for collection, and quota errors that arrive as HTTP 200 with an error body — handled explicitly and surfaced on source health.
[!question]- ADR-007 — A paid chain, not a paid source Alternatives: one paid vendor; no paid vendor; query every vendor every time and merge. Chosen: an ordered chain — NewsAPI.ai, then NewsAPI.org — where the first success wins, wrapped in retry and a circuit breaker, over a keyless baseline that always runs. Why: a keyed vendor is a single point of silent failure. Quota exhaustion and rate limiting return an empty result set, which is indistinguishable from a customer having a quiet fortnight — so the brief gets thinner and nobody learns why. Two vendors with independent failure surfaces, plus a baseline that needs no key at all, means the system degrades in steps instead of going dark. Querying everything every time was rejected: it doubles cost for heavily overlapping coverage. Tradeoff: more moving parts in collection, and the breaker is process-local — on serverless it is per-invocation, so it bounds waste within a run rather than across the week. Recorded here because it is a real limit, not a hidden one.
[!question]- ADR-005 — Persist every rejected candidate Chosen: write all judged articles to
candidateswith adropped_reason. Why: the product is the filter. An unauditable filter cannot be trusted or tuned, and "why didn't I hear about X?" is unanswerable without it. Tradeoff: table growth; needs retention policy at scale.
[!question]- ADR-006 — Drop
COMPANY WEBSITEscraping Chosen: removed from the design. Why: ~35 different layouts, JS-rendered newsrooms, bot walls, robots.txt questions — and it breaks silently, in exactly the way that matters: a scraper returning zero is indistinguishable from a customer having a quiet week. Unattended, it degrades to nothing within weeks and nobody notices. The signal it would carry (press releases, leadership changes) is already better covered by aggregators. Revisit if: a small set of strategic accounts justifies per-account maintained adapters with explicit health checks.
| Not automated | Why |
|---|---|
| Routing to account owners | No CRM access, so any mapping is a guess. Wrong routing burns trust once and permanently. |
| Acting on a signal | It suggests a next step; it never drafts outreach or opens a task. A wrong automated action is worse than no action. |
| Padding a thin week | The editor may return zero. Three real items beat four with a filler. |
| Auto-tuning thresholds from feedback | Would overfit to a handful of clicks. A human moves the dial. |
| Human review before send | Considered and rejected: a review queue nobody staffs is worse than none, and it defeats "assume you are not around". Mitigated instead by traceable sources and hallucination guardrails. |
| Req | Requirement | Where |
|---|---|---|
| R1 | Weekly, unattended | Scheduler + cron tick |
| R2 | Risk / Opportunity split | brief_items.category |
| R3 | Four-minute read | max_items, editorial cut |
| R4 | Independently forwardable | One card per item, self-contained |
| R5 | Selects, does not aggregate | Drop ladder + editor may return zero |
| R6 | The five signal archetypes | SIGNAL_TYPES taxonomy + triage prompt |
| R7 | Tied to a tracked account | brief_items.account_id |
| R8 | Live source URL + date | ✅ built, |
| R9 | The "so what" | why_it_matters, suggested_action |
| R10 | No repeats across editions | ⬜ Not built (D9) |
| R11 | Wire copies collapse to one | similar() clustering |
| R12 | Entity resolution controlled | |
| R13 | Defined zero-signal edition | Empty-section copy |
| R14 | Runs independent of the author | Vercel + Postgres |
| R15–R19 | Everything editable in a UI | Control Plane |
| R20 | On-demand run / re-send | Run now, Send again |
| R21 | Push to inbox | Resend |
| R22 | Access controlled | |
| R23 | Failures visible, non-silent | |
| R25 | Bounded cost | Per-run cost_usd |
| R26 | Durable config | Postgres |
| R27 | Inspectable run record | runs + candidates |
| R28 | Answers "is it working" | Observability Plane |
| R29 | One-click per-item feedback | HMAC verdict links |
| R30 | Real list from locus.sh | 35 seeded accounts |
| R37 | Recipient timezone | settings.timezone |
- Which of the 35 are live contracts? A logo wall is not a book of business. Tier is currently a public-information guess.
- Renewal dates and ARR. The two strongest prioritisation inputs, and neither is public.
- Where does an item go after the CRO forwards it? If it should land in a CRM or a Slack channel, that changes the Delivery Plane.
- Non-English coverage for LATAM / MENA / SEA accounts — worth the added noise, or accept the blind spot? (D11)
- D1 push alerting for failed runs, missed periods, partial sends
- D2 duplicate-send guard, per-period run lock
- D3 model failures no longer masquerade as quiet weeks
- D4 per-recipient delivery state and retry
- D5 catch-up window, D12 DST tolerance, D13 zoned pause, D14 midnight
- D6 stale-run sweep and correct failure marking
- D7 source-URL validation against the candidate set
- D8/D9 deterministic shortlist with tier, renewal, escalation and novelty
- D10 per-source health, D11 locale fan-out
-
account_owner,monitoring_keywords,exclude_terms,renewal_date,arrin the control plane - Real multi-article event clustering
- Retention policy for
runs.logandcandidates - Route each item to its account owner once a CRM mapping exists
Both providers read a pool of up to five keys (OPENROUTER_API_KEY_1 … _5,
NEWSAPI_AI_KEY_1 … _5), implemented once in src/lib/keypool.ts.
[!question]- ADR-009 — A provider chain, not a provider Alternatives: one provider; abstract behind a gateway product. Chosen: an ordered, individually-toggleable chain — OpenRouter, then Azure OpenAI — each with its own key pool, model mapping and error classifier. Why: key-level and provider-level failures are different problems with different fixes. Five keys do not help when the provider itself is down, the deployment was deprecated, or billing is suspended. Both providers speak the OpenAI wire format, so the second one costs an adapter rather than an integration. Tradeoff: two model configurations to keep straight, and quality changes when the chain falls back — so
usage.servedByandusage.degradedare recorded per call, and falling back raises an alert rather than passing silently. Azure also does not report per-call cost, so spend is marked unmeasured rather than estimated.
[!question]- ADR-008 — Key pools, not a key Alternatives: one key per provider; a paid plan large enough not to need it. Chosen: an ordered pool per provider with per-key health, failover, and classification of failures into transient vs fatal. Why: the two most common ways an API call dies — rate limit and exhausted credit — are both fixed by using a different key, and quota is per key. On a limited tier one key covers the first dozen accounts and then the brief silently thins out. Getting the transient/fatal split right is the crux: marking a
429fatal burns a healthy key, and marking a402transient retries into the same wall on all 35 accounts. Tradeoff: state is process-local, so on serverless it bounds waste within a run rather than across the week. Dead keys are pushed out as alerts so they survive the invocation.
| Behaviour | Detail |
|---|---|
| Resolution | Bare name plus _1…_5 and 1…5, de-duplicated |
| Ordering | Lowest slot first, preferring keys with fewer recent failures |
| Transient failure | 20s cooldown, next key tried immediately |
| Fatal failure | 6h cooldown, source_down alert raised once per key |
| Single-key case | Waits up to ~25s for a cooldown to lift instead of failing the account |
| All keys spent | AllKeysExhaustedError, surfaced as a run failure with the upstream message |
Verified by npm run check:keys — 11 assertions covering resolution,
de-duplication, failover order, the transient/fatal split, once-only dead-key
reporting, and exhaustion.
Model access goes through OpenRouter (src/lib/llm.ts), one key for any model.
That indirection buys something specific: the pipeline makes two very different kinds of call. Triage runs once per account — roughly 35 times a week — over a large, repetitive payload where the job is mostly "is this the right company and is this a corporate event". The editorial cut runs once and is the hardest judgement in the product. Routing them to different tiers is most of the cost story, and with OpenRouter it is two environment variables rather than two integrations.
| Call | Env var | Default | Volume |
|---|---|---|---|
| Triage | OPENROUTER_TRIAGE_MODEL |
google/gemini-2.5-flash |
~35 per run |
| Editorial cut | OPENROUTER_EDITOR_MODEL |
anthropic/claude-sonnet-5 |
1 per run |
Free OpenRouter models work but are a real downgrade, and were measured rather
than assumed: nvidia/nemotron-nano-9b-v2:free returns valid structured output
at zero cost, openai/gpt-oss-20b:free returned JSON that failed schema
validation, and z-ai/glm-5.2:free was rate limited. The schema guard caught the
bad one, which is the point of validating rather than trusting.
Two implementation details worth recording:
- Strict schemas, validated twice. Zod schemas are converted to JSON Schema
with every property required and
additionalProperties: false, sent asresponse_format: json_schema, and then the response is re-validated against the Zod schema. A model that ignores the schema fails loudly instead of returning a plausible-looking empty brief — that is D3, and it is the failure mode most likely to go unnoticed. - Real cost, not an estimate. OpenRouter reports actual spend per call, so
runs.cost_usdis money rather than a hardcoded price table that goes stale the next time a model is repriced.