Inspector V2 Working Group Meeting - Sept 9, 2026 #3340
cliffhall
started this conversation in
Meeting Notes - Inspector V2 WG
Replies: 1 comment
|
The rules vs. procedures split around skills is a nice change. Also liked the point that the repo holds up under agent-driven change mostly because of the boring stuff: architecture, CI gates, documentation, and explicit constraints. The “agent magic” part is much smaller than it looks from the outside. That feels increasingly important as agents start making larger changes. Good boundaries do more work than a smarter prompt. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Inspector V2 Working Group Meeting - Sept 9, 2026
Agenda
Attendees
Discussion
A release-day session. Most of it went on the machinery around the release — replacing Dependabot's pull requests with issues, folding
npm auditinto the release procedure, and a nightly watch on the SDKs — then on what a week of real work had shown about skills, two usability reports from @BobDickinson's gateway work, and a longer closing exchange on why this codebase holds up under agent-driven change.Skills shipped, and the composable test server carried it again. @BobDickinson opened on it: "the skill stuff looks nice." The SDKs don't yet support the server side of skills, and @cliffhall credited the composable test servers with simulating it in their absence — "just really a boon for us." Several people in the skills discussion are already working on SDK support, so within weeks there should be a wave of servers trying the functionality, and "it'll be nice that there is an in-org answer to being able to see if they've done it right."
Where the change list lives. Asked whether there's a rolling change list, @cliffhall pointed at the release page: every issue in the release, the full commit-level change log, and the smoke-test ledger attached to it. The
releaseskill had covered the steps but "didn't cover them exactly like they need to happen," so he formalized the procedure the night before. A second attachment — what skills adds and how it works, directory reading, the CLI, a TUI screenshot — was held back until it could be checked against what actually went out the door. @BobDickinson's suggestion: post it to the skills interest group.Dependabot now files issues, not PRs. The underlying mismatch:
mainis the default branch, so it's what Dependabot watches, and it holds the current release — what you get fromnpx— while the work lands onv2/main. And Dependabot only produces PRs, when the whole flow runs through the board, where every PR needs an issue because issues are what travel across it. @cliffhall had considered retargeting each Dependabot PR ontov2/mainwith a generated issue, but it was a hassle and it was unclear whether Dependabot would tolerate the branch changing underneath it. So PRs are off and alerts stay on, and a scheduled job reads the alerts and files an issue for each, worked like any other.The release now clears the decks with
npm audit. Alerts are computed againstmain, so mergingv2/maindown would surface anything not yet addressed. The release procedure therefore starts with annpm auditat high severity, before the version bump. Findings are fixed by PR againstv2/mainand merged into the release branch, so what lands onmainstays byte-identical tov2/mainand nothing ever needs merging back.npm audit fixis ruled out — it can resolve an advisory inside a caret range by silently downgrading. The expected steady state is that the alert sweep comes back empty.This release deferred one item: a dev-only, not-high vitest advisory, with an issue and PR already sitting ready for 2.7.0 — "not something to jam in there on release night." It's called out at the bottom of the ledger.
@BobDickinson agreed from experience. Because releases are punctuated, dependency updates matter at release time and should be staged and tested on
v2/main. The alert sweep againstmainis still worth having as a backstop: if something critical appears between releases, it may justify an early one, and "something changed in the world" can raise an alert even whenmainhasn't changed. And it removes a real cost he lives with at work: Dependabot running whenever it likes, five updates at ten minutes of CI each, and an hour where nobody can check anything in without rebasing. An alert that shows up while a release is being assembled simply goes to the next one.A nightly watch on the SDKs. Ola had long ago filed an issue about keeping up with SDK updates, from when the SDK was converging on 2.0 alongside the Inspector. @cliffhall had a card he kept bumping milestone to milestone, waiting on the TypeScript SDK and
ext-apps, and checking their versions by hand. There's now a nightly job that checks both upstreams for a new release and files an issue to update — "whenever they move, we'll catch it."Localhost subdomains need an SDK change. A reporter set on running both the Inspector and their MCP server on
localhostsubdomains got a lot of work done here but couldn't close it out: the SDK deliberately disallows it. @cliffhall added context to the existing SDK issue for whoever takes it up. In the meantime, they can set allowed origins on the MCP server side and get at least that part.The queue is shrinking. Around 30 issues went into this milestone and fewer remain; perhaps a quarter to a third were filed by @cliffhall himself as follow-ups. Triage runs every other day, and new issues aren't "coming in fast and furious." npm weekly downloads are holding and, per @BobDickinson, have climbed lately. @cliffhall — "a bugs before features guy" — is looking forward to releases made mostly of things the group decided to build, though "it was impossible to not put skills into this."
Modern-era servers mostly just changed the SDK. @BobDickinson's disappointment from the gateway side: almost everyone who has moved to the modern era did only the SDK upgrade and none of the modern work. Tool list caching gives it away. Its two settings — public vs. private (does everyone get this list, or does it depend on the token, as with GitHub?) and the cache timeout — are exactly what a gateway wants to know. Yet he can name a server's SDK from them, because the defaults differ: Go servers come back no-timeout and private, TypeScript servers no-timeout and public. "The intent was to actually spend five minutes and think about this and put the right flags in there, not just use the SDK default." Skills may fare better, since a server that wants them has to wire them in deliberately.
Notifications may be closer than they look. @BobDickinson asked whether notifications — triggers, events; "they use like three different words" — are on the short list. @cliffhall hasn't been tracking it closely. @BobDickinson's read echoes his earlier one on triggers: the SEP looked mid-discussion a month ago, but a launch-day presentation made it look essentially decided and simply not yet in the docs.
Extensions shouldn't need SDK changes. @cliffhall's understanding of the original idea was that extensions require no protocol modification, so he's unsure what SDK changes skills would actually need; any new methods, like listing skills, should come from the extension's own SDK. @BobDickinson agreed and filled in the mechanism: every SDK pairs high-level wrappers (
tools/list,tools/call) with a raw "compose the method and params and send it" path. An extension's methods can go through that path from any modern-era client or server, and SDKs may later promote the popular ones into wrappers.His narrower point was about the UI: Connection Info listed skills under extensions but not apps. That turned out to be the fixture — that composable server simply had no apps — while the client side still advertised
ui. @cliffhall noted a follow-up PR since added a directory section there as well.Long tool names truncate in the Tools list. Gateways prefix every tool with its server's name (
github_copilot_…then a double underscore), so the shared prefix fills the left-hand tool list and the real name is cut off. There's no hover and no way to widen the column, so the only way to tell tools apart is to click each one. @cliffhall first took it as the monitoring sidebar — closing that gives the protocol/network/console views full width with filtering — and noted that panel needs its own smallest-useful layout, plus a rule that PR screenshots be taken wide enough not to crop. On the actual tools list, both agreed a hover would help.A likely regression behind HTTPS-intercepting proxies. In @BobDickinson's corporate environment, a man-in-the-middle security proxy swaps certificates. Browsers and the shell trust it through the macOS keychain; Node doesn't, unless started with a CA-trust flag. Moving away from the old browser-side proxy to the API model put those requests in the Node process, so what used to work doesn't. He'll pin down the exact flag and open an issue; whether to enable it by default is open. @cliffhall: "if you've stubbed your toe on it, then I'm sure others will."
A week of skills in practice. This was the first full week after
AGENTS.mdwas split into skills, with a similar volume of work, and "the agents have done nothing weird" — they load skills when needed. The evals now also test second-order loading:issue-createsays an issue needs a priority and that the rubric lives in another skill, and the test checks that the second hop actually happens, in the evals and in practice. The finding he stressed: the description is the lever — it needs honing, starting with the verb ("use this when…") — and the eval prompts need the same tuning, since a description that fires 100% on questions nobody asks proves little.His conclusion reverses the one that kept him away from skills — a published report that models load them hit-or-miss. He'll apply the approach to a couple of client repos in a maintenance milestone next week. @BobDickinson: "you should write a blog about that." @cliffhall intends to, and owes the AIF some writing from his ambassadorship; an M5 Max studio on order should free up the cores he keeps pegging.
What makes the codebase hold. @cliffhall's view: the repo sustains a high volume of change without breaking because of the up-front work — the spec, an architecture that holds water, a record of what not to do learned from v1, and agent guidance that says how this repo writes TypeScript and maintains components — alongside how far the models have come. @BobDickinson's framing: what makes a codebase work for AI is what makes it work for humans — clean architecture, gates, CI coverage, documentation — and "it's 80% just basic hygiene and 20% agent magic." Humans move slowly enough to absorb the mess, so it shows less. @cliffhall added the four-dimension 90% coverage gate in CI as the regression guard — "like a diode" — and noted that unlike a new team member, "every single new agent is 100% up when they begin each new PR."
Humans still earn their keep. @BobDickinson spent two hours with Opus 4.8 on a subtle bug spanning four or five interconnected systems; it kept rat-holing, and he found the cause himself and suspects he'd have found it in ten minutes by reading the code. @cliffhall agreed that human knowledge of what keeps code robust over time is still the difference. He still reads the code so it doesn't get away from him, but treats a lapse — "a bare button with styles on it" — as a signal to fix the skill's wording, not to type in all caps. The skills extraction itself showed that most instructions had required a two-hop read the agent sometimes skipped; now the hops are factored out and testable.
2.6.0 is ready. @BobDickinson: "release looks good." @cliffhall will merge it, post it in the usual places, and let the skills group know, to catch anything missed quickly.
Next Steps
v2/main, with the alert sweep onmainas the backstop andnpm auditas a release step; watch whether notifications/triggers is stable enough to pick up.Transcript
Operational Details
v2/main- v2 Inspector, merged tomainfor milestone releasesv1/main- deprecated v1 Inspector, urgent security patches onlyv2v2Repo contributions
WG Leads:
All reactions