Skills Over MCP Working Group - August 25th 2026 Meeting Notes #3307
Replies: 1 comment
|
On section 7 — a third thing that sits between the digest and the provenance, with numbers rather than intuition. Arsham's split is the right one: a digest says the files were not modified, provenance says whether you trust the publisher. What neither covers is that the publisher can change after both are established, while the bytes stay identical. Ownership of a package moves, the repository link is removed, a name is unpublished and republished by someone else — the digest still matches, a signature made before the handover still verifies, and the entity that controls the next version is different. We run a daily external diff of the whole MCP registry (25,122 entries, series from 2026-07-30) and this is measurable:
That is a direct measurement of what Arsham said in passing — that packages from trusted maintainers have been modified and backdoored. It is not folklore; it is roughly two in five of the handovers we see, and none of them would trip a digest check. On Jonathan's point about digests for skills specifically, our data suggests the case is stronger for skills than for servers: For a skill the instruction text is the executable surface — it enters the model's context and steers behaviour. One in seven rewrites we observed moved no version at all, so a consumer pinning by version sees nothing. For MCP servers, where versioning is universal and well-formed, a version pin is already a decent proxy; for skills it demonstrably is not. That asymmetry seems worth carrying into the sidecar design rather than treating both surfaces the same. On Vijaydeep's question (enterprise skills served alongside skills pulled from a central repository): the same asymmetry bites there. A per-skill digest tells the client which bytes it got, but if the central repository's upstream changed hands, nothing in the served artifact reflects that. Whatever carries provenance probably needs a way to say when the publisher assertion was made, so a client can notice it is old — a signature with no freshness is indistinguishable from a fresh one. No ask attached. Not proposing a mechanism and not proposing ourselves as one — Peter asked whether implementors had run into the trust-boundary breaking down, and this is what it looks like measured from the outside. The underlying series is public and free ( |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Repo: modelcontextprotocol/experimental-ext-skills
Charter: Skills Over MCP WG
Discord: #skills-over-mcp-wg
SEP: SEP-2640 (Skills Extension) (V1 proposal, Extensions Track)
Context: Peter walked through the substantive changes made to SEP-2640 in response to core-maintainer feedback and asked for blockers before putting it up for vote. The rest of the session covered how the new progressive discovery effort relates to this group, what happens to this group after V1 ships, skill provenance and signatures, and a set of open review comments on the PR.
Attendees
(Note: other attendees were present but didn't appear in the transcript - please leave a reply if you want to be added)
1. SEP-2640 updates from core-maintainer feedback
Peter summarized the changes made to the proposal since the last round of core-maintainer review. Most were clarifications, but three are substantive enough to flag:
SKILL.mdis fetched when the skill is loaded and supporting files when they are read. Hosts should not materialize the whole skill manifest onto the file system or anywhere else, and should cache what they do retrieve, validated against the entry digests.SKILL.mdincluded) and 16 MiB total per skill, and eachresourcesentry carries a requiredsizefield so both limits can be checked from the listing before any file is fetched. Peter surveyed roughly 10,000 popular skills and the limits covered essentially all of them; about two did not fit, one because it shipped a large video file.Other changes that landed on the branch between August 19 and 22, for anyone who last read the SEP before this session: an explicit
"resources": "dynamic"marker for skills that cannot publish stable digests (commit); a definition of the window in which a host is "acting on" a skill (commit); write-isolation or re-verification requirements for disk caches (commit); a clarification that a plainresources/readofSKILL.mddoes not load the skill (commit); andresultTypeon all sample responses per the 2026-07-28 spec. The full design rationale was moved out of the SEP into the WG repo as docs/rationale.md, in PR #124, which also synced the baseline copy with the SEP text.Peter asked for any concerns and said that absent blockers he would put the SEP up for a core-maintainer vote today. No objections were raised.
Update after the meeting: Peter posted in Discord that the extension is back up for the core-maintainer vote, with a result expected in a week.
2. Bulk download: normative text, no enforcement
Vijaydeep asked whether the "don't download everything" guidance is a restriction or a recommendation, since a client can still do it. Peter's answer: the SEP text is a MUST NOT, but the group is not going to police network requests, so in practice it functions as a recommendation. A client that tries this against a server with a large number of skills should expect heavy rate limiting (a server with two skills will be fine). The useful part of having it in the extension is that when client A causes problems for server B, the server author can point at the extension text and say the client is doing it wrong.
3. On-demand loading and latency
Arsham asked for clarification on the expected exchange: the server lists skills, the client knows they exist without having them on the machine, and the client downloads a given skill later based on what it is doing. Peter confirmed that is the model. He expects most clients to list skills up front, put them into an in-context index, and then use an "activate skill" or "load skill" tool (which most harnesses already have) to pull the
SKILL.mdand supporting files when needed. The SEP's host integration sketch describes this flow.Arsham raised latency: if the model is mid-task and has to go fetch a skill over a slow network or from an overloaded server, it sits and waits. Peter's response: clients may cache downloaded files and use the digests to keep caches in sync, which helps on every fetch after the first. First-fetch latency does go up, but most skills are very small and the round trip should be a couple hundred milliseconds or less, especially in region. For a very large skill on first call there could be noticeable latency, which is not really different from a tool call.
4. Are the size limits configurable?
Arsham asked whether clients are expected to make the 16 MiB limit configurable, for example so a large enterprise could serve a 100 MB skill internally. Peter: clients can lift the limit as much as they want. The prescribed limit exists so server authors know what constraints to work within to get the broadest client support. Clients will set their own limits, but those should be above what the spec prescribes. (The SEP text: hosts MUST support skills up to the limits and MAY support larger ones; servers SHOULD NOT exceed them.)
5. Search, tagging, namespacing, and progressive discovery
Search. Answering a question Samuel raised in chat, Peter said the extension supports an ad hoc pattern in the interim: a server exposes a tool such as
search_skillthat takes a query, runs the search server-side, and returns a skill URI the model can then load. An official mechanism will come from the progressive discovery work, which is spinning up as part of the 2026 roadmap refresh. On the roadmap (updated August 22), progressive discovery sits under priority area 4, "Improved Primitives", owned by the Core Primitives WG, which is forming during this roadmap period. Ola noted it is the reboot of what was formerly the primitive grouping effort and will send links.Namespacing and identity. Arsham asked whether identifying skills by URI changes how a client like Claude Code namespaces skills and tools per server. Peter explained that Claude Code's read-resource tool takes two parameters, the server name and the URI, so URI fetches are already namespaced that way. The identity matters in two places: cache keys should include the server ID, and slash-command invocation needs namespacing (something like
/github-create-project) because names can collide across servers. Samuel had also raised that nothing in the SEP disallows two skills with the same name within a single server, so clients have to disambiguate that case as well; the SEP's Names section requires hosts to disambiguate rather than discard or prefer one.Tagging. Tara asked in chat whether the extension supports tagging skills for filtering. Peter: not within this extension. It is something he expects the progressive discovery work to look at.
How the two efforts relate. Tara asked how the progressive discovery work will feed back into skills over MCP. Her team Implemented skills over MCP last year with some differences from the current SEP, and they rely on namespacing, tagging, and versioning, so they want to understand how those primitives will apply to skills specifically. Peter clarified the structure: this group is focused on getting the SEP accepted, hopefully going to vote today and becoming an official extension this week. Progressive discovery is one of the problems maintainers want solved by the end of the year, with another version of the MCP spec tentatively targeted for December. The group doing that work is an open working group like this one, focused primarily on tool calling but expected to consider skills and other primitives too.
Ola added that there is a related set of questions about what metadata belongs in the skill itself. Right now the extension passes front matter through without controlling its spec, and some of the follow-ups belong in a skills-spec discussion rather than here. One current example is the
dependenciesfrontmatter proposal raised on the PR, which has since been cross-posted to the Agent Skills repo. Bloomberg's implementation, which Sam has described in past meetings, does a lot with custom front matter, and Ola pointed people at the group's past meeting notes and the WG repo's skill_metakeys doc for details. Peter encouraged anyone with an implementation that is working well to feed it into the progressive discovery work, and said clients and servers adding their own metadata tags in the meantime is exactly how the group learns what works in practice.Update after the meeting: Ola posted the links in Discord: the Primitive Grouping interest group, which is the group being re-branded, the Core Primitives working group channel, and the roadmap post covering the topics that group is an umbrella for.
6. Roadmap and the future of this group
Peter, seeing no major objections, confirmed he will put the SEP up for vote. Ola summarized the roadmap she shared previously: get the extension accepted, then follow up on documentation and gathering post-acceptance feedback. Much of the follow-up work will split out into the other MCP groups. She wants to identify who is interested in carrying specific items forward, especially people who have context from implementing this or being in these discussions.
In the Discord agenda post before the meeting she named the two areas needing the most help next: example implementations, especially from client and host implementers, and updating the repo's host implementation guidelines (governance around access and approvals, threat model) once the SEP text is final. The roadmap she referred to is a Google doc; she asked whether to put a copy in the repo for commenting and later reference.
Ola expects the working group to be retired in its current form and transitioned to an Interest Group once V1 ships, with a channel kept open (as with other extensions) and continued work on documentation and implementation. The goal is to keep open threads moving forward and to stay connected to the broader problems being solved. (Group lifecycle rules, including retirement and IG formation, are in the Working and Interest Groups governance page.)
Arsham asked whether the file system work ties into the download story. Peter: it could. The groups are all very new and directions are not completely set, but he expects overlap; for example, that group might add batch download per primitive that this extension could build on. Server-hosted file systems are what that group is looking at, so it will intersect with this work. On the roadmap this is the File Uploads WG, which continues on scoped file operations and filesystem-like resource semantics (range reads, hierarchical listing). Ola noted this extension had to add a directory listing because none existed, which is really a virtual file system problem and belongs with that more general group (see why a directory read method in the rationale doc).
7. Provenance, signatures, and digests
Peter raised a topic from earlier core-maintainer discussions: provenance, meaning who authored a skill and whether to trust them. Today trust is implicit, so trusting a server means trusting its skills. That breaks down when the server is a hosting platform (GitHub serving skills from other people's repos) where blanket trust is not what you want. The extension currently says it is not a package manager, but this could be a direction where it adds value. He asked whether implementers have hit this. Ola said she has heard a lot of interest in the topic and is interested in it herself.
Aditya mentioned the question originated in the context of gateways that host many skills. The question is whether the trust boundary sits with a provider that signs a skill ("this is a GitHub skill, we signed it for you") or with the MCP server layer, where the gateway is responsible for serving the right skills to end systems. The concrete ask was to allow-list a set of digests rather than only deny-list, and more precisely, a signature rather than a digest. This comes up with large banks and consulting firms that don't want unvetted skills exposed even if an employee wrote one; they want to be able to say these are approved and everything else is not.
Arsham separated the two problems. A digest tells you the files you downloaded were not modified in transit, on the client, or on the server. Provenance is whether you trust the publisher, which is harder and more subjective, and as the open source world has shown, trusted maintainers' packages have still been modified to include back doors. Peter agreed: the digest tells you that you got what you asked for. (The SEP says this explicitly: digests are unsigned and supplied by the same server that supplies the content, and hosts MUST NOT treat a digest match as a security boundary.) He suggested a solution might live in the Agent Skills specification itself, for example as a signature in front matter.
Ola pointed to C2PA content provenance, which has mostly been discussed for images but now has additions for markdown, and noted someone is looking to formalize how to carry that kind of information in MCP metadata. A markdown-oriented version applied to skills could be interesting, and some people are already experimenting with it.
Jonathan mentioned that digests and signatures for skills have come up several times in the Agent Skills repo discussions, and he has commented on a few. It is a useful feature, but with caveats:
Vijaydeep asked how publisher certification works when an enterprise writes its own skills and serves them from the same MCP server as skills pulled from a central repository, so a client can tell which to trust. Arsham restated the scenario: an MCP server with some skills published by the enterprise itself and others from an external source, and the question is how to filter which are trusted. Jonathan's position: no spec should require a signature or digest; it has to be fully optional. In other words: the spec can say where signatures go, but what to accept comes down to the organization or product configuration. Aditya agreed with the balance: too prescriptive and you lose adoption, too permissive and enterprises can't use it at all.
Next steps
Questions or follow-ups? Join us in #skills-over-mcp-wg on Discord.
Prepared from an auto-generated meeting transcript; please reply with any corrections.
All reactions