Key Takeaways
- Of the four workflow components Touil names — hooks, MCP servers, sub-agents, and skills — skills are the only one where, in his words, "all of your knowhow is actually at the skills level."
- He built a six-month simulation of 15 teams of 5 to 12 people each, scoring skills per engineer, daily skill utilization, cross-team duplication ratio, and a quality-and-security ratio, to compare governed and ungoverned outcomes.
- On his timeline, Anthropic published the first skills article about eight months before the talk, an open standard landed two months later, and by February most agent harnesses had adopted it.
- Ungoverned skills create "a new class of technical debt" across seven failure modes: duplication, quality decay, discoverability, ownership, composability, security, and permissions.
Imad Touil, a distinguished engineer at QuantumBlack, opened his AI Engineer talk by asking the room for three shows of hands: who has created and used a skill, who shares skills within their team, and who governs and maintains skills across their organization. The first two drew enough hands for him to call it “amazing” and “great.” The third left, in his words, “a few hands” — and that gap is the whole talk. His argument: skills are where an organization’s know-how becomes executable, and almost nobody manages them like the shared infrastructure they have quietly become.
Four workflow components, and only one carries your know-how
Touil frames the agentic software stack as two loops. The inner loop is the coding-agent harness — context manager, tools and MCPs, memory and state, skills loader. The outer loop is workflows, built from four parts: hooks, MCP servers, sub-agents, and skills. Underneath sit enablement components: an environment sandbox, an MCP gateway, a model gateway spanning local and frontier models, a knowledge graph over core IT systems and the skills registry, and a workflow marketplace.
Then he strips three of the four workflow components of their glamour. Hooks are pre-triggers on events. MCP servers are mostly consumed rather than built: “tell me like who actually build a lot of MCPs — we just use MCP tools that is actually provided by the tool that we used to use before,” so “we don’t really own” them. Sub-agents exist “just to minimize the context window.” The protocol layer, in other words, is plumbing — not proprietary knowledge.
What is left is skills. “At the end of the day you will find all of your knowhow is actually at the skills level,” Touil said. “And if you don’t have the right structure of your skills then you’re not really having a deterministic workflow.” Workflows themselves he calls harness blueprints — artifacts that shape how your coding harness behaves at runtime.
The workflow you draw versus the one you actually have
Most coding agents are shaped around a four-step loop: specify, plan, break into tasks, implement. Touil’s point is that this familiar loop is one square on a much larger board — “this is like building a product increment.”
The real lifecycle, as he walks it, starts with product strategy — what to build, success metrics, roadmap — fed by market research, competitive analysis, and customer interviews. Then discovery: problem statements, solution design, validation, user stories. Then, before anyone writes a feature, data preparation and delivery: cleaning the catalog, wiring endpoints into core systems, building pipelines. Only then does the product increment loop run, followed by platform engineering and ops, launch, and incident resolution.
And that is one lifecycle. Drawing on 18 years serving organizations, Touil says every org he has worked with runs several in parallel — a mobile app SDLC here, a different platform there, an internal tool, a customer-facing product. “It is not like a one workflow that can actually build anything you want for your organization.” Even the sprawling diagram he shows is, by his estimate, “probably 10, 20% of what it is.” That mismatch is the same one behind an AI-native development team’s operating model.
Skills inherit their design rules from microservices
Touil is explicit that this is not new: “we have solved this with the microservices kind of movements.” Skills should be reusable, modular, and discoverable, so an engineer on one team can find what another built without asking around. Portable across workflows and harnesses, because everyone adopted the same standard: “If I’m having a skill on Claude Code and I want to move it to Cursor, it’s going to just work.”
They should be specialized, not monolithic — “you should not build like a one skill like a monolith” — and composable, so pulling several together does not create duplication or conflicts. And cost efficient: skills partly exist to solve a context-window problem, using progressive disclosure to load the right skills, in the right amount, at the right time. The result, he says, is a new unit that makes an organization’s know-how “executable, portable and cheap.”
His worked example is regulatory. A data retention policy skill sits alongside disclosure standards, GDPR rules, and fill-in templates; a regulatory disclosure review workflow pulls the relevant ones at runtime, and the output is deterministic — an audit report you can store, plus improvement findings that loop back into the codebase.
On evidence, Touil pointed to a recent skills benchmark that ran the same engineering and cybersecurity tasks against the latest models with and without skills. Without skills, models “did well,” he said, and that baseline “is going to continue to be improving day after day.” With skills applied, runs were more deterministic and “the outcome was clearly higher.” He cited direction rather than figures, so treat the magnitude as unquantified.
Seven ways ungoverned skills turn into debt
“If we don’t govern skills, we will start creating a new class of technical debt,” Touil said, then enumerated the failure modes.
Duplication — teams on the same stack, not talking to each other, build the same skill repeatedly. Quality — untested skills degrade, not only against the task but against each new model release. Discoverability — he reaches for the internal developer portal analogy: Backstage solved “who owns this microservice” with a catalog you query instead of a human you interrupt. Ownership — without an owner, nobody maintains a skill. Composability — it takes governance and a domain-driven approach to the catalog.
Security is where he lingers. Public skills get widely experimented with, and “some skills may have some prompt injection.” Skills also ship scripts — precisely the deterministic part of them — so without a checking pipeline “you may be pulling something that is insecure.” Permissions close the list: “not every skill is actually something that anyone in the organization should access,” since some encode sensitive business logic.
Individual, team, platform — then governance
Touil’s adoption ladder has three rungs. Individuals create, test, improve, and use skills — structured, on a tool the org has agreed on, not randomly. Teams then share and improve them, which happens fast because teammates build on the same stack. Third and critical: a centralized platform, with a metadata catalog that makes skills searchable, an MCP plugged into that catalog, and a CLI to pull skills into a local IDE or a sandbox. It also needs dependency tracking, versioning so an agent can notice a newer version and pull it, access control, and evaluation plus observability. The sequencing pairs with our guide on adopting coding agents at work.
Wrapping all of it is governance — “and this is where technology stops solving the problem.” Ownership depends on your org chart, but Touil names architects, engineering, infra, and cyber leads each holding part of the domain, keeping skills aligned with internal policy.
What the 15-team simulation showed
To make it concrete, Touil built a simulation: roughly 15 teams of 5 to 12 people, scored on skills contributed per engineer, daily skill utilization, cross-team duplication ratio, and a quality-and-security ratio, run forward six months.
In the ungoverned run, teams do create and use skills — “this is already happening within your organization” — but with no visibility, and skills are tightly coupled to productivity uplift. His example: without a regulation skill, “someone is vibe coding back and forth and trying to figure out exactly how to steer the agent to implement it properly,” burning tokens and time instead of getting it right in one shot. Maturity varies sharply between teams — he points to one at medium productivity, another at low-to-medium productivity with medium quality and security but cost that is “really high.”
Governed, the picture changes, though he refuses to oversell it: some teams still split off, and “this is reality, it isn’t going to be perfect.” What you get is common ground. Publish one skill, and when the next engineer starts building something similar, the harness finds the existing skill and pulls it instead. His caveat: skills are one component of workflows, and the same centralization argument applies to workflows themselves.
What happens next
Touil framed all of this as early — “this is just the start,” covering maybe six to eight months of ecosystem movement — and named three things to watch.
A skills registry is first, and he suggests you should already have one. The vendors that solved the internal developer portal problem are centralizing this capability, so “if you don’t have it today, maybe in a couple of months you will see it coming.”
Skills evaluation is second, and openly unsettled — there is “still kind of a discussion on what is the right approach.” The most valuable thing he has found is a static check against Anthropic’s own best practices: “If the skill is not invoked properly, if the skill is not structured properly, there’s a high chance that it’s not going to be high quality.” A cheap first gate, not a substitute for the harder measurement work in our evals guide for teams.
Auto-evolving skills is third, and he is pointedly ambivalent: “this is what everyone kind of like is the next hype right now,” a closed loop that evolves skills automatically. Start that machine without guardrails, he warns, and the impact is far larger than today’s mess. Auto-evolution amplifies whatever discipline, or lack of it, already exists in your catalog.
Quick poll
Where is your organization on the skills ladder?
When Touil asked his live audience who governs and maintains skills across their organization, he counted "a few hands."
FAQ
Who is Imad Touil and what was the talk? A distinguished engineer at QuantumBlack. His AI Engineer talk, “AI-Native Organisations Run on Skills: How to Structure and Scale Them,” argues that skills hold an organization’s executable know-how and need the governance microservices eventually got.
Why do skills matter more than MCP servers or sub-agents? Because the other components are generic. Hooks are event pre-triggers, MCP servers are mostly built by vendors rather than owned by you, and sub-agents mainly protect the context window. Skills are where domain-specific know-how lives.
How should you evaluate skills today? Touil says the right approach is still being debated. His practical starting point is statically checking skills against Anthropic’s published best practices, on the reasoning that a skill which is not invoked or structured properly is unlikely to be high quality.