CWA is a specification for assembling everything a model call needs — instructions, state, evidence, history, the live turn — into one payload where every item has an owner, an authority, a lifecycle, a budget, and a place. The assembly is deterministic, and every step of it can be inspected.
Every model call is assembled by a function, and in most systems that function has no name. It has grown arguments for a year, it concatenates whatever the route happens to hold, and it knows nothing about who is allowed to direct the model, how old each piece is, what to shorten when the window is full, or what it sent last time. The failures this produces are not random: they repeat, they have recognisable shapes, and they can be named.
Where an item sits in the window matters too, and that question has evidence of its own: Why placement is a profile, not a rule.
"Evidence polluted governance" and "history displaced tool results" are sentences a team can act on. Named planes and slots turn vague prompt debates into specific bug reports.
Given the same request state, source versions, selection rules, and budget, the assembler emits the same payload in the same order. That is a claim a test can check. "The model will answer the same way" is not, and CWA never makes it.
Tokens are finite, so when every item is measured and tiered the assembler knows what to protect, what to compress, what to drop, and when to stop and report that the evidence did not fit.
Rule zero (R-5): a governance item may state a rule, but the application enforces it outside the model. No assembler relies on context to keep an invariant.
CWA gives you four planes and eleven typed slots, and no single request fills them all. The assembler admits the items each request actually needs, fits them to a finite budget in tier order, and renders them in a profile-defined order. This simulation uses estimated tokens and illustrative outcomes, not model evaluations; each system's route is taken to require evidence when the system uses an evidence slot. Remove the instructions or the query and it shows the refusal R-4 requires instead of a payload. A refused assembly has no payload, but its trace still records every decision made before the refusal (R-17); a snapshot that fails its schemas or checks is rejected before assembly, with neither a payload nor a trace. Watch seven very different systems assemble, then pull a slot out or squeeze the budget and see what degrades first.
{{ ucSymptom }}
{{ ucNote }}
A plane says why a piece of context exists. A slot is a typed collection inside a plane, with its own admission rules. An item is one governed object in a slot — a policy, a retrieved chunk, a tool observation. Newcomers learn the planes; implementers use the slots; the assembler moves the items.
{{ pl.blurb }}
{{ layer.purpose }}
{{ layer.value }}
Real systems put three retrieved chunks, two tool observations, and a memory summary plus a task episode into the same request. Each one is a separate item with its own metadata, and without that metadata even a clean slot order produces unpredictable payloads.
{
"id": "refunds-eu:v17#p4",
"slot": "evidence.knowledge",
"source": "policy-corpus",
"source_version": "2026-09-10",
"authority": "reference_only",
"trust": "verified",
"freshness": "2026-09-12T15:30:00Z",
"expires": "2026-12-12T15:30:00Z",
"scope": {
"tenant": "acme",
"task": "refund_request"
},
"relevance": 0.91,
"token_budget": 420,
"tier": "compressible",
"variants": [],
"conflict_policy": "defers",
"injection_risk": "untrusted_content",
"lineage": "verbatim",
"eligibility": "support-chat/illustrative/v1: tenant acme; rerank at least 0.5",
"body": "Pro plans refund in full within 30 days of purchase."
}
A fixed order reproduces a payload, but it says nothing about who is allowed to direct the model, what wins when sources disagree, whether this order is right for this model, or how anyone would know. The normative rules, the producer contract, the placement evidence and the implementations each have a page of their own.
"Retrieval goes in slot three" is memorable, but it is too coarse. Between the vector store and the payload there are eight steps: four the retriever does before assembly, and four the assembler does, each one a rule the trace can point at. A route can also opt a slot into supersession of stale observations, exact deduplication and a per-source cap (R-25, R-24, R-26); a slot that doesn't ask for them skips those steps. For high-stakes answers the slot holds evidence items, each carrying its chunk plus source, version, relevance and authority, rather than undifferentiated "knowledge."
{{ s.text }}
The profile picks each slot's tag and order, and a renderer the spec publishes writes the request. Message roles are not a style choice: only governance takes the system and tool roles, prior turns stay a transcript inside their tag, and the query is the only live user turn. Every item names its slot and authority, evidence arrives as separate items with provenance, and the trace records the profile that ordered them. The spec publishes two tokenizers, fixture-whitespace/v1 and estimate-utf8/v1, and two renderers, fixture-xml/v1 and cwa-messages/v1, and every implementation provides all four; an application may count with a tokenizer of its own under an ID no published tokenizer uses (R-16).
Each member's text as the model receives it. On the wire the request is one RFC 8785 JSON object, and each system and tools entry also carries its item's id. The token count covers each entry's text and the message content; roles, ids and JSON punctuation are left to the route's budget.margin_percent, 15 here (R-16).
Your existing tools supply the materials, and CWA is the assembly step between them and the model.
Retrieval produces candidates for the Evidence plane. CWA decides which are admitted, where they're placed, and what they may never overrule.
Read →MCP is the transport for tools and resources. CWA governs what it returns — tool specs proposed to the route's capability policy and admitted into Governance, observations in Evidence, each with freshness and authority.
Read →Frameworks orchestrate the loop. CWA assembles the payload each step of that loop actually sends, and emits the trace for it.
Read →Prompting is craft at the sentence level and CWA is structure at the payload level, so you need both.
Read →Your framework is the harness: it runs the loop, calls the tools, keeps the state, and retries the failures. It is deliberately unopinionated about the one thing that decides answer quality, which is what actually lands in the window on each pass. CWA fills that gap. Pick your stack:
The getting-started guide takes you in six steps from your current prompt to a payload with owners. It includes a migrator that proposes slots for your prompt and five files to copy: an integration sketch, a rendering template, a complete item example and JSON Schema, and a cwa.md for coding agents.
Name the slots you already send, give every item an authority and a source, and see how much of your payload had no owner. Three assemblers, in Python, TypeScript and Go, implement this spec, covering admission, fitting, profiles and traces, and pass every published conformance case and reject every published rejection; none is released yet, and we want to hear what they must do. That report is what an assembler's conformance claim rests on, while a producer's or an application's rests on its word, and an application must render with a renderer the published cases cover (§1).