# Donkey Development Kit — full documentation
> An SDK for consuming Agent Fabric governance from Python agent code, without adopting Mule.
Source: https://docs.donkey-kit.dev/ · Index: https://docs.donkey-kit.dev/llms.txt
---
Source: https://docs.donkey-kit.dev/index.md
Governed by the gateway. Understood by your code. DDK
brings Omni Gateway awareness into the agents you already write — in
LangGraph, Google ADK, Strands, CrewAI, LlamaIndex, the OpenAI Agents SDK,
the Anthropic SDK or Microsoft Agent Framework — so governance is not
just enforced on your agents, it is respected by them.
>
}
media={
}
actions={[
{ label: 'Quickstart', href: '/quickstart', primary: true },
{ label: 'Feature overview', href: '/feature-overview' },
{ label: 'Examples', href: '/examples' },
]}
/>
## Why gateway awareness
Enterprises put a gateway in front of their AI traffic for good reasons: one
place to authenticate every caller, meter every token, block personal data,
stop prompt injection and attribute cost to the team that spent it. MuleSoft
**Omni Gateway**, managed through **Agent Fabric**, does exactly that — for
LLMs, MCP servers, APIs and agents alike.
But a gateway that governs alone only solves half the problem. The agent on the
other side of the wire sees an opaque `403` or `429` and does what code does
with errors it does not understand: it retries a PII block as if it were a
network blip, burns the remaining budget on calls that will be refused, crashes
a nightly batch at 2am, and leaves no trace that connects its own run to the
gateway's audit log. Control is one-sided, and the cost of that shows up as
wasted tokens, broken runs and incident tickets.
**DDK makes governance a collaboration.** It brings the gateway's view of the
world into your agent code:
- a PII refusal arrives as a typed `PIIDetected`, not a generic HTTP error, so
the agent can redact and continue instead of retrying;
- the remaining token budget is an object the agent can **pace** against,
instead of a limit it discovers by failing;
- every call carries a correlation ID and emits an OpenTelemetry span, tagged
with your cost dimensions, that lines up with the gateway's own audit trail;
- and you can rehearse all of it on your laptop, against a local simulator,
before an agent ever meets production policy.
The gateway stays the enforcement point — it always has the final word. What
changes is that your agents become **good citizens** of the platform: aware of
the rules, efficient within them, and observable end to end. That is the
enterprise vision behind DDK — control at the proxy, efficiency in the agent —
built for AI engineers, developers and the AI teams who have to run their
agents in production.
Authentication, cost attribution, PII blocking and budget limits are set
once at the gateway and inherited by every app that uses DDK.
The same MuleSoft API management plane that already governs thousands of
enterprise APIs, now covering model traffic. No second gateway to buy,
staff or audit.
One import, your framework, your IDE, your code. DDK returns native
framework objects, not wrappers — eight frameworks, three lines to eject.
Governed model access today, with governed tool access, agent-to-agent
calls and publishing to the Agent Fabric registry on the same foundation.
**DDK is an open-source, community-driven project** (Apache-2.0). It is **not
an official Salesforce or MuleSoft product** and is not supported by
Salesforce. "Agent Fabric", "Anypoint", "MuleSoft" and "Omni Gateway" are
Salesforce trademarks; DDK uses them only to describe the platform it
connects to. Meet the people behind it on the [Team](https://docs.donkey-kit.dev/community/team.md) page.
## Architecture
DDK sits inside your agent process. Today it speaks to the platform on two
fronts: governed calls through the gateway, and telemetry to your
observability stack. A third front, publishing assets to the control plane,
is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md).
**The AI control plane** is where the platform team manages the AI estate:
the agent registry, cost control, gateway federation, and governance and
observability across every runtime.
**Omni Gateway** is the single data-plane entry point for APIs, MCP servers,
LLMs and agents. Its policies do the enforcing — authentication and identity,
token budgets and rate limits, PII detection, prompt-injection and
content-safety guardrails, model routing and fallback, audit trails, contract
drift and tool-poisoning detection — before traffic reaches the managed
upstreams: enterprise APIs and integrations, MCP servers, LLM providers, and
other AI apps and agents.
**DDK** is the developer-side half of that picture. It is wired into the
framework client your agent already uses and adds:
- **Governed calls and typed refusals** — every model request goes through the
gateway with consumer credentials, correlation and attribution headers
injected; every policy rejection comes back as a typed exception such as
`PIIDetected` or `TokenBudgetExceeded`, never confused with an auth error.
- **Token budget awareness** — the gateway's rate-limit headers become a
`Budget` object with `remaining`, `pace()` and `wait_for_reset()`.
- **Local testing** — `donkey mock` and `donkey.simulate()` replay real
gateway rejection shapes on your laptop, and a pytest conformance suite
proves your agent handles each one before it ships.
- **OpenTelemetry GenAI spans** — each call emits a span carrying the policy
decision, policy type, budget and correlation ID, exported to whatever
observability stack you run (Grafana, Datadog, Jaeger, and others) and
joined to the gateway's audit record through the correlation ID.
- **Registry and agent-to-agent** Roadmap —
planned, not shipped: scanning your code to publish tools and agent cards
to the control plane, discovering governed MCP tools, and serving or
exposing your agent to other agents over A2A.
The division of labour is deliberate. The gateway enforces; DDK makes the
enforcement legible and actionable inside the agent. Nothing in DDK
re-implements a policy client-side, and nothing in your process can override
the gateway.
## Before and after
Take the most ordinary piece of agent code there is: one model call.
**Without DDK**, a stock client talks straight to the model provider. It works
— and it is invisible. There are no centralised controls or policy
enforcement, no usage tracking, no record of which team or agent spent which
tokens, and nothing an auditor can follow.
```python
import openai
client = openai.OpenAI()
completion = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Capital of Switzerland?"}],
)
print(completion.choices[0].message.content)
```
**With DDK**, the call is the same shape and the client is the same native
`openai.OpenAI` — but it now goes through your organisation's Omni Gateway.
The request is authenticated and attributed, policy is applied, the agent
knows how much budget it has left, and the platform team sees the usage per
model and per consumer in Agent Fabric.
```python
from donkey_kit import Donkey
with Donkey.from_env() as donkey:
client = donkey.llm.client(sync=True) # a real openai.OpenAI, governed
reply = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Capital of Switzerland?"}],
)
print(reply.choices[0].message.content)
print(donkey.budget.remaining, "tokens left")
```
Three lines changed. Governed, observable and attributed — without leaving
your framework.
## Make a governed call
From Python, DDK hands you your framework's own objects. From any other
language, call the same governed proxy over its OpenAI-compatible HTTP API
with the `client_id` / `client_secret` header pair.
```python
from donkey_kit.integrations.langgraph import chat_model
# A real langchain_openai.ChatOpenAI, already pointed at your governed proxy.
model = chat_model("gpt-4o", temperature=0)
reply = await model.ainvoke([("user", "Explain quantum computing in simple terms.")])
print(reply.content)
```
```python
from donkey_kit import Donkey
# A real openai.OpenAI, already pointed at your governed proxy — no event loop.
with Donkey.from_env() as donkey:
client = donkey.llm.client(sync=True)
reply = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms."}],
)
print(reply.choices[0].message.content)
```
```typescript
const base = process.env.DONKEY_LLM_PROXY_URL!; // ends in "/", no /v1
const resp = await fetch(`${base}chat/completions`, {
method: "POST",
headers: {
"content-type": "application/json",
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
body: JSON.stringify({
model: "gpt-4o",
messages: [{ role: "user", content: "Explain quantum computing in simple terms." }],
}),
});
console.log((await resp.json()).choices[0].message.content);
```
```bash
curl "${DONKEY_LLM_PROXY_URL}chat/completions" \
-H "content-type: application/json" \
-H "client_id: ${DONKEY_LLM_PROXY_CLIENT_ID}" \
-H "client_secret: ${DONKEY_LLM_PROXY_CLIENT_SECRET}" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Explain quantum computing in simple terms."}]
}'
```
## Native objects, never wrappers
> **Adapters return the framework's native object — never a wrapper.**
`donkey.langgraph.chat_model("gpt-4o")` returns a real `langchain_openai.ChatOpenAI`.
`donkey.llamaindex.llm("gpt-4o")` returns a real `OpenAILike`. Hand them
straight to `create_agent`, a LlamaIndex query engine or a Strands `Agent`.
And if you ever want to drop DDK, you eject to three lines of native
constructor code — every [framework page](https://docs.donkey-kit.dev/frameworks.md) shows exactly which
three.
A stock client with a `base_url` and two headers can reach the gateway. What
DDK adds is the **single place in your process** where every request enters
and every response leaves — which is where typed refusals, budget pacing,
correlation IDs, cost tags, spans and simulation all attach without you
wiring each one.
## Next steps
Your first governed call in minutes — no gateway or credentials needed.
Every capability, its purpose, and what it saves you.
Install and quickstart for each of the eight supported frameworks.
Runnable demos, from a first governed call to a full LangGraph agent.
What is available now and what is coming next, phase by phase.
DDK is open source — issues, docs, examples and adapters welcome.
---
Source: https://docs.donkey-kit.dev/quickstart.md
# Quickstart
Make your first governed model call, then see what a plain `base_url` can't
give you — a **typed refusal**, a **budget** and a **span** — in a few
minutes, on your laptop. You start against DDK's **local gateway simulator**,
so you need no Anypoint account and no credentials. Switching to your real
Omni Gateway later is an environment change, not a code change.
## What you can use
| Capability | Status | Learn more |
|---|---|---|
| Governed model access (the raw client, through the proxy data plane) | Live | [Frameworks](https://docs.donkey-kit.dev/frameworks.md) |
| Framework adapters for 8 frameworks | Conformance-tested LangGraph · Live-verified ADK `gemini()` · Signature-confirmed the rest | [Frameworks](https://docs.donkey-kit.dev/frameworks.md#supported-frameworks) |
| Typed refusals | Live | [Typed refusals](https://docs.donkey-kit.dev/errors.md) |
| Budget & pacing | Live | [Budget & pacing](https://docs.donkey-kit.dev/budget.md) |
| OpenTelemetry spans & cost attribution | Live | [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md) |
| Local simulator & `simulate()` | Live | [Local simulator](https://docs.donkey-kit.dev/simulator.md) |
| Conformance testing with pytest | Live | [Testing & conformance](https://docs.donkey-kit.dev/testing.md) |
| CLI (`init`, `doctor`, `mock`, `test`) & decorators | Live | [CLI & decorators](https://docs.donkey-kit.dev/cli.md) |
| Governed tool access (MCP) | Roadmap | [Tool access](https://docs.donkey-kit.dev/tool-access.md) |
| A2A agents (`serve`, `expose`, `dev`) | Roadmap | [A2A agents](https://docs.donkey-kit.dev/a2a.md) |
| On-behalf-of identity | Roadmap | [Identity](https://docs.donkey-kit.dev/identity.md) |
| Human-in-the-loop | Roadmap | [Human-in-the-loop](https://docs.donkey-kit.dev/hitl.md) |
| Scan & publish to the registry | Roadmap | [Scan & publish](https://docs.donkey-kit.dev/publishing.md) |
| Policy handshake | Roadmap | [Policy handshake](https://docs.donkey-kit.dev/policies.md) |
## Your first governed call
### Install
```bash
pip install "donkey-kit[llm,local,otel]"
```
`llm` adds the OpenAI client, `local` the gateway simulator, `otel` the
OpenTelemetry SDK. Quote the extras — zsh treats unquoted brackets as a glob.
### Start the local gateway
In one terminal:
```bash
donkey mock
```
The simulator listens on `127.0.0.1:8080` and replays real gateway responses,
including every refusal shape. It enforces no policy and ignores
credentials, and every response carries `x-donkey-simulator: true`.
Port `8080` already taken? Run `donkey mock --port 9000` — see
[Choosing a port or host](https://docs.donkey-kit.dev/simulator.md#choosing-a-port-or-host).
"Local gateway" here means DDK's pure-Python simulator, not Omni/Flex
Gateway running in Local Mode — that deployment target isn't supported
(see the [unsupported boundary](https://docs.donkey-kit.dev/reference/unsupported-boundary.md)).
### Point DDK at it
In a second terminal:
```bash
export DONKEY_LLM_PROXY_URL="http://127.0.0.1:8080"
export DONKEY_LLM_PROXY_CLIENT_ID="local" # placeholder — the simulator ignores auth
export DONKEY_LLM_PROXY_CLIENT_SECRET="local"
```
### Make the call
```python filename="hello.py"
from donkey_kit import Donkey
with Donkey.from_env() as donkey:
client = donkey.llm.client(sync=True) # a real openai.OpenAI, governed
reply = client.responses.create(
model="gpt-5.1",
input="Tell me a one-sentence bedtime story about a unicorn.",
)
print(reply.output_text)
print("budget remaining:", donkey.budget.remaining, "tokens")
```
```text
A sleepy unicorn named Luma painted soft silver stars across the night sky with her glowing horn, then curled up on a moonbeam so all the children below could fall asleep beneath her gentle, sparkling light.
budget remaining: 99500 tokens
```
`donkey.llm.client()` returns the OpenAI SDK's own client, routed through
DDK — so credentials, correlation and attribution headers are injected, and
`donkey.budget` is updated from the gateway's response. The simulator replays
a response captured from a real gateway, so you get this story whatever you
ask; against your Omni Gateway the model answers your actual prompt.
Prefer one command? Run `python -m examples.quickstart.main` from the
`python/` directory of the
[SDK repository](https://github.com/Donkey-Development-Kit/donkey-development-kit/tree/main/python/examples/quickstart):
it boots the simulator for you and runs the steps above end to end.
## Catch a typed refusal Live
Ask the simulator for a specific rejection by using the model id
`donkey-sim/`, then bridge the raw response into DDK's taxonomy with
`classify()`:
```python
import openai
from donkey_kit import PIIDetected
from donkey_kit.core.errors import classify
try:
client.responses.create(
model="donkey-sim/pii-detected",
input="Email the report to jane.doe@example.com.",
)
except openai.APIStatusError as exc:
governed = classify(exc.response)
if isinstance(governed, PIIDetected):
print("blocked, entities:", governed.entities)
```
```text
blocked, entities: ['Email']
```
A PII block is not an auth error, and a policy `429` must never be retried —
typed refusals let the agent react correctly. The framework adapters
(`donkey.langgraph`, …) classify for you. See [Typed refusals](https://docs.donkey-kit.dev/errors.md).
## Pace against the budget Live
The gateway's token window is a first-class object. Keep a reserve in hand
and wait for the window to reset instead of failing:
```python
from donkey_kit import BudgetReserveReached
async with Donkey.from_env() as donkey:
while True:
try:
async with donkey.budget.pace(reserve=0.05):
await enrich(batch)
break
except BudgetReserveReached:
await donkey.budget.wait_for_reset()
```
See [Budget & pacing](https://docs.donkey-kit.dev/budget.md).
## See a span Live
With `otel` installed, every governed call emits an OpenTelemetry GenAI span.
Print them to the console:
```python
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import ConsoleSpanExporter, SimpleSpanProcessor
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(ConsoleSpanExporter()))
trace.set_tracer_provider(provider) # set before the first call
```
Put these lines at the top of `hello.py` and run it again. Alongside the
story, the console prints the span:
```text
{
"name": "donkey.llm.chat",
"context": {
"trace_id": "0x604e2f805df7f12cfdc402f83a66c8f4",
"span_id": "0x212735c846ec8f18",
"trace_state": "[]"
},
"kind": "SpanKind.INTERNAL",
"parent_id": null,
"status": {
"status_code": "UNSET"
},
"attributes": {
"gen_ai.request.model": "gpt-5.1",
"gen_ai.system": "openai",
"gen_ai.response.model": "gpt-5.1",
"donkey.routing.type": "ModelBased",
"donkey.routing.fallback": false,
"gen_ai.usage.input_tokens": 17,
"gen_ai.usage.output_tokens": 51,
"donkey.usage.cached_tokens": 0,
"donkey.usage.cache_write_tokens": 0,
"donkey.usage.reasoning_tokens": 0,
"donkey.policy.decision": "allow",
"donkey.budget.remaining": 99000,
"donkey.correlation_id": "e2fec0b686694055846da9446fa24c96"
},
...
}
```
Each `donkey.llm.chat` span carries `gen_ai.usage.*` token counts,
`donkey.policy.decision`, `donkey.budget.remaining` and the correlation ID.
Set `OTEL_EXPORTER_OTLP_ENDPOINT` and DDK exports over OTLP with no code at
all. See [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md).
## Attribute cost to a run Live
Group one logical task under a single correlation ID and cost tags. The
correlation ID joins the gateway record, your logs and your spans; the cost
tags land on your spans:
```python
async with donkey.run(id=ticket.id, team="support", project="triage"):
await triage_agent.ainvoke(ticket)
```
Or fold it into one line with a decorator:
```python
@donkey.governed(team="support")
async def handle_ticket(ticket): ...
```
See [CLI & decorators](https://docs.donkey-kit.dev/cli.md).
## Test the refusal branch Live
Inject a real gateway refusal in-process — no server, no network:
```python
from donkey_kit import PIIDetected
with donkey.simulate(PIIDetected):
await agent.ainvoke(...) # the next call fails as a real PIIDetected
```
Then grade your own agent against every refusal shape with the pytest
conformance suite. See [Testing & conformance](https://docs.donkey-kit.dev/testing.md).
## Connect to your Omni Gateway
The same code runs against your real gateway — stop the simulator and set
three values. The proxy authenticates on a `client_id` / `client_secret`
header pair (consumer auth), not a bearer token.
```bash
export DONKEY_LLM_PROXY_URL="https:////" # note: no /v1
export DONKEY_LLM_PROXY_CLIENT_ID=""
export DONKEY_LLM_PROXY_CLIENT_SECRET=""
```
Where the values come from:
- **[Create the model proxy](https://docs.mulesoft.com/general/model-proxy-create-model-proxy)**
on Omni Gateway with **Format=OpenAI**. Its consumer endpoint is your
`DONKEY_LLM_PROXY_URL`.
- **[Request access](https://docs.mulesoft.com/exchange/to-request-access)**
to the proxy in Exchange. That registers a client application and issues the
`client_id` / `client_secret` pair.
Then check the setup:
```bash
donkey doctor
```
```text
[ok] config env (3 fields)
[i] llm endpoint (env)
[i] control plane anypoint.mulesoft.com (default)
[ok] gateway reachable, responded
[ok] credentials client_id accepted
[ok] model accepted by the proxy
[i] budget 99,000 / 100,000 remaining, resets in 59s, observed 0s ago
```
`donkey doctor` tells a wrong URL from wrong credentials from a model that is
not on the proxy's allow-list, instead of one opaque failure.
## Next steps
- **[Pick your framework](https://docs.donkey-kit.dev/frameworks.md)** — get a native LangGraph, ADK,
Strands, LlamaIndex, CrewAI, OpenAI Agents SDK, Anthropic SDK or Agent
Framework object in three lines.
- **[Examples](https://docs.donkey-kit.dev/examples.md)** — runnable demos for every capability on this page.
- **[Scenarios](https://docs.donkey-kit.dev/scenarios.md)** — support triage, a nightly batch and an internal
copilot, built end to end.
---
Source: https://docs.donkey-kit.dev/feature-overview.md
# Feature overview
DDK is **gateway-aware**: control stays at the proxy, efficiency moves into the
agent. The gateway enforces policy; DDK makes each decision visible and
actionable in your code, so the agent reacts to a refusal, paces its budget and
reports what it did.
Everything below hangs off a single shared transport inside your process — the
one place where every request enters and every response leaves. Each capability
attaches there once, so you never wire it call by call.
## Model access
**Purpose:** point any of eight agent frameworks at your governed Omni
Gateway proxy. **Advantage:** you get your framework's own object —
`ChatOpenAI`, `LiteLlm`, `OpenAIModel`, `crewai.BaseLLM` … — with credentials,
correlation, attribution and retry policy injected. No wrapper to code
around, three lines to eject.
**Purpose:** use the OpenAI SDK directly. **Advantage:**
`donkey.llm.client()` returns a native `AsyncOpenAI` (or `OpenAI` with
`sync=True`) — Chat Completions and Responses, streaming included — governed
on identical terms.
| Framework | Call | Returns |
|---|---|---|
| LangGraph | `donkey.langgraph.chat_model("gpt-4o")` | `langchain_openai.ChatOpenAI` |
| Google ADK | `donkey.adk.model("gpt-4o")` | `LiteLlm` |
| Google ADK on a `Format=Gemini` proxy | `donkey.adk.gemini("gemini-2.5-flash")` | `google.adk.models.Gemini` |
| Strands | `donkey.strands.model("gpt-4o")` | `OpenAIModel` |
| MS Agent Framework | `donkey.agent_framework.chat_client("gpt-4o")` | Agent Framework chat client |
| LlamaIndex | `donkey.llamaindex.llm("gpt-4o")` | `OpenAILike` |
| OpenAI Agents SDK | `donkey.openai_agents.model("gpt-4o")` | `OpenAIChatCompletionsModel` |
| Anthropic SDK | `donkey.anthropic.client()` | `anthropic.AsyncAnthropic` |
| CrewAI | `donkey.crewai.llm("gpt-4o")` | `crewai.BaseLLM` |
Every adapter offers the same governed connection three ways — a factory on a
shared `Donkey` (`donkey.langgraph.chat_model(...)`), a module-level factory
(`from donkey_kit.integrations.langgraph import chat_model`), or
`connection_kwargs()` when you want to build the native object yourself.
## Governance
**Purpose:** turn every gateway rejection into a typed exception —
`PIIDetected`, `TokenBudgetExceeded`, `PromptInjectionBlocked`,
`ContentSafetyBlocked`, `AuthError`, `GatewayUnavailable` and more.
**Goal:** branch on the governance outcome, not on a parsed error body.
**Advantage:** a PII block is never mistaken for an auth failure, and a
policy `429` is never retried.
**Purpose:** expose the gateway's token window as a `Budget` object.
**Goal:** stop *before* the limit, not after it. **Advantage:**
`pace(reserve=)` and `wait_for_reset()` let an overnight batch slow down
instead of dying at 2am.
**Purpose:** on-behalf-of token exchange. **Goal:** per-user policy
reaches the gateway. **Advantage:** requests never silently fall back to
the service identity.
**Purpose:** one vocabulary for "pause and ask a human". **Advantage:**
mapped onto each framework's native interrupt, so approval flows look the
same everywhere.
**Purpose:** read the policy set in force. **Advantage:** skip calls that
are certain to be refused. Advisory only — the gateway always wins.
## Observability
**Purpose:** one span per governed call, using the GenAI semantic
conventions plus a stable `donkey.*` namespace — policy decision, policy
type, budget, routing and token usage. **Advantage:** refused calls still
produce a span, streaming produces exactly one, and prompt content stays
out by default. Zero-config OTLP export.
**Purpose:** `donkey.run(id=…, team=…, project=…)` binds one correlation
ID and validated cost tags to every call in a task. **Advantage:** your
log line, your span and the gateway's audit record join on the same ID.
**Purpose:** `donkey.last_call` records which gateway served the call,
how it was routed and what it used. **Advantage:** detect a model
substitution — or make it raise — instead of discovering it in a bill.
## Developer tooling
**Purpose:** `donkey mock` replays real gateway responses and refusal
shapes on `127.0.0.1`. **Advantage:** build and demo against governance
without an Anypoint account or credentials.
**Purpose:** `donkey.simulate()` injects a refusal in-process, and a
pytest plugin grades *your* agent against every refusal shape.
**Advantage:** the PII branch is tested before production, not in it.
**Purpose:** `donkey init`, `doctor`, `mock` and `test`, plus
`@donkey.governed` and `@donkey.tool`. **Advantage:** `doctor` tells wrong
credentials from wrong URL from model-not-allowed; one decorator gives a
function a run scope, span and typed refusals.
**Purpose:** the docs are published as `llms.txt` and per-page markdown.
**Advantage:** Cursor, Claude Code and other assistants write correct DDK
code from the source.
## Registry & catalog
**Purpose:** discover governed MCP tools from the catalog and bind them as
native framework tools. **Advantage:** allow/deny filtering, pinning and a
lockfile — only governed tools reach your agent.
**Purpose:** `serve`, `expose` and `dev` make your agent callable by other
agents, on the official `a2a-sdk`. **Advantage:** inbound tasks are
governed with the same correlation, spans and refusals.
**Purpose:** derive a manifest and agent card from your code and register
them in the Agent Fabric registry. **Advantage:** the catalog stays in sync
from CI, not by hand.
## What DDK leaves to the platform
DDK makes the platform's capabilities reachable and typed; it does not
reproduce them. Policy enforcement, semantic caching, provisioning, agent
scanners, kill switch, trusted agent identity, approval UIs and evaluation all
stay with Agent Fabric and Omni Gateway. Where the platform exposes a signal,
such as a semantic-cache hit or the routing decision, DDK surfaces it to your
code. See the [Roadmap](https://docs.donkey-kit.dev/roadmap.md#what-ddk-will-not-build) for the full list.
---
Source: https://docs.donkey-kit.dev/frameworks.md
# Model access
Governed model access from eight agent frameworks. Each adapter returns the
framework's **own native object**, pointed at your Omni Gateway LLM proxy with
consumer auth and attribution headers already set. Nothing wraps the object you
get back, and every page shows the plain-framework code you can switch to at
any time.
**Using TypeScript or another language?** The SDK is Python. From TypeScript
or any other language, call the proxy's OpenAI-compatible HTTP API directly —
every framework page has a **TypeScript** tab showing the official `openai`
npm client (or `@anthropic-ai/sdk` for Anthropic) pointed at the same proxy.
## Supported frameworks
LangGraph is the **deep** adapter: it runs the full conformance suite in CI,
including graph-level scenarios against a compiled `StateGraph`. The other
seven are **supported at `connection_kwargs()`**: the governed connection
settings are tested, and each exposes factory methods that return the native
object.
Each card shows what has been proven about that adapter, in the terms the
[verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md) uses:
| Status | Means |
|---|---|
| Conformance-tested | Runs the conformance suite against the [local simulator](https://docs.donkey-kit.dev/simulator.md) in CI. |
| Live-verified | Has made a real round-trip through a governed proxy. |
| Signature-confirmed | The factory builds the native object against the installed framework, checked offline by `python scripts/verify_frameworks.py`. No live round-trip yet. |
The proxy data plane every adapter calls (base URL, credential headers,
streaming, rejection shapes) is live-verified.
`chat_model()` → `langchain_openai.ChatOpenAI`
`model()` → `google.adk … LiteLlm` (signature-confirmed); `gemini()` → `google.adk.models.Gemini` (live-verified)
`model()` → `strands … OpenAIModel`
`chat_client()` → Agent Framework chat client
`model()` → `agents.OpenAIChatCompletionsModel`
`client()` → `anthropic.AsyncAnthropic`
`llm()` → `crewai.BaseLLM` (`OpenAICompletion`)
`llm()` → `OpenAILike` (`is_chat_model=True`)
`donkey.llm.client()` → `openai.AsyncOpenAI` (or `OpenAI` with `sync=True`)
## The shape is the same everywhere
```bash
pip install "donkey-kit[]"
export DONKEY_LLM_PROXY_URL=… DONKEY_LLM_PROXY_CLIENT_ID=… DONKEY_LLM_PROXY_CLIENT_SECRET=…
```
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
model = donkey..("gpt-4o") # native object at the proxy
```
Each framework page shows the factory name, the native class you get back, the
three ways to construct it, and **the manual equivalent** — the plain framework
constructor call the factory makes for you.
### Printing `connection_kwargs()` hides the secrets
`connection_kwargs()` returns a `dict` that prints `'***'` in place of
`api_key`, the `client_secret` header, `Authorization` and other credential
keys, including inside nested header mappings. The framework still receives the
real values, and `json.dumps`, `dict(...)`, `{**kwargs}` or `kwargs.items()`
still expose the top-level `api_key`. Some framework objects built from the
kwargs (LangGraph's `ChatOpenAI`, LlamaIndex's `OpenAILike`, CrewAI's
`OpenAICompletion`) print credentials themselves. See
[What printed output hides](https://docs.donkey-kit.dev/reference/configuration.md#what-printed-output-hides)
for exactly what is and isn't masked.
Where the framework's dependencies are installed, `connection_kwargs()` also
carries the SDK's HTTP client in the form that framework takes: `http_client`
and `http_async_client` (LangGraph), `http_client` and `async_http_client`
(LlamaIndex), `async_client` (MS Agent Framework), `client` (ADK's `model()`),
or an `interceptor` (CrewAI). Pass them through with the rest of the kwargs.
Each HTTP client is a non-owning view of the SDK's shared client: it sends
through the shared client, and closing it (as Strands does after every call, or
`async with` on an OpenAI client) leaves the shared client open. Only
`donkey.aclose()` / `donkey.close()` end the connection pool. `donkey.http_client()`
returns the same view if you build a framework client by hand.
```python
kwargs = donkey.llamaindex.connection_kwargs()
print(kwargs["default_headers"]) # {'client_id': 'my-client-id', 'client_secret': '***'}
OpenAILike(model="gpt-4o", **kwargs) # receives the real secret
```
## Match the adapter to your proxy's wire format
Every adapter on this page except Anthropic and ADK's `gemini()` — and the raw
`donkey.llm.client()` — speaks the **OpenAI wire format**. The format your proxy accepts is the
**Format** (OpenAI / Anthropic / Gemini) chosen when the proxy was provisioned.
It is a property of the proxy, not an SDK setting, so there is no config field
for it: pick the adapter that matches your proxy.
| Proxy ingress **Format** | Use |
|---|---|
| **OpenAI** | `donkey.llm.client()` or any framework adapter. Default DDK proxies are `Format=OpenAI`. |
| **Anthropic** | `donkey.anthropic.client()` (native `AsyncAnthropic`). The proxy serves the native Messages route at `POST //v1/messages`; OpenAI-shape `/chat/completions` returns 404. |
| **Gemini** | `donkey.adk.gemini("gemini-2.5-flash")` (ADK's native `Gemini` model — see [Native Gemini](https://docs.donkey-kit.dev/frameworks/adk.md#native-gemini)). The proxy serves `POST //models/:generateContent` and `:streamGenerateContent`. There is no standalone `google-genai` adapter; you can also reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). |
**Ingress Format is not the same as the upstream provider.** The ingress
Format is the wire protocol *your request* speaks to the proxy. The upstream
provider is the model the proxy routes *to* after accepting it. A
model-based-routing proxy with OpenAI ingress already fans out to OpenAI,
Gemini, Azure OpenAI, Bedrock Anthropic, and NVIDIA upstreams, selected by the
`model` value in your request body.
So to use Gemini or Claude models you don't need a Gemini- or Anthropic-format
proxy: send an OpenAI-format request naming that model to an OpenAI-format
proxy.
### Decision models: TypeSafe Jev Roadmap
[TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
is a *System One* model: instead of generating text it answers typed questions
(yes/no probability, a choice among options, or a score on a scale) with
calibrated confidence. Its API is **not OpenAI-compatible**, so none of the
adapters above, and no OpenAI-format proxy, can call it. Planned support
returns TypeSafe's own client pointed at Jev behind Omni Gateway, with the same
auth, correlation, typed refusals, spans, budget and simulator support as LLM
calls.
## Injection depth differs by framework
How much of the SDK's HTTP layer reaches the request depends on what each
framework's constructor accepts. With **header injection**, the proxy auth and
attribution headers are sent. With **transport injection**, the SDK's shared
HTTP client is also used, which adds per-run correlation IDs, retries, spans,
`donkey.last_call`, the `jwt`-mode JWT, and
[credentials only to checked endpoints](https://docs.donkey-kit.dev/reference/configuration.md#credentials-go-only-to-checked-endpoints).
| Framework | Header injection | Transport injection | Notes |
|---|---|---|---|
| LangGraph | ✅ | ✅ | `default_headers` plus the SDK's async client (`ainvoke`) and blocking client (`invoke`). |
| Strands | ✅ | ✅ | Via `client_args`. |
| OpenAI Agents SDK | ✅ | ✅ | The adapter builds the `AsyncOpenAI` client itself. |
| Anthropic SDK | ✅ | ✅ | Returns a bare `client()`, not a model-bound object. On `anthropic` 1.0 and later, transport injection goes through a bridged `httpx2` client — see the [Anthropic page](https://docs.donkey-kit.dev/frameworks/anthropic.md). |
| LlamaIndex | ✅ | ✅ | Via `http_client` (sync) and `async_http_client`. `is_chat_model=True` is forced. |
| MS Agent Framework | ✅ | ✅ | Via an `async_client` built on the SDK's client. |
| Google ADK — `model()` | ✅ (`extra_headers`) | ✅ | LiteLLM gets a pre-built OpenAI `client` that sends through the SDK's client. |
| Google ADK — `gemini()` | ✅ | ✅ | Via `HttpOptions.httpx_async_client`, on a `Format=Gemini` proxy. |
| CrewAI | ✅ (`extra_headers`) | ❌ | CrewAI's native OpenAI provider builds its own HTTP client: correlation is per client and `donkey.last_call` is not populated. An `interceptor` keeps credentials to checked endpoints. |
## Retries happen once, in the SDK
The SDK's transport retries `502`, `503` and `504` with backoff, up to
`max_retries` times, and never retries a `4xx`. On the proxy a `429` is a
token-budget refusal (`TokenBudgetExceeded`), so sending it again would only
spend more of a budget that is already gone. Every adapter therefore turns off
the provider SDK's own retries (`max_retries=0`). The transport also marks every
final `4xx` with `x-should-retry: false`, which the `openai` and `anthropic`
SDKs honour, so a client you build yourself from `connection_kwargs()` with its
own retry setting doesn't re-send a refusal either.
Some frameworks retry above the provider SDK, where the SDK can't reach:
| Framework | A budget `429` is sent | A persistent `503` is sent | What to do |
|---|---|---|---|
| LangGraph, OpenAI Agents SDK, Anthropic SDK, LlamaIndex, MS Agent Framework, Google ADK | once | `max_retries + 1` times | Nothing. |
| Strands | once from the model; up to 6 times from a default `Agent` | `max_retries + 1` times | Build the agent with `Agent(retry_strategy=None)`. |
| CrewAI | 3 times | once | No setting turns it off. See the [CrewAI page](https://docs.donkey-kit.dev/frameworks/crewai.md#notes). |
Transport injection also decides whether `jwt` mode works: the rotating JWT is
attached only by the SDK's shared async client. CrewAI can't carry it, so it
raises `ConfigError` in `jwt` mode. Sync calls such as LangGraph's `invoke()`
also raise `ConfigError` instead of sending. See the
[`jwt` mode note](https://docs.donkey-kit.dev/reference/configuration.md#jwt--model-wallet-auth-mode).
[`bearer` mode](https://docs.donkey-kit.dev/reference/configuration.md#bearer-token-auth-mode) has the same
reach; there, CrewAI raises `ConfigError` instead of sending no token.
A URL override passed to a factory (`base_url`, `api_base`, `openai_api_base`,
or Strands' `client_args["base_url"]`) must pass the same
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https) as the
configured proxy URL; it then receives the configured credentials.
See the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md) for how each constructor
signature the adapters depend on is checked.
---
Source: https://docs.donkey-kit.dev/frameworks/langgraph.md
# LangGraph
LangGraph (and LangChain more broadly) gets a governed `ChatOpenAI` pointed at
your Agent Fabric LLM proxy. LangGraph is the **deep adapter**: every proxy
header and the SDK's shared async transport reach the native client, and the
adapter runs the full conformance suite in CI.
**What you get**
- A native `langchain_openai.ChatOpenAI` — nothing LangGraph-specific wraps it.
- Per-run correlation IDs that reach every graph node.
- Typed gateway refusals (`PIIDetected`, `TokenBudgetExceeded`, …) inside nodes.
- A conformance suite you can run against your own graph.
## Install
```bash
pip install "donkey-kit[langgraph]"
```
## Quickstart
```python
from donkey_kit.integrations.langgraph import chat_model
llm = chat_model("gpt-4o")
```
`llm` is a real `langchain_openai.ChatOpenAI` instance. Drop it straight into
your graph nodes or chains.
Call the proxy's OpenAI-compatible API with the official `openai` npm client.
The same base URL and `client_id`/`client_secret` headers also work with
**LangChain.js** (`ChatOpenAI`, via `configuration.baseURL` +
`defaultHeaders`).
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance** (reuses one HTTP client and lifecycle
across every adapter you use in a run):
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.langgraph.chat_model("gpt-4o")
```
The adapter is also callable: `donkey.langgraph("gpt-4o")` is the same as
`donkey.langgraph.chat_model("gpt-4o")`.
**2. Module-level factory** (shortest — uses a cached, env-configured default
`Donkey`):
```python
from donkey_kit.integrations.langgraph import chat_model
llm = chat_model("gpt-4o")
```
**3. Governed kwargs, native constructor** (you call `ChatOpenAI` yourself):
```python
from donkey_kit import Donkey
from langchain_openai import ChatOpenAI
async with Donkey.from_env() as donkey:
llm = ChatOpenAI(model="gpt-4o", **donkey.langgraph.connection_kwargs())
```
## Manual equivalent
The factories make this native constructor call for you:
```python
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="gpt-4o",
base_url=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
default_headers=..., # client_id / client_secret header pair, not bearer
http_async_client=..., # the SDK's shared httpx async client (ainvoke)
http_client=..., # the SDK's shared blocking client (invoke)
max_retries=0, # the SDK retries in its own transport layer
use_responses_api=True, # /responses, the route the raw client and the simulator use
)
```
`connection_kwargs()` returns exactly these keys, so you can drop the factory
and construct `ChatOpenAI` by hand at any time. An OpenAI-format proxy serves
both the Responses API (`/responses`) and `/chat/completions`. The adapter uses
`/responses` because the raw `donkey.llm` client does, and because the local
simulator, which the conformance suite runs against, serves only that route.
Pass `use_responses_api=False` to `chat_model(...)` to call
`/chat/completions` instead. A `base_url` / `openai_api_base` passed to
`chat_model(...)` must pass the
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https).
## Graph-level features
### Correlation IDs reach every node
Bind a run ID once with `donkey.run(id=…)` and every node sees it via
`current_correlation_id()`, with nothing threaded through graph state.
LangGraph runs nodes on `asyncio` tasks that copy the current context, so the
ID propagates on its own:
```python
from donkey_kit.core.telemetry import current_correlation_id
async def prepare(state):
logger.info("handling", extra={"correlation_id": current_correlation_id()})
return {}
async with donkey.run(id=ticket.id):
await graph.ainvoke({"messages": [("user", ticket.text)]})
```
### Typed refusals inside a node
On a proxy refusal, LangChain raises its own wrapped
`OpenAIPermissionDeniedError`, not the SDK's typed exception. Wrap the model
call in `typed_refusals()` and a gateway refusal comes back through the
[error taxonomy](https://docs.donkey-kit.dev/errors.md) instead:
```python
from donkey_kit.integrations.langgraph import typed_refusals
async def call_model(state):
with donkey.langgraph.typed_refusals(): # or: with typed_refusals():
reply = await model.ainvoke(state["messages"])
return {"messages": [reply]}
```
A PII block now propagates out of `graph.ainvoke(...)` as `PIIDetected`, a
budget block as `TokenBudgetExceeded`, and so on — each carrying the
correlation and call IDs the client sent. Transport-level errors with no HTTP
response (connection failures, timeouts) pass through unchanged.
The typed error is raised without a chained cause, because LangChain's error
message repeats the gateway's rejection text, which for a PII block includes
the flagged values, and a traceback prints every chained exception. The
original LangChain error is on `err.framework_error`; `err.__cause__` is
`None`. No frame in the traceback holds it as a local variable, so reporters
that print frame locals (Sentry, `pytest -l`) don't show it.
### `interrupt()` composes with typed refusals
A human-in-the-loop `interrupt()` and a typed refusal don't interfere: the
graph pauses cleanly at the interrupt, and on resume a refusal in a downstream
model node still surfaces as its typed exception.
On Python 3.10, `interrupt()` only works when the graph runs with
`graph.invoke()`. Under `graph.ainvoke()` it raises `RuntimeError: Called
get_config outside of a runnable context`, whether the node is async or sync.
LangGraph reads the run's config from a context variable that asyncio tasks
can't carry before Python 3.11, and `interrupt()` takes no config argument to
pass it in. Use Python 3.11 or later for async graphs that interrupt.
### Run the conformance suite against your own graph
The suite that tests this adapter is also a pytest plugin you can point at
your own agent:
```bash
pytest --donkey-conformance --agent=my_app:build
```
`build` returns an object with an awaitable `run(text)`. The suite checks that
it doesn't retry a budget refusal, surfaces `PIIDetected` typed, carries the
correlation ID into its logs, and tolerates a response with no budget headers.
The [`examples/langgraph`](https://github.com/Donkey-Development-Kit/donkey-development-kit/tree/main/python/examples/langgraph)
factory has exactly this shape. See [Testing](https://docs.donkey-kit.dev/testing.md).
### Which provider served this?
A call routed by the gateway to Gemini, Anthropic, or Bedrock still comes back
with `response_metadata["model_provider"] == "openai"` on the LangChain
message. That is not a routing bug — LangChain stamps `model_provider` from
the **client class** (`ChatOpenAI`, OpenAI-compatible), not from whichever
upstream the gateway actually picked. Don't use it for routing attribution;
use one of the following instead.
**Per context — `donkey.last_call`.** Right after `ainvoke`, read the record
the gateway reported for the most recent call:
```python
reply = await model.ainvoke(state["messages"])
r = donkey.last_call
r.served_provider # e.g. "gemini"
r.served_model
r.routing_type
r.fallback
r.substituted
```
`last_call` is contextvar-scoped: it is the last call **in the current task**,
so it does not survive into graph state and does not leak back to a parent
that gathered parallel branches (each branch runs on its own `asyncio` task
with its own copy). Read it immediately after the call it describes.
LangChain sends an `ainvoke` request from a task of its own, so a model built by
`donkey.langgraph(...)` carries a callback that brings the record back to the
task that called `ainvoke`. A `ChatOpenAI` you build from `connection_kwargs()`
has no such callback: after its `ainvoke` (but not `invoke` or `astream`),
`last_call` stays `UNOBSERVED`. A batch (`abatch`, or `agenerate` with several
inputs) runs its requests side by side, so it leaves no record in the caller
either. See the
[full field reference](https://docs.donkey-kit.dev/reference/last-call.md#routing--fallback) for every
field.
**Per message — opt-in headers.** For attribution that needs to travel with
the message itself (into checkpoints, across the boundary where `last_call`
would go stale), pass `include_response_headers=True` — this is
`langchain_openai.ChatOpenAI`'s own field, forwarded through
`chat_model(**kwargs)` like any other native kwarg:
```python
model = donkey.langgraph.chat_model("gpt-4o", include_response_headers=True)
msg = await model.ainvoke(...)
headers = msg.response_metadata["headers"]
headers.get("x-llm-proxy-llm-provider") # e.g. "gemini"
headers.get("x-llm-proxy-llm-model")
headers.get("x-llm-proxy-routing-fallback")
headers.get("x-llm-proxy-routing-type")
```
Because this is langchain-openai's own behavior, not the SDK's, its exact
shape tracks that package's version:
- It copies **all** response headers into `response_metadata["headers"]` on
every message — and therefore into every checkpoint that stores the
message.
- On a streamed response, the headers land on the **first chunk only**.
- Whether headers are captured on the chat-completions + `response_format`
path has varied by langchain-openai version — verify against the version
you have installed rather than assuming either way.
**Traces.** The OTel GenAI span for the call already carries the served
provider and model under `gen_ai.system` / `gen_ai.response.model` (see
[Telemetry](https://docs.donkey-kit.dev/telemetry.md)) — but a LangSmith-style tracer that reads LangChain's
own fields instead of the span will still show `openai`.
## Notes
- `base_url`, `api_key`, `default_headers`, `http_async_client` and
`http_client` are all forwarded, so proxy auth headers and the SDK's
transport (retries, correlation IDs) reach every request, from `ainvoke()`
and `invoke()` alike, so `donkey.simulate()` and `donkey.last_call` cover
both. In `jwt` mode only `ainvoke()` / `astream()` carry the JWT: `invoke()`
and `stream()` raise `ConfigError` before sending anything.
- Printing the model shows `client_secret`: `ChatOpenAI`'s own `repr()` /
`str()` include `default_headers`. Don't print or log it; see
[What printed output hides](https://docs.donkey-kit.dev/reference/configuration.md#what-printed-output-hides).
- `max_retries=0` is intentional: retries live in the SDK's transport layer,
so the SDK and the OpenAI client don't both retry.
- The proxy is OpenAI-compatible but not the full OpenAI API: the base URL has
no `/v1` prefix, there is no `/models` endpoint, and auth is a
`client_id`/`client_secret` header pair rather than a bearer token.
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface as typed
exceptions, and [Which provider served this?](#which-provider-served-this)
for reading the served provider/model instead of LangChain's own
`model_provider`.
---
Source: https://docs.donkey-kit.dev/frameworks/adk.md
# Google ADK
Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy in one
of two ways, depending on the proxy's ingress **Format**:
- `model()` — ADK's `LiteLlm` model wrapper, for a `Format=OpenAI` proxy (the
default). The adapter translates the governed connection into LiteLLM's own
model-string and kwarg conventions for you.
- `gemini()` — ADK's native `Gemini` model, for a `Format=Gemini` proxy. The
SDK's shared HTTP client is injected, so every governed feature works: per-run
correlation, spans, usage and `donkey.last_call`. See
[Native Gemini](#native-gemini).
**What you get**
- A native `google.adk.models.lite_llm.LiteLlm` or `google.adk.models.Gemini`,
with the proxy auth and attribution headers set.
- For `model()`: the `openai/` model prefix and LiteLLM kwarg names handled
automatically. Supported at `connection_kwargs()`. LiteLLM gets a pre-built
OpenAI `client` that sends through the SDK's shared HTTP client, so per-run
correlation, retries, spans, `donkey.last_call` and the `jwt`-mode JWT apply
(see [Notes](#notes)).
- For `gemini()`: the shared client injected through
`HttpOptions.httpx_async_client`, round-trip verified live against a
`Format=Gemini` proxy.
## Install
```bash
pip install "donkey-kit[adk]"
```
## Quickstart
```python
from donkey_kit.integrations.adk import model
llm = model("gpt-4o")
```
`llm` is a real `google.adk.models.lite_llm.LiteLlm` instance. The model string
is prefixed with `openai/` before it reaches LiteLLM (`openai/gpt-4o`), which
is the prefix LiteLLM's OpenAI-compatible route expects — you don't add it
yourself.
Call the proxy's OpenAI-compatible API with the official `openai` npm client.
The same base URL and `client_id`/`client_secret` headers also work with
**ADK for TypeScript** (`@google/adk`).
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.adk.model("gpt-4o")
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.adk import model
llm = model("gpt-4o")
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from google.adk.models.lite_llm import LiteLlm
async with Donkey.from_env() as donkey:
llm = LiteLlm(model="openai/gpt-4o", **donkey.adk.connection_kwargs())
```
## Manual equivalent
```python
from google.adk.models.lite_llm import LiteLlm
llm = LiteLlm(
model="openai/gpt-4o",
api_base=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
extra_headers=..., # client_id / client_secret header pair
client=..., # an AsyncOpenAI on the SDK's shared client
max_retries=0, # the SDK retries in its own transport layer
)
```
LiteLLM uses `api_base` and `extra_headers`, not `base_url` /
`default_headers` — `connection_kwargs()` already translates for you.
`client` is present when the `openai` package is installed; LiteLLM's OpenAI
route uses it in place of the client it would build. `max_retries=0` has to
be passed to LiteLLM itself: LiteLLM sets the client's retry count on every
call, and its default is 2. An `api_base` /
`base_url` passed to `model()` must pass the
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https), and the
adapter builds `client` on that URL.
## Native Gemini
A proxy provisioned with **Format = Gemini** exposes the native Gemini API at
`POST //models/:generateContent` (and
`:streamGenerateContent`). ADK's own `Gemini` model speaks that API, so
`gemini()` returns a native `google.adk.models.Gemini` bound to the proxy, with
the SDK's shared HTTP client injected.
The three forms mirror `model()`. Pass the bare Gemini model id — there is no
prefix, because the URL path carries the model:
```python
from donkey_kit import Donkey
from google.adk.models import Gemini
async with Donkey.from_env() as donkey:
# 1. Off a shared Donkey instance
llm = donkey.adk.gemini("gemini-2.5-flash")
# 3. Governed kwargs, native constructor
llm = Gemini(model="gemini-2.5-flash", **donkey.adk.gemini_connection_kwargs())
```
```python
# 2. Module-level factory
from donkey_kit.integrations.adk import gemini
llm = gemini("gemini-2.5-flash")
```
**Pointing at the Gemini proxy.** `DONKEY_LLM_PROXY_URL` usually names a
`Format=OpenAI` proxy. If your Gemini proxy is a different one, pass its URL
as `base_url`; the `client_id`/`client_secret` pair must be contracted on that
proxy, and the URL must pass the
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https):
```python
llm = donkey.adk.gemini("gemini-2.5-flash", base_url="https://…/ddk-gemini-inbound/")
```
Any other keyword is passed to ADK's `Gemini` and overrides the governed
default — for example `retry_options`, or your own `client_kwargs`.
**Manual equivalent.** This is what `gemini_connection_kwargs()` returns:
```python
from google.adk.models import Gemini
llm = Gemini(
model="gemini-2.5-flash",
base_url=..., # the Format=Gemini proxy URL
client_kwargs={
"api_key": ..., # a placeholder; google-genai requires one
"http_options": {
"base_url": ..., # same URL
"api_version": "", # the proxy path has no /v1beta segment
"headers": ..., # client_id / client_secret header pair
"timeout": ..., # milliseconds
"httpx_async_client": ..., # the SDK's shared client
},
},
)
```
`client_kwargs` replaces ADK's default HTTP options wholesale, so all five
`http_options` keys are passed together. `google-genai` requires an API key and
always sends it as `x-goog-api-key`; the gateway authenticates on the
`client_id`/`client_secret` pair and ignores it.
**`donkey.last_call`.** Usage (`input_tokens`, `output_tokens`,
`total_tokens`, `cached_tokens`, `reasoning_tokens`) is read from Gemini's
`usageMetadata`, and `requested_model` from the URL path. Gemini's
`total_tokens` includes the thinking tokens it also reports as
`reasoning_tokens`. A `Format=Gemini` proxy is a passthrough, so it sends no
routing headers: `served_provider`, `served_model` and `routing_type` stay
`None`, and `substituted` is `False`. `request_id` and `api_instance_id` are
populated.
`last_call` is contextvar-scoped, and ADK's `Runner` makes the model call in a
task of its own, so the caller of `runner.run_async(...)` reads `UNOBSERVED`.
Read it in an `after_model_callback`, which runs in the same task as the call:
```python
from google.adk.agents import LlmAgent
def record_usage(callback_context, llm_response):
r = donkey.last_call
print(r.requested_model, r.input_tokens, r.output_tokens)
return None # keep the model's response
agent = LlmAgent(
name="assistant",
model=donkey.adk.gemini("gemini-2.5-flash"),
after_model_callback=record_usage,
)
```
When streaming, the callback runs once per partial response; usage lands on the
last one.
**Errors.** `google-genai` raises its own `google.genai.errors.APIError`
(`ClientError` for a 4xx) and the SDK does not wrap it. The error's `.response`
is the proxy's HTTP response, so `classify()` gives you the typed Donkey error:
```python
from donkey_kit.core.errors import classify
from google.genai import errors as genai_errors
try:
async for event in runner.run_async(...):
...
except genai_errors.APIError as exc:
err = classify(exc.response) # e.g. AuthError on 401, UpstreamRequestError on 404
```
**Reusing a model across `asyncio.run(...)` calls works.** The shared HTTP
client keeps one connection pool per event loop, so a sync app that wraps each
run in `asyncio.run(...)` can build the `Donkey` (or call the module-level
`gemini()`) once and reuse it.
## Notes
- **Both factories send through the SDK's shared HTTP client**, so the
correlation ID bound by `donkey.run()` reaches every request.
- **`donkey.last_call` is set in the context that made the call.** ADK's
`Runner` calls the model in a task of its own, so read it in an
`after_model_callback` (see [Native Gemini](#native-gemini)). Until
`gemini()` has been called on a `Donkey`, a cold read on a `Donkey` that
resolved only ADK reports `UNAVAILABLE` rather than `UNOBSERVED`, because
`model()` is still listed as not observing calls. Aligning that, and the
matching conformance exemptions, is tracked in
[#740](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/740).
- **Refusals on the `model()` path aren't typed.** LiteLLM raises its own
error without the response headers, so `classify()` has nothing to read; see
the [ADK examples](https://docs.donkey-kit.dev/examples/adk.md).
- **`gemini()` needs `google-adk>=2.4`**, the first release whose `Gemini`
accepts `client_kwargs`; the `adk` extra declares that floor. On an older
ADK, `Gemini` drops the governed client without an error and talks to Google
directly, so `gemini()` raises `NotImplementedError` instead of returning a
model that bypasses the gateway.
- `google-adk` requires `litellm>=1.84` as a floor, not a ceiling — pin your
own upper bound if you need one.
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface through
ADK's `LiteLlm` model, and [Native Gemini](#native-gemini) for `gemini()`.
---
Source: https://docs.donkey-kit.dev/frameworks/strands.md
# Strands Agents
Strands Agents gets a governed `OpenAIModel` pointed at the Agent Fabric LLM
proxy. The connection details travel in Strands' `client_args`, which Strands
passes straight through to its underlying OpenAI client.
**What you get**
- A native `strands.models.openai.OpenAIModel`.
- Full header **and** transport injection through `client_args` — per-run
correlation IDs and `donkey.last_call` work, as with LangGraph.
- Supported at `connection_kwargs()`.
## Install
```bash
pip install "donkey-kit[strands]"
```
## Quickstart
```python
from donkey_kit.integrations.strands import model
llm = model("gpt-4o")
```
`llm` is a real `strands.models.openai.OpenAIModel` instance — pass it to your
`Agent` as you would any other Strands model.
Call the proxy's OpenAI-compatible API with the official `openai` npm client.
The same base URL and `client_id`/`client_secret` headers also work with the
**Strands TypeScript SDK** (`@strands-agents/sdk`).
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.strands.model("gpt-4o")
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.strands import model
llm = model("gpt-4o")
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from strands.models.openai import OpenAIModel
async with Donkey.from_env() as donkey:
llm = OpenAIModel(model_id="gpt-4o", **donkey.strands.connection_kwargs())
```
## Manual equivalent
```python
from strands.models.openai import OpenAIModel
llm = OpenAIModel(
model_id="gpt-4o",
client_args={
"base_url": ..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
"api_key": ...,
"default_headers": ..., # client_id / client_secret header pair
"http_client": ..., # a non-owning view of the SDK's shared client
"max_retries": 0, # the SDK retries in its own transport layer
},
stream=False, # see Notes
)
```
Everything the SDK injects lives inside the single `client_args` dict that
Strands forwards to its internal OpenAI client.
## Notes
- **Build the agent with `retry_strategy=None`.** A Strands `Agent` retries a
throttled model call by default (up to 6 attempts), and Strands treats every
`429` as throttling. On the proxy a `429` is a budget refusal, so turn the
agent's retry off and let the SDK's transport handle the transient `5xx`:
```python
from strands import Agent
agent = Agent(model=donkey.strands.model("gpt-4o"), retry_strategy=None)
```
The model itself has `max_retries=0`, so the OpenAI client under it doesn't
retry either.
- Strands forwards `client_args` verbatim to the underlying OpenAI client, so
both header injection (`default_headers`) and transport injection
(`http_client`) are available.
- **The client stays open.** Strands opens and closes an OpenAI client for
every request (`async with AsyncOpenAI(**client_args)`). The `http_client` it
gets is a view whose close is a no-op, so the SDK's shared client survives
every call. Only `donkey.aclose()` ends the connection pool.
- **Streaming is off by default.** The governed model sets `stream=False`. A
proxy routing to a Gemini upstream answers a streamed request with one whole
`chat.completion` and no chunk deltas, and Strands fails on it. Pass
`donkey.strands.model("gpt-4o", stream=True)` on routes that stream.
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface as typed
exceptions, and the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md) for
the current status of every constructor signature this adapter depends on.
---
Source: https://docs.donkey-kit.dev/frameworks/agent-framework.md
# Microsoft Agent Framework
Microsoft Agent Framework gets a governed chat client pointed at the Omni
Gateway LLM proxy, plus policy middleware that ends a run on a governance
rejection with the SDK's typed refusal.
**What you get**
- A native `agent_framework.openai.OpenAIChatClient`, which calls the
Responses API (`POST /responses`), checked against agent-framework 1.19.0.
Pass `api="chat_completions"` for the Chat Completions client (see
[Which API](#which-api)).
- `policy_middleware()`, which ends a run on a proxy refusal with a typed
`PolicyViolation` such as `PIIDetected`.
- Supported at `connection_kwargs()`. The client gets an `async_client` that
sends through the SDK's shared HTTP client, so per-run correlation, retries,
spans, `donkey.last_call` and the `jwt`-mode JWT apply (see [Notes](#notes)).
## Install
```bash
pip install "donkey-kit[agent_framework]"
```
## Quickstart
```python
from donkey_kit.integrations.agent_framework import chat_client
llm = chat_client("gpt-4o")
```
`llm` is a real `agent_framework.openai.OpenAIChatClient` instance.
Microsoft Agent Framework ships for .NET, Python, and Go, not TypeScript. From
TypeScript, call the proxy's OpenAI-compatible API directly with the official
`openai` npm client:
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.agent_framework.chat_client("gpt-4o")
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.agent_framework import chat_client
llm = chat_client("gpt-4o")
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from agent_framework.openai import OpenAIChatClient
async with Donkey.from_env() as donkey:
llm = OpenAIChatClient(
model="gpt-4o",
**donkey.agent_framework.connection_kwargs(),
)
```
## Manual equivalent
```python
from agent_framework.openai import OpenAIChatClient
llm = OpenAIChatClient(
model=..., # `model`, not `model_id`
base_url=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
default_headers=..., # client_id / client_secret header pair
async_client=..., # an AsyncOpenAI on the SDK's shared client
)
```
`async_client` is present when the `openai` package is installed. The
constructor uses it as given, so a `base_url` passed to `chat_client()` gets a
client built on that URL; the URL must pass the
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https).
The same kwargs work with `OpenAIChatCompletionClient`, the Chat Completions
client.
## Which API
agent-framework-openai has two OpenAI-compatible chat clients, and
`chat_client()` picks one with `api=`:
| `api=` | Client | Request |
| --- | --- | --- |
| `"responses"` (default) | `OpenAIChatClient` | `POST /responses` |
| `"chat_completions"` | `OpenAIChatCompletionClient` | `POST /chat/completions` |
The default is the Responses API because it is the proxy route the SDK has
verified live, the same one `donkey.llm` and the LangGraph adapter use. It is
verified on an OpenAI upstream only. A proxy can route a model to any
upstream, and not every upstream serves `/responses`: a route to Azure OpenAI
answers it with `404 Resource not found`, which reaches you as a
`ChatClientException`. For such a route, ask for the Chat Completions client:
```python
llm = chat_client("azure/gpt-4.1-mini", api="chat_completions")
```
Both clients work with `policy_middleware()`. The local simulator serves
`/responses` only, so keep the default when you run against it.
## Policy middleware
`donkey.agent_framework.policy_middleware()` returns a chat middleware for
`Agent(..., middleware=[...])`:
```python
from agent_framework import Agent
from donkey_kit import Donkey, PIIDetected
async with Donkey.from_env() as donkey:
agent = Agent(
client=donkey.agent_framework.chat_client("gpt-4o"),
middleware=[donkey.agent_framework.policy_middleware()],
)
try:
await agent.run("...")
except PIIDetected as err:
print(err.correlation_id, err.entities)
```
Without it, a proxy refusal reaches you as Agent Framework's generic
`ChatClientException`. With it, the refusal goes through
`donkey_kit.core.errors.classify()` and the run ends with the typed error,
for example `PIIDetected` or `TokenBudgetExceeded`. The error carries the
correlation and call ids that were sent, and keeps the original exception on
`.framework_error`. This also works for streaming runs
(`agent.run(..., stream=True)`): the typed error is raised while you iterate
the stream. The chat client sends with retries off, so a refused request is
sent once.
Errors that have no proxy response behind them, such as a connection failure,
pass through unchanged. The middleware is marked with Agent Framework's
`@chat_middleware` decorator. That is confirmed offline against
agent-framework 1.19.0, with no live round-trip yet. If the decorator is
missing from your installed version, `policy_middleware()` raises a
`NotImplementedError` naming it.
## Notes
- **Constructor signature.** `OpenAIChatClient` and
`OpenAIChatCompletionClient` both take `model`, `base_url`, `api_key`, and
`default_headers` (agent-framework 1.19.0; `model_id` is not accepted). If the import fails or an upstream release renames a kwarg,
`chat_client()` raises a `NotImplementedError` naming the class path or
signature to check, rather than a raw `ImportError` or `TypeError`.
- **`donkey.last_call` is set in the context that made the call.** A cold read
(no call yet in this context) on a `Donkey` that resolved only adapters like
this one still reports `status == LastCallStatus.UNAVAILABLE` rather than
`UNOBSERVED`, because MS Agent Framework is still listed as not observing calls. Aligning
that, and the matching conformance exemptions, is tracked in [#740](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/740).
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for the full `PolicyViolation` hierarchy
that `policy_middleware()` raises.
---
Source: https://docs.donkey-kit.dev/frameworks/openai.md
# OpenAI Agents SDK
The OpenAI Agents SDK (pip package `openai-agents`) gets a governed
`OpenAIChatCompletionsModel` backed by a pre-built `AsyncOpenAI` client. The
adapter builds that client itself, with the SDK's shared HTTP client and proxy
headers, and hands it to the Agents SDK ready-made.
**What you get**
- A native `agents.OpenAIChatCompletionsModel`.
- Full header **and** transport injection — both travel together in one
`AsyncOpenAI` object.
- Supported at `connection_kwargs()`.
## Install
```bash
pip install "donkey-kit[openai-agents]"
```
## Quickstart
```python
from donkey_kit.integrations.openai_agents import model
llm = model("gpt-4o")
```
`llm` is a real `agents.OpenAIChatCompletionsModel` instance — pass it to
`Agent(model=...)` as you would any other Agents SDK model.
Call the proxy's OpenAI-compatible API with the official `openai` npm client.
The same base URL and `client_id`/`client_secret` headers also work with the
**OpenAI Agents SDK for JS** (`@openai/agents`).
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.openai_agents.model("gpt-4o")
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.openai_agents import model
llm = model("gpt-4o")
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from agents import OpenAIChatCompletionsModel
async with Donkey.from_env() as donkey:
llm = OpenAIChatCompletionsModel(
model="gpt-4o",
**donkey.openai_agents.connection_kwargs(),
)
```
## Manual equivalent
```python
from openai import AsyncOpenAI
from agents import OpenAIChatCompletionsModel
async_client = AsyncOpenAI(
base_url=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
default_headers=..., # client_id / client_secret header pair
http_client=..., # the SDK's shared httpx client
)
llm = OpenAIChatCompletionsModel(
model="gpt-4o",
openai_client=async_client,
)
```
`connection_kwargs()` returns exactly one key, `openai_client`, holding this
pre-built `AsyncOpenAI` instance.
## `donkey.last_call`
Every model call goes through the SDK's shared HTTP client, so the SDK records
it. `last_call` is contextvar-scoped, though, and `Runner` sends a run's first
turn from a task of its own (`run_streamed` sends every turn that way). After a
one-turn `Runner.run`, or any `run_streamed`, the caller reads `UNOBSERVED`.
After a longer `Runner.run` it reads the last turn's call. To read every call,
use `RunHooks.on_llm_end`, which runs in the same task as the call:
```python
from agents import Agent, RunHooks, Runner
class LastCallHooks(RunHooks):
async def on_llm_end(self, context, agent, response):
r = donkey.last_call
print(r.request_id, r.input_tokens, r.output_tokens)
agent = Agent(name="assistant", model=donkey.openai_agents.model("gpt-4o"))
await Runner.run(agent, "hello", hooks=LastCallHooks())
```
## Notes
- **A pre-built client is the preferred integration point.** When a framework
accepts a ready-made `AsyncOpenAI` instead of loose kwargs, the shared
transport and every proxy header travel together as one object, with no
risk of a kwarg being dropped. That's why injection is full here even though
the model object never sees `base_url` or `default_headers` directly.
- `openai-agents` is distinct from the plain `openai` package:
`agents.OpenAIChatCompletionsModel` lives in the Agents SDK. Installing
`donkey-kit[openai-agents]` pulls it in for you. For the raw governed client
with no framework, use `donkey.openai()` (from `donkey-kit[llm]`) instead.
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface as typed
exceptions, and the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md) for
the current status of every constructor signature this adapter depends on.
---
Source: https://docs.donkey-kit.dev/frameworks/anthropic.md
# Anthropic SDK
The Anthropic SDK gets a governed `AsyncAnthropic` client pointed at the Omni
Gateway LLM proxy, with the SDK's shared transport and proxy headers passed
straight into the client constructor.
**What you get**
- A native `anthropic.AsyncAnthropic` client.
- Full header **and** transport injection, on `anthropic` 0.x and 1.x.
- Supported at `connection_kwargs()`.
**Requires a `Format=Anthropic` proxy.** The native Anthropic Messages route
(`POST //v1/messages`) is only served by a proxy provisioned with
the Anthropic ingress Format. Default DDK proxies are `Format=OpenAI`: there,
`/v1/messages` returns 404 and Claude is reachable only as an upstream
provider through the OpenAI-compatible adapters. See
[Model access](https://docs.donkey-kit.dev/frameworks.md) for how ingress Format works.
## Install
```bash
pip install "donkey-kit[anthropic]"
```
This installs the newest `anthropic` release. `anthropic` 0.x (from 0.40) works
too, so a project that holds it below 1.0 needs no change.
## Quickstart
```python
from donkey_kit.integrations.anthropic import client
llm = client()
```
`llm` is a real `anthropic.AsyncAnthropic` instance. Unlike the other
adapters, the factory takes no `model` argument — pass the model ID per call,
as the Anthropic SDK expects:
```python
reply = await llm.messages.create(
model="claude-...",
max_tokens=1024,
messages=[{"role": "user", "content": "Say hi in three words."}],
)
```
Use the official **`@anthropic-ai/sdk`** client pointed at a `Format=Anthropic`
proxy, with the same `client_id` / `client_secret` header pair and the model ID
passed per call:
```typescript
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // a Format=Anthropic proxy
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.messages.create({
model: "claude-...",
max_tokens: 1024,
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.anthropic.client()
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.anthropic import client
llm = client()
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from anthropic import AsyncAnthropic
async with Donkey.from_env() as donkey:
llm = AsyncAnthropic(**donkey.anthropic.connection_kwargs())
```
## Manual equivalent
```python
from anthropic import AsyncAnthropic
llm = AsyncAnthropic(
base_url=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
default_headers=..., # client_id / client_secret header pair
http_client=..., # sends through the SDK's shared transport (see below)
max_retries=0, # the SDK retries in its own transport layer
)
```
`connection_kwargs()` returns exactly these keys, so you can drop the factory
and construct `AsyncAnthropic` by hand at any time.
Which `http_client` you get depends on the installed `anthropic`:
| `anthropic` | Built on | `http_client` |
| --- | --- | --- |
| 0.x | `httpx` | A non-owning `httpx.AsyncClient` view that sends through the shared client. |
| 1.0 and later | `httpx2` | An `httpx2.AsyncClient` whose transport sends every request through the shared client. |
`anthropic` 1.0 moved to `httpx2`, Pydantic's continuation of `httpx`, and
rejects any `httpx` client with a `TypeError`. The bridged `httpx2` client keeps
one HTTP stack: the governed headers, retries, the GenAI span, budget tracking
and `donkey.last_call` all run in the shared client either way. If you bring
your own `http_client` instead, use an `httpx2.AsyncClient` on 1.0 and later.
## Notes
- **`client()`, not `model(...)`.** The other adapters return a framework
object already bound to a model ID, because their native constructors accept
`model`. `AsyncAnthropic` is a bare client and the model ID is an argument to
`.messages.create()`, so `donkey.anthropic.client()` takes no model argument.
- **Async only.** The adapter returns `AsyncAnthropic`; there is no governed
sync `anthropic.Anthropic`. `connection_kwargs()` carries the SDK's async
client, so spread it only into `AsyncAnthropic`. A sync `Anthropic` you build
yourself does not go through the SDK's transport.
- **Proxy Format.** MuleSoft Model Proxy offers three ingress Formats (OpenAI /
Gemini / Anthropic), fixed when the proxy is created
([MuleSoft docs](https://docs.mulesoft.com/general/model-proxy)). A
`Format=Anthropic` proxy returns a native Anthropic body from
`/v1/messages` and 404s an OpenAI-shaped `/chat/completions` request. Auth is
the same `client_id` / `client_secret` header pair as every other proxy.
To use the native surface, set `DONKEY_LLM_PROXY_URL` (or `llm_proxy_url`)
to a `Format=Anthropic` proxy.
- **What `donkey.last_call` reads.** On a native Anthropic proxy,
`request_id` is Anthropic's own `request-id` header. `cached_tokens` and
`cache_write_tokens` come from `cache_read_input_tokens` and
`cache_creation_input_tokens`. Streamed calls fill `input_tokens` and the
cache counts from `message_start` and `output_tokens` from `message_delta`.
Anthropic reports no total, so `total_tokens` is `None`. Its `input_tokens`
excludes both cache counts, while OpenAI's includes `cached_tokens`. Add the
cache counts back in before you compare cost across providers.
- **Closing the client.** `AsyncAnthropic.close()`, and leaving
`async with AsyncAnthropic(...)`, closes its `http_client`. On both stacks
that is not the shared client (the bridge on 1.0 and later, the view on 0.x),
so the shared client stays open and later `client()` calls keep working. To
end the connection pool, close the `Donkey` (`async with Donkey.from_env()`
or `await donkey.aclose()`).
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface as typed
exceptions, and the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md) for
the current status of every constructor signature this adapter depends on.
---
Source: https://docs.donkey-kit.dev/frameworks/crewai.md
# CrewAI
CrewAI gets a governed native LLM object pointed at the Agent Fabric LLM
proxy. The adapter translates the governed connection into CrewAI's own
model-string and kwarg conventions for you.
**What you get**
- A native `crewai.BaseLLM` — concretely `OpenAICompletion`, CrewAI's native
OpenAI provider. `crewai.LLM`'s own `__new__` factory routes an
`openai/`-prefixed model with an explicit `base_url` to that provider rather
than returning an `LLM` instance itself. The proxy auth and attribution
headers are set.
- The `openai/` model prefix and CrewAI's kwarg names handled automatically.
- Supported at `connection_kwargs()`. CrewAI's native provider builds its own
HTTP client rather than using the SDK's, so there is no run correlation ID
and `donkey.last_call` is not populated (see [Notes](#notes)).
## Install
```bash
pip install "donkey-kit[crewai]"
```
## Quickstart
```python
from donkey_kit.integrations.crewai import llm
model = llm("gpt-4o")
```
`model` is a real `crewai.BaseLLM` instance (`OpenAICompletion`). The adapter
prefixes the model string with `openai/` (`openai/gpt-4o`) — you don't add it
yourself. That prefix, together with the proxy `base_url`, is what routes
`crewai.LLM`'s factory to its native OpenAI provider; the provider strips the
prefix again, so the proxy receives the bare model id (`gpt-4o`).
CrewAI is Python-only. From TypeScript, call the proxy's OpenAI-compatible API
directly with the official `openai` npm client:
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
model = donkey.crewai.llm("gpt-4o")
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.crewai import llm
model = llm("gpt-4o")
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from crewai import LLM
async with Donkey.from_env() as donkey:
model = LLM(model="openai/gpt-4o", **donkey.crewai.connection_kwargs())
```
## Manual equivalent
```python
from crewai import LLM
model = LLM(
model="openai/gpt-4o",
base_url=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
extra_headers=..., # client_id / client_secret header pair
max_retries=0, # don't let the OpenAI client retry a refusal
)
```
`connection_kwargs()` fills in `base_url`, `api_key`, `extra_headers` (the
`client_id` / `client_secret` header pair), and `max_retries=0` for you. When CrewAI is installed it
also adds an `interceptor`: with one set, CrewAI's provider builds HTTP clients
that don't follow redirects, and the interceptor removes the credential headers
from any request to an origin other than the proxy (see
[Credentials go only to checked endpoints](https://docs.donkey-kit.dev/reference/configuration.md#credentials-go-only-to-checked-endpoints)).
A `base_url` / `api_base` passed to `llm()` must pass the
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https).
## Notes
- **No run correlation ID.** CrewAI sends requests through its native OpenAI
provider's own HTTP client rather than the SDK's shared one, so the
correlation ID bound by `donkey.run()` doesn't reach the request. The auth
and attribution headers are still sent on every request. The conformance
suite checks this as a documented behaviour.
- **`jwt` mode isn't supported.** The JWT is added only by the SDK's shared
client, which CrewAI's provider doesn't use, so `donkey.crewai.llm()` and
`connection_kwargs()` raise `ConfigError` in `jwt` mode. Use client-id auth
with CrewAI; see the
[`jwt` mode note](https://docs.donkey-kit.dev/reference/configuration.md#jwt--model-wallet-auth-mode).
- **A budget refusal is sent 3 times.** CrewAI wraps every LLM call in its own
rate-limit retry (3 attempts, with a 1s then 2s wait) and treats any `429` as
a rate limit. On the proxy a `429` is a budget refusal, so CrewAI re-sends it
twice before raising. CrewAI has no setting to turn this retry off. The
OpenAI client underneath has `max_retries=0`, and transient `5xx` errors are
not retried at all, because the SDK's transport isn't used.
- **Printing the model shows the API key.** `OpenAICompletion`'s own `repr()` /
`str()` include `api_key`. Don't print or log it.
- **`donkey.last_call` is unavailable.** Because the response is handled by
CrewAI's own client, gateway identity, routing, and usage fields can't be
observed. When every adapter resolved on a `Donkey` is like this one,
`donkey.last_call` reports
`status == LastCallStatus.UNAVAILABLE` and `available == False`, and names the
resolved adapters in `surface`.
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface through
CrewAI.
---
Source: https://docs.donkey-kit.dev/frameworks/llamaindex.md
# LlamaIndex
LlamaIndex gets a governed `OpenAILike` LLM pointed at the Agent Fabric LLM
proxy, with the chat-model flag a chat-only gateway requires already set.
**What you get**
- A native `llama_index.llms.openai_like.OpenAILike`.
- `is_chat_model=True` and `is_function_calling_model=True` set for you, and
`max_retries=0` so `OpenAILike` doesn't retry on top of the SDK's transport.
- Supported at `connection_kwargs()`. Sync and async calls send through the
SDK's HTTP clients (`http_client`, `async_http_client`), so per-run
correlation, retries, spans and `donkey.last_call` apply (see
[Notes](#notes)).
## Install
```bash
pip install "donkey-kit[llamaindex]"
```
## Quickstart
```python
from donkey_kit.integrations.llamaindex import llm
model = llm("gpt-4o")
```
`model` is a real `llama_index.llms.openai_like.OpenAILike` instance, ready to
hand to any LlamaIndex query engine, chat engine, or agent.
Call the proxy's OpenAI-compatible API with the official `openai` npm client.
The same base URL and `client_id`/`client_secret` headers also work with
**LlamaIndex.TS**.
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: process.env.DONKEY_LLM_PROXY_URL, // no /v1
apiKey: "unused", // required slot; proxy uses the headers below
defaultHeaders: {
client_id: process.env.DONKEY_LLM_PROXY_CLIENT_ID!,
client_secret: process.env.DONKEY_LLM_PROXY_CLIENT_SECRET!,
},
});
const reply = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hi in three words." }],
});
console.log(reply.choices[0].message.content);
```
## Three ways to construct
**1. Off a shared `Donkey` instance:**
```python
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
model = donkey.llamaindex.llm("gpt-4o")
```
**2. Module-level factory** (shortest):
```python
from donkey_kit.integrations.llamaindex import llm
model = llm("gpt-4o")
```
**3. Governed kwargs, native constructor:**
```python
from donkey_kit import Donkey
from llama_index.llms.openai_like import OpenAILike
async with Donkey.from_env() as donkey:
model = OpenAILike(model="gpt-4o", **donkey.llamaindex.connection_kwargs())
```
## Manual equivalent
```python
from llama_index.llms.openai_like import OpenAILike
model = OpenAILike(
model="gpt-4o",
api_base=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
default_headers=..., # client_id / client_secret header pair
is_chat_model=True, # required — see below
is_function_calling_model=True,
http_client=..., # the SDK's blocking client
async_http_client=..., # the SDK's async client
max_retries=0, # the SDK's transport retries
)
```
LlamaIndex uses `api_base` rather than `base_url`; `connection_kwargs()`
already translates for you. An `api_base` passed to `llm()` must pass the
[`https://` rule](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https).
## Notes
- **Always set `is_chat_model=True`.** `OpenAILike` defaults to
`is_chat_model=False`, which routes requests to the completions endpoint
instead of chat — and that fails against a chat-only proxy like the Omni
Gateway LLM proxy. `connection_kwargs()` always sets it (and
`is_function_calling_model=True`); set it yourself if you construct
`OpenAILike` outside the adapter.
- **`llm()` sets model defaults from the bare name.** LlamaIndex looks models
up by their exact OpenAI name, so a provider-prefixed name like
`openai/gpt-5-mini` would miss its reasoning-model handling and its
context-window table. `llm()` strips one `/` prefix before both
lookups, and still sends the prefixed name to the proxy:
- `context_window` comes from LlamaIndex's table (400,000 for `gpt-5-mini`).
For a name the table doesn't list, such as a Gemini or Bedrock model, it
stays at `OpenAILike`'s default of 3,900 tokens. That default also caps
agent memory and RAG prompt packing, so pass `context_window=` for those
models.
- A prefixed reasoning model (gpt-5, o-series) gets `temperature=1.0`, sends
`max_tokens` as `max_completion_tokens`, and sends `reasoning_effort`.
LlamaIndex already does the same for the bare name.
Any `temperature=`, `context_window=` or `additional_kwargs=` you pass takes
precedence. `connection_kwargs()` carries no model, so if you build
`OpenAILike` yourself, set these yourself.
- **`donkey.last_call` is set in the context that made the call.** A cold read
(no call yet in this context) on a `Donkey` that resolved only adapters like
this one still reports `status == LastCallStatus.UNAVAILABLE` rather than
`UNOBSERVED`, because LlamaIndex is still listed as not observing calls. Aligning
that, and the matching conformance exemptions, is tracked in [#740](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/740).
- **`jwt` mode is async-only.** `acomplete()` / `achat()` carry the JWT; sync
`complete()` / `chat()` go through the blocking client, which can't fetch
one, so they raise `ConfigError` before sending anything.
- **Printing the model shows credentials.** `OpenAILike`'s own `repr()` /
`str()` include `client_secret` and `api_key`. Don't print or log it; see
[What printed output hides](https://docs.donkey-kit.dev/reference/configuration.md#what-printed-output-hides).
See the [error taxonomy](https://docs.donkey-kit.dev/errors.md) for how proxy rejections surface as typed
exceptions, and the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md) for
the current status of every constructor signature this adapter depends on.
---
Source: https://docs.donkey-kit.dev/errors.md
# Governed error taxonomy
Live
The proxy doesn't just pass model calls through — it enforces policy. When it
rejects a call, DDK turns the response into a **typed exception** so you
branch on the governance outcome instead of parsing bodies.
## The rejection shapes `classify()` types
`classify()` types nine rejection shapes. **Neither the status code nor the
shape of the `error` value alone is a sufficient discriminator** — a `403` can be
PII, a killed agent, a regex-guard block, a content-safety block (all policy
blocks) *or* auth,
and the same nested-object envelope is emitted by both the upstream provider and
a gateway policy. The authoritative discriminator is the error **`type`** plus
specific headers.
| Rejection | HTTP | Discriminator | Maps to |
|---|---|---|---|
| Client-ID enforcement (auth) | `401` | flat `{"error":"…"}` + `www-authenticate: Client-ID-Enforcement` | `AuthError` |
| PII detected | `403` | nested `{"error":{type:"pii_detected"}}`, **no** `www-authenticate` | `PIIDetected` (parses `entities`) |
| Agent kill switch | `403` | nested `{"error":{code:"agent_killed"}}`, **no** `type`, **no** `www-authenticate` | `AgentKilled` |
| Injection protection | `400` | header `x-injection-protection: blocked` (**not** the status) | `PromptInjectionBlocked` |
| Regex prompt guard | `403` | top-level `matched_patterns` list (flat `error`) | `PromptInjectionBlocked` (`policy="regex-prompt-guard"`) |
| Content safety / guardrails | `403` | header `x-llm-proxy--…-action: reject` (Azure Content Safety / Bedrock Guardrails) | `ContentSafetyBlocked` (parses `categories`) |
| Token rate limit | `429` | **empty body**; `x-token-limit`/`-remaining`/`-reset` headers (ms) | `TokenBudgetExceeded` (`retry_after` derived) |
| Content moderation (undiscriminated) | `4xx` | falls through — no nested `error`, no injection/guard/safety discriminator | generic `PolicyViolation` |
| Upstream provider 4xx | `4xx` | nested `error` object **with** `code`/`type`/`param` — in an OpenAI-style object envelope `{"error":{…}}` **or** a Gemini-style list envelope `[{"error":{…}}]` (`status`→`error_type`) | `UpstreamRequestError` |
| Upstream 5xx | `5xx` | status range (no competing discriminator) | `UpstreamModelError` (retryable) |
| Bare model name on a multi-provider proxy | `400` | flat `{"error":"…"}` saying the model "is not in the known unique model map" | `ModelNotRoutable` (`.model`) |
`PIIDetected`, `AgentKilled`, the regex-prompt-guard check, and the content-safety
check are all evaluated **before** the generic 401/403→auth rule, because each is
a `403` (or `4xx`) that is *not* an auth failure. `AgentKilled` is also checked
before the upstream-4xx rule: its nested `error` object would otherwise read as
an upstream request mistake, when the upstream was never called. Likewise the injection check gates on the
`x-injection-protection` header, so an ordinary malformed `400` stays an
ordinary refusal.
Client-ID enforcement (`401`) is a **consumer-auth** case, not one of the nine
policy-rejection rows.
The Injection Protection shape is live-verified against a deployed proxy
(#669); `classify()` keys on the header discriminator alone for it, not the
body.
## The exception tree
All importable from `donkey_kit`, along with `classify`:
```
DonkeyError # base of the whole tree
├─ ConfigError # misconfiguration or client misuse (raised locally, no request sent)
├─ AuthError # rejected data-plane credentials or control-plane auth
├─ PolicyViolation # base for every governance rejection
│ ├─ PIIDetected # 403, type=pii_detected; .entities, .gateway_message
│ ├─ AgentKilled # 403, code=agent_killed — the Agent Kill Switch blocked this agent
│ ├─ TokenBudgetExceeded # 429; .retry_after (seconds)
│ ├─ PromptInjectionBlocked # x-injection-protection: blocked, or regex matched_patterns
│ └─ ContentSafetyBlocked # Azure Content Safety / Bedrock Guardrails vendor reject header; .categories
├─ GatewayUnavailable # transport failure — gateway unreachable, NO response; .base_url/.cause (ungoverned)
├─ UpstreamRequestError # upstream 4xx; .code/.error_type/.param
├─ ModelNotRoutable # 400, bare model name on a multi-provider proxy; .model
├─ UpstreamModelError # upstream 5xx — provider error, retryable
├─ BudgetReserveReached # client-side, from budget.pace(); .fraction_used/.reserve/.reset_at
├─ ModelSubstituted # client-side, opt-in; .requested_model/.served_model/.served_provider
├─ PlatformTeamOnly # Governance.apply() without the platform-team opt-in; also a PermissionError
└─ ToolInvocationError, RegistryError, PublicationDrift, ProvisioningError, GovernanceDrift
# tool access, registry, publishing and provisioning (Roadmap surfaces)
```
`GatewayUnavailable` is deliberately **not** under `PolicyViolation`: it is the one
*ungoverned* failure in the tree (see below). Everything under `PolicyViolation` is
something the gateway told the SDK; `GatewayUnavailable` is the gateway not being
there to tell it anything.
## Cookbook: every exception — discriminator, retryable, next step
One row per exception you can catch, with the three facts you need to write a
handler: **what tells it apart** (the discriminator), **whether retrying it can
ever succeed**, and **the next step** its `.remediation` names. "Retryable" here
means *by you* — the transport already retries the only class that is safe to
(`UpstreamModelError`), and treats every governance refusal as terminal so it can
never burn an exhausted budget or replay a blocked prompt.
| Exception | Discriminator | Retryable? | Next step (`.remediation`) |
|---|---|---|---|
| `AuthError` | Data plane: `401`, or `403` + `www-authenticate`. Control plane: Anypoint auth-provider or connected-app token acquisition fails. | **No** — terminal. In `jwt` mode the async data-plane client refreshes the wallet JWT and retries **once** on a `401`, then surfaces it; in client-id mode a data-plane `401` surfaces immediately. Control-plane token acquisition surfaces immediately and never affects model calls. | Data plane: check the consumer `client_id` / `client_secret` pair and API Manager authorization. Control plane: check the configured auth provider; for a connected app, verify `ANYPOINT_CLIENT_ID` / `ANYPOINT_CLIENT_SECRET` and the required scopes. |
| `PIIDetected` | `403`, nested `type: "pii_detected"`, **no** `www-authenticate` | **No** — a `PolicyViolation`, never retried. | Remove or redact the flagged values (`.entities`), or relax the policy's entity list in API Manager. |
| `AgentKilled` | `403`, nested `code: "agent_killed"`, **no** `type` | **No** — a `PolicyViolation`, never retried; the agent stays blocked until an administrator restores it. | Ask an administrator to restore this agent's model access in Governance > Security. |
| `TokenBudgetExceeded` | `429`, empty body, `x-token-*` headers | **Not immediately** — never auto-retried; only worth retrying *after* the window resets. | Wait for `.retry_after` (seconds) / the reset, then retry — or request an increase in API Manager. |
| `PromptInjectionBlocked` | header `x-injection-protection: blocked`, **or** a top-level `matched_patterns` list (regex prompt guard) | **No** — a `PolicyViolation`, never retried. | Review and sanitise the untrusted input, or adjust the policy's sensitivity / deny-list in API Manager. |
| `ContentSafetyBlocked` | `403` + `x-llm-proxy--…-action: reject` (Azure Content Safety / Bedrock Guardrails) | **No** — a `PolicyViolation`, never retried. | Revise the flagged content (`.categories`), or adjust the policy's categories / severity thresholds in API Manager. |
| `PolicyViolation` (generic) | a `4xx` matching **no** known rejection shape | **No** — terminal. | Inspect `.response`; file an issue with the status/headers/body so the shape can be typed. |
| `UpstreamRequestError` | non-`429` `4xx`, nested `error` with `code`/`type`/`param` (object **or** Gemini list envelope) | **No** — a client-side request mistake passed through the gateway, terminal. | Fix the flagged model or parameter (`.code` / `.param`); if `model_not_found`, request the model in API Manager. |
| `ModelNotRoutable` | `400`, flat `error` saying the model "is not in the known unique model map" (model-based routing with more than one provider) | **No** — a client configuration mistake, terminal; the upstream was never called. | Use the `provider/model` form, e.g. `openai/gpt-5-mini` instead of `gpt-5-mini`. |
| `UpstreamModelError` | `5xx` | **Yes** — the transport already retries `502` / `503` / `504`; a persistent `5xx` is safe for you to retry too. | Transient provider failure — retry, then escalate if it persists. |
| `GatewayUnavailable` | transport failure — DNS, refused connection, TLS, timeout — with **no** HTTP response | **Not automatically** — terminal here; you may retry or fall back. | Check host reachability, `.base_url`, and network egress; run [`donkey doctor`](https://docs.donkey-kit.dev/cli.md). |
Two more `DonkeyError`s are **client-side signals**, not gateway refusals, so
they sit outside the retry question. `BudgetReserveReached` is raised *before* a
call by [`donkey.budget.pace()`](https://docs.donkey-kit.dev/budget.md) and is meant to be recovered from
(`await donkey.budget.wait_for_reset()`, then continue) when its `.reset_at` is
known. If `.reset_at` is `None`, propagate or handle it instead — waiting returns
immediately and an unconditional retry would spin. Its `.remediation` carries
that branch as an inspectable next step, so you don't have to parse the
exception message. `ModelSubstituted` reports that a call *succeeded* against a
different model than requested (opt-in via `on_model_substitution="raise"`).
`ConfigError` is raised locally, before any request. It reports every missing
field at once, and it also reports an endpoint that may not receive the
configured credentials: a non-`https://` URL, or a URL from the project's config
files paired with credentials from elsewhere. See
[Which credentials a URL receives](https://docs.donkey-kit.dev/reference/configuration.md#which-credentials-a-url-receives).
Fix the config and re-run. The transport raises `ConfigError` for two lifecycle
mistakes as well, so neither escapes as a bare `RuntimeError`: a call on a
`Donkey` whose HTTP client is closed (after `aclose()`, or after a framework
closed the client it was given), and a call whose pooled connections belong to
an event loop that has closed. The SDK's own connection pools are per event
loop, so a second `asyncio.run()` on one `Donkey` works; this one comes from a
transport you passed in and reused across `asyncio.run()` calls. Its
`.remediation` names the fix for each. Through the OpenAI SDK it arrives
wrapped, like `GatewayUnavailable`: catch `openai.APIConnectionError` and read
the `ConfigError` from `e.__cause__`.
`AuthError.remediation` follows the plane that failed. Errors classified from
an LLM-proxy response use the canonical consumer-credential guidance that
[`donkey doctor`](https://docs.donkey-kit.dev/cli.md) also prints. Control-plane token failures override that
default with guidance for the provider that failed: connected-app errors point
to the Anypoint credentials and scopes, while an exhausted `ChainedAuth` points
to each configured provider's credential or token source.
## When the gateway can't be reached at all
Every rejection above describes something the gateway *told* the SDK.
`GatewayUnavailable` is the opposite: a transport-level failure — DNS, refused
connection, TLS error or timeout — with **no HTTP response** behind it. It is the
one *ungoverned* failure the taxonomy names, so a long-running agent can tell
"lost the gateway" apart from any other network fault and react — checkpoint,
queue, shed load, or fall back to a non-AI path — instead of pattern-matching a
raw `httpx` exception.
`DonkeyAsyncClient` and its blocking twin both raise it, so the async and sync
surfaces behave identically. It is terminal and **not retried**. It carries:
- `.base_url` — the origin that failed, on the exception, not only in the message.
- `.cause` — the underlying `httpx` exception (also chained via `raise … from`).
- `.request_id` — always `None`; there was no response to read the upstream provider's id from.
- `.correlation_id` / `.call_id` — the run and per-call ids the client sent, carried
even though no response came back, so the failure joins your logs like any other.
Its `.remediation` names the three real causes — an unreachable host, a wrong
base URL, or blocked network egress — and points at `donkey doctor` for
connectivity diagnosis.
## Every refusal names a next step
Every `DonkeyError` — not just every refusal — carries a non-empty,
human-readable `remediation`, so `except DonkeyError as e: log(e.remediation)`
is always safe. The constructor **raises** if you try to build one with a blank
remediation, every class accepts a `remediation=` override, and each class ships
a canonical default. For the refusals:
- `PIIDetected` → remove or redact the flagged values, or relax the policy's
entity list in API Manager.
- `AgentKilled` → ask an administrator to restore this agent's model access in
Governance > Security.
- `TokenBudgetExceeded` → wait for the window to reset (see `retry_after`) or
request an increase.
- `PromptInjectionBlocked` → review and sanitise the untrusted input, or adjust
the policy's sensitivity.
- `ContentSafetyBlocked` → revise the flagged content, or adjust the policy's
categories / severity thresholds.
The text names the **action you can take**, not the policy that fired. Because
each default lives on the exception class, it is a single source of wording that
[`donkey doctor`](https://docs.donkey-kit.dev/cli.md) reuses for its own failure output, so the CLI and the
exception never disagree.
## Refusal messages don't repeat blocked content
The PII policy's rejection text repeats every value it flagged. `PIIDetected`
builds its own message instead, from the entity types, their count and their
character offsets. So `str(exc)`, `repr(exc)` and `exc.args` never contain the
blocked value, and neither does a log line, a traceback or
[`donkey doctor`](https://docs.donkey-kit.dev/cli.md):
```
Request blocked: personally identifiable information detected (403): 1 entity (Email at chars 12-32). Values withheld; the gateway's text is on .gateway_message.
```
`.entities` lists the flagged types (`["Email"]`). The gateway's own text is on
`.gateway_message`, and the raw body is on `.response`. Neither is rendered by
`str()` or `repr()`. Both carry the blocked content, so handle them like the
prompt itself. The other refusal messages are built from status codes,
headers, category names and policy pattern names, never from the request.
A traceback also prints every chained exception, and a framework's own error
usually repeats the gateway's text. So when the SDK maps a framework error to
a typed one, as LangGraph's
[`typed_refusals()`](https://docs.donkey-kit.dev/frameworks/langgraph.md#typed-refusals-inside-a-node) does,
it raises it without a chained cause: `exc.__cause__` is `None`, and the
framework error is on `exc.framework_error` (`None` when there was none). No
frame in the traceback holds the framework error as a local variable, so error
reporters that print frame locals (Sentry, `pytest -l`) don't show it either.
Python still keeps it on the suppressed `exc.__context__`, so treat
`framework_error` and `__context__` like `gateway_message`.
This applies to policy refusals. `UpstreamRequestError` messages include the
upstream provider's own error text, which can quote parts of your request. For
everything the SDK does and doesn't hide in printed output, see
[What printed output hides](https://docs.donkey-kit.dev/reference/configuration.md#what-printed-output-hides).
## The ids every `DonkeyError` carries
Every exception in the tree carries three ids so you can join a failure to your
logs and to the gateway's own record:
| Attribute | What it is | Provenance |
| --- | --- | --- |
| `.correlation_id` | The **run** id, shared by every call in a [`donkey.run()`](https://docs.donkey-kit.dev/telemetry.md#correlation-ids) block | The `X-Correlation-Id` request header the client sent — always equals what went on the wire. |
| `.call_id` | The **per-call** id, unique per logical request and stable across that request's retries | The `X-Donkey-Request-Id` request header the client sent. Present **even when the request fails before any response** (a transport error). |
| `.request_id` | The **upstream provider's own** id, passed through by the gateway | Read back from a **response** header whose name varies by provider (`x-request-id` for OpenAI, `x-amzn-requestid` for Bedrock, `apim-request-id` for Azure, `request-id` for a native Anthropic proxy). Quote it to the provider's support team. Absent on a transport error, or on a route where the provider forwarded none. |
`classify(response)` fills `.correlation_id` and `.call_id` from the response's
own request, so bridging an `openai` error (below) needs no extra wiring — the
correlation id on the exception equals the header that was actually sent. (If you
overrode the header names in config, pass the ids to `classify()` explicitly.)
## Bridging from the raw client
`donkey.llm.client()` is the **OpenAI SDK**, so on an HTTP failure it raises
`openai.APIStatusError`, **not** a `DonkeyError`. Bridge into the taxonomy by
applying `classify()` to the error's `.response`.
```python
import openai
from donkey_kit import PIIDetected, TokenBudgetExceeded, AuthError, classify
try:
resp = await client.chat.completions.create(model="gpt-4o", messages=msgs)
except openai.APIStatusError as e:
governed = classify(e.response) # -> a DonkeyError subclass
if isinstance(governed, PIIDetected):
print("blocked, entities:", governed.entities)
elif isinstance(governed, TokenBudgetExceeded):
print("slow down; retry after", governed.retry_after, "s")
elif isinstance(governed, AuthError):
print("bad credentials:", governed)
else:
print(f"{type(governed).__name__}: {governed}")
except openai.APIConnectionError as e:
print("could not reach the proxy:", e)
```
The blocking client from `donkey.llm.client(sync=True)` behaves identically here
— drop the `await`. It is the same OpenAI SDK raising the same
`openai.APIStatusError`, and `classify()` reads the response the same way.
## Retry behaviour
Both clients retry only transient upstream/gateway failures (502/503/504) and
treat every 4xx as terminal — **including a 429**: on this proxy a 429 is a
token-budget refusal (`TokenBudgetExceeded`), so retrying it would only burn the
same already-exhausted window. `retry_after` is still surfaced for you to pace
against, but the transport never silently retries it.
In `jwt` mode the async client additionally refreshes the wallet JWT and
retries **once** on a 401. In the default client-id mode the data-plane client
holds no token, so a 401 is terminal, as it is on the blocking client, and
surfaces immediately as `AuthError`. The Anypoint control-plane credential
lives on a separate client: model calls never fetch, send or refresh it. See
[What the SDK sends where](https://docs.donkey-kit.dev/reference/configuration.md#what-the-sdk-sends-where).
## Unrecognised shapes
Any content-moderation or federated-guardrail response that matches none of the
discriminators above falls through to a generic `PolicyViolation` rather than an
invented type. DDK only types a refusal by a discriminator it can identify
reliably; everything else stays inspectable via `.response`.
One `400` that used to land here is now typed: a bare model name (`gpt-5-mini`)
sent to a model-based proxy with more than one provider. The gateway rejects it
before any upstream call, and `classify()` returns `ModelNotRoutable`. It keeps
the gateway's text and is not a `PolicyViolation`.
---
Source: https://docs.donkey-kit.dev/budget.md
# Budget & pacing
Live
Once a token-rate-limit policy is applied, the governed proxy reports your
token budget on its responses, in **two shapes** depending on the response:
| Response | Header | Example |
|---|---|---|
| Success `200` (and a `403` refusal) | `x-llm-proxy-ratelimit`, as prose | `Token rate limit: 10000 tokens remaining of 10000 limit. Reset in 56711ms.` |
| Budget refusal `429` | `x-token-limit`, `x-token-remaining`, `x-token-reset` | numeric values |
Parsing those yourself in every call site is tedious, and most code skips
it — so the first time budget matters is the moment it runs out.
With DDK **you never parse a header.** Every response that carries either shape
updates a `Budget` object on the `Donkey` instance. Where both are present the
numeric values win, and the prose header fills any field they leave unset.
Until the first such response, every field is `None`, never a misleading zero.
## The object
```python
donkey.budget.limit # int, tokens per window
donkey.budget.remaining # int, from the last response
donkey.budget.reset_at # datetime, converted from the ms-to-reset value — not raw
donkey.budget.observed_at # when headers were last seen (staleness)
donkey.budget.fraction_used # 0.0–1.0
```
And two helpers that use it:
```python
await donkey.budget.wait_for_reset() # sleeps until reset_at
async with donkey.budget.pace(reserve=0.10): # raises BudgetReserveReached at 90%
...
```
`pace()` raises **before** issuing the request that would cross your reserve —
not after a `429` comes back.
The budget object is per-`Donkey`, not global: two instances with different
credentials do not share state.
## Example: a batch job that finishes by itself
50,000 product records, enriched overnight against a governed model, budget
window resetting every hour, no human awake.
Without a budget object, the script runs flat out, takes a `429` at record
31,000, crashes, and someone re-runs it from record 0 in the morning —
spending the budget twice to do the same work. With `pace()`:
```python
for batch in chunks(records, 200):
while True:
try:
async with donkey.budget.pace(reserve=0.05):
await enrich(batch)
except BudgetReserveReached as exc:
if exc.reset_at is None:
raise # waiting cannot make progress without a reset time
await donkey.budget.wait_for_reset()
continue
break
checkpoint(batch)
```
Once `reset_at` has elapsed, the old observation is stale and `pace()` no longer
refuses, so the job continues unattended without a manual budget observation. A
later response updates the observed fields only when it carries a recognised
budget signal, and a fresh future `reset_at` makes the guard active again.
If a partial observation reports usage without a `reset_at`, the loop above
re-raises `BudgetReserveReached` after one attempt instead of calling
`wait_for_reset()` and spinning at zero delay. Preserve the last checkpoint and
escalate rather than crossing the reserve.
## Example: a dashboard that prevents the outage
`fraction_used` is per-agent, so it can be graphed. The owner of a support
agent sees it climbing at 14:00 and asks for an increase before the 16:00
peak — rather than explaining an outage afterwards.
## Budget is observed in-band
The gateway reports budget on response headers; there is **no endpoint that
answers "what is my remaining budget?"**. So `remaining` is only as fresh as
your last call, and a brand-new process knows nothing until its first request
completes.
This is why `observed_at` is part of the public surface: a dashboard reading
`remaining` without checking `observed_at` is reporting history, not state.
A budget-query endpoint on the gateway would make this object live rather than
last-known-good — see [Roadmap](https://docs.donkey-kit.dev/roadmap.md).
## Semantic cache steering
Live
When the proxy is fronted by the Anypoint **semantic-caching** policy, the
gateway can answer a request from a stored completion when a semantically
similar prompt was seen before — no provider round-trip, no fresh token spend.
DDK caches nothing and computes no embeddings itself (that stays on the
[do-not-build](https://docs.donkey-kit.dev/roadmap.md) list); it lets you **steer** the gateway's cache per
block and **surfaces** the outcome.
```python
# Skip the cache for a block where a fresh answer matters:
with donkey.cache(skip=True):
await agent.run(task)
# Or tighten the match and shorten entry lifetime:
async with donkey.cache(threshold=0.9, ttl=60):
...
```
`donkey.cache(...)` is a dual sync/async context manager — like
[`donkey.run(...)`](https://docs.donkey-kit.dev/telemetry.md#correlation-ids), the controls bind to a context variable, so they
reach every governed call in the block (including calls on framework-spawned
`asyncio` tasks) with no threading through framework state. The five controls:
| Control | Type | Effect |
|---|---|---|
| `skip` | `bool` | Bypass the cache policy entirely (passthrough to the provider). |
| `no_store` | `bool` | Look up, but do not write the result on a miss. |
| `ttl` | `int` | Override the entry time-to-live, in seconds (a positive int). |
| `threshold` | `float` | Override the similarity threshold, in `[0.0, 1.0]`. |
| `principal_id` | `str` | Override the id the similarity filter partitions on. |
The controls travel as `x-cache-*` request headers on model requests to the
LLM proxy only. Control-plane requests inside the block (the connected-app
token request, registry and tool calls) never carry them, so `principal_id`
doesn't leave the data plane.
An invalid control (a negative `ttl`, a `threshold` outside `[0.0, 1.0]`, a
`principal_id` with a control character) raises `ConfigError` **at the call
site**, not on the first request. The **outcome** of each call is on
[`donkey.last_call.cache_status`](https://docs.donkey-kit.dev/reference/last-call.md#semantic-cache) /
`.cache_score` and the OTel span. The same
[degradation](https://docs.donkey-kit.dev/frameworks.md) as `donkey.run(...)` applies: a
`connection_kwargs()`-only adapter that does not route through the shared
transport does not see the context variable, so its calls are not steered.
## Behaviour guarantees
- The reset value (milliseconds *until* reset, in both `x-token-reset` and the
prose `Reset in …ms`) is converted to a `datetime` anchored to
`observed_at`, accurate to the second.
- `pace()` raises before the request that would cross the reserve, never after
a `429`.
- After `reset_at`, `pace()` no longer refuses, so a wait-and-retry loop can
continue without a manual budget observation.
- If the reserve is reached without a known `reset_at`, the retry loop raises
once instead of spinning at zero delay.
- A semantic-cache **hit** is a verbatim replay with no provider round-trip, so
it never advances the budget window — the replayed `usage` is not fresh spend.
- Budget state is per-`Donkey` instance.
---
Source: https://docs.donkey-kit.dev/identity.md
# Identity
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned
design.
By default the gateway sees your **service**, not the person the agent is
acting for. That is fine until policy depends on who is asking. DDK will let you
run a block of agent code on behalf of a user:
```python
async with donkey.as_user(id_token=slack_user_oidc_token):
await hr_bot.answer(question) # the gateway sees the user, not just the service
```
Inside that block, DDK acquires a user-scoped token via RFC 8693 token exchange
and attaches it to every governed call, so the gateway's Trusted Agent Identity
layer can apply per-user policy. The exact token-exchange endpoint and header
the gateway expects will be pinned to the gateway's documented contract.
## The problem this solves
An HR bot must not answer a manager's salary question about a *different*
manager's report.
That decision belongs to the gateway — it is the component that holds the
identity policy, the org-chart relationships, and the audit obligation. DDK's
job is to get the user's token onto the request correctly so the gateway can
make the call it is already able to make.
## Where the boundary sits
**DDK does not implement authorisation logic.** Trusted Agent Identity is a
MuleSoft gateway feature. DDK does the token-exchange plumbing and header
placement — nothing more. "Check the user's role in the SDK" is client-side
enforcement, and it is on the [will-not-build list](https://docs.donkey-kit.dev/roadmap.md): code in your
process can be bypassed by code in your process.
## No silent fallback
The failure mode to design against is a token exchange that quietly fails and
lets the call proceed under the **service** identity. Per-user policy would stop
applying while everything still appears to work.
So entering `as_user()` will either attach a user-scoped token or raise. It never
degrades to the service identity without telling you.
---
Source: https://docs.donkey-kit.dev/hitl.md
# Human-in-the-loop
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned
design.
Refunds over €100 need a human. Today that is bespoke code, per team, per
framework. DDK will let you declare the approval requirement on the tool itself:
```python
@donkey.tool(approval="required", risk="financial")
async def issue_refund(ticket_id: str, amount: float): ...
```
When the agent calls `issue_refund`, DDK raises `ApprovalRequired` (or triggers
the framework's own interrupt), records the pending approval against the
correlation ID, and resumes when you resolve it:
```python
await donkey.approvals.resolve(approval_id, approved_by=reviewer.id)
```
## Mapped onto what your framework already has
DDK does not introduce a new pause mechanism. It maps one vocabulary onto the
primitive each framework already ships:
| Framework / protocol | Native primitive |
|---|---|
| LangGraph | `interrupt()` / `Command(resume=…)` |
| OpenAI Agents SDK | tool-approval / guardrail hooks |
| Google ADK | before/after tool callbacks |
| Strands | hooks |
| MCP | elicitation |
| Omni Gateway | Trusted Agent Identity step-up (MFA) |
## What DDK adds
Every framework already has human-in-the-loop. DDK adds three things on top,
none of which is a new mechanism:
1. **Normalisation** — one vocabulary across frameworks, so approval policy is
not rewritten when a team switches from ADK to LangGraph.
2. **Auditability** — the pending approval appears in the
[span](https://docs.donkey-kit.dev/telemetry.md), and the approver's identity lands in the audit trail
next to the correlation ID. "Who approved this refund?" becomes a query.
3. **Gateway routing** — a high-risk approval can be routed through the
gateway's identity layer for step-up MFA, rather than trusting a click in
your own UI.
## Out of scope
DDK will not ship an approval queue or an approval UI — both are on the
[will-not-build list](https://docs.donkey-kit.dev/roadmap.md). It integrates with whatever you already run:
Slack, ServiceNow, or LangGraph's own checkpointer.
A queue would mean owning a durable store, an escalation model, and a
notification system — a product rather than a feature.
---
Source: https://docs.donkey-kit.dev/policies.md
# Policy handshake
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned
design. It depends on a policy-discovery endpoint on the gateway, which does not
exist today.
On first connection, DDK will fetch the policy set in force for your
credentials and expose it:
```python
donkey.policies.models_allowed # ["gpt-4o", "claude-sonnet"]
donkey.policies.tools_allowed # [...]
donkey.policies.budget # the same Budget object, now live
donkey.policies.pii.mode # "block" | "mask" | "log"
donkey.policies.content_safety.on # True
donkey.policies.observed_at # when this view was fetched
```
## What it enables
**Fewer wasted calls.** An agent that requests a model outside its allow-list
costs one round-trip and one refusal, every time. With the handshake, the
adapter picks from `models_allowed` at construction time and the refusal never
happens. Multiply by a few thousand tickets a day.
**Better UX.** If `pii.mode == "mask"`, the gateway will redact rather than
reject — so your bot can say *"some details were redacted"* instead of
*"request failed."* Same gateway behaviour, a very different experience, and
today the client has no way to know which mode is in force.
**Warm start.** It also makes [budget](https://docs.donkey-kit.dev/budget.md) live rather than
last-known-good: a fresh process currently knows nothing about its budget until
its first response comes back.
## The handshake is advisory
**The handshake never makes an access decision.** It exists to avoid *wasted*
calls and to improve UX. The gateway still evaluates every request.
If the client's cached view and the gateway disagree, **the gateway wins** and
the client learns from the refusal. Skipping a gateway call *"because the
handshake said it's fine"* is client-side enforcement, which is on the
[will-not-build list](https://docs.donkey-kit.dev/roadmap.md).
This is also why `observed_at` is part of the surface, as it is on the budget
object: a policy view is a snapshot, and code that cannot tell a snapshot from
live state will eventually make a decision it should not have.
---
Source: https://docs.donkey-kit.dev/telemetry.md
# Telemetry & cost
Live
DDK answers two questions about every governed call: *what happened?* and *who
pays for it?* It does so with OpenTelemetry GenAI spans, per-run correlation
IDs, cost-attribution tags, and routing and resilience signals.
## OpenTelemetry GenAI spans
Every governed call produces a span following the OpenTelemetry **GenAI
semantic conventions** — `gen_ai.system`, `gen_ai.request.model`,
`gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens` — plus attributes
for the governance layer that generic instrumentation cannot know about:
```
donkey.policy.decision = allow | refuse
donkey.policy.type = pii_detected | token_budget | injection | agent_killed | …
donkey.budget.remaining = 18450
donkey.correlation_id = …
donkey.cost.team = support
donkey.cost.project = triage-v2
donkey.cost.env = prod
donkey.cost.enduser.id = user-42
donkey.routing.type = ModelBased # how the gateway routed
donkey.routing.fallback = false # did it fail over?
gen_ai.response.model = gpt-5.1 # the model that actually served
donkey.cache.status = hit # semantic-cache outcome (on a cached proxy)
donkey.cache.score = 0.9518 # similarity score, hit only
donkey.usage.cached_tokens = 512 # omitted when the provider reports none
donkey.usage.cache_write_tokens = 128
donkey.usage.reasoning_tokens = 96
```
The three `donkey.usage.*` counts carry the cost-relevant detail tokens the
semantic conventions have no pinned key for — cached / cache-write prompt tokens
and reasoning-model thinking tokens. They are read from the response `usage`
block's detail sub-objects and are **omitted, never `0`,** when the provider
reports no detail counts. When the response passes through the SDK's shared HTTP
client, the same counts are exposed per-call on `donkey.last_call`.
A **refused** request still produces a span, with `donkey.policy.decision=refuse`
and `otel.status_code=ERROR`. A streaming response produces **exactly one** span,
with token counts filled in at stream end.
Export goes over OTLP to wherever you already send spans. **Nothing in the
emit path is Anypoint-specific**, so if your team already runs Grafana, Elastic,
Honeycomb, New Relic, Datadog, Dynatrace, Langfuse or Phoenix, "policy refusals
per hour by type" and "tokens per ticket" show up in the dashboard you already
have. See [Send spans to your observability backend](#send-spans-to-your-observability-backend).
### Zero-config export
Set the standard OpenTelemetry endpoint env var and spans flow — there is **no
SDK-specific variable**.
```bash
pip install "donkey-kit[otel]"
export OTEL_EXPORTER_OTLP_ENDPOINT=https://langfuse.acme.internal
export OTEL_SERVICE_NAME=support-triage # standard OTel var, honoured for free
python -m my_app # spans flow, refused calls included
```
`Donkey.from_env()` reads `OTEL_EXPORTER_OTLP_ENDPOINT` (or the traces-specific
`OTEL_EXPORTER_OTLP_TRACES_ENDPOINT`) and, when one is set, installs an OTLP
exporter behind a batch processor. The network flush runs on that background
thread, off your request path, keeping instrumentation overhead under
1 ms per call. The `[otel]` extra ships the **http/protobuf** exporter;
`OTEL_EXPORTER_OTLP_PROTOCOL=grpc` is honoured only if you also install
`opentelemetry-exporter-otlp-proto-grpc`.
**With no endpoint set, the export path is inert and silent** — no exporter is
built, nothing connects, nothing is printed. If your process already configures
its own OpenTelemetry provider (say via `opentelemetry-instrument`), DDK rides it
rather than replacing it, so your spans flow through the pipeline you already
set up.
Opt out of telemetry entirely with a single flag:
```bash
export DONKEY_TELEMETRY=false # or telemetry = false in .donkey-kit.toml
```
### Send spans to your observability backend
Any backend that accepts OTLP works with the same three standard variables:
`OTEL_EXPORTER_OTLP_ENDPOINT`, `OTEL_EXPORTER_OTLP_HEADERS` for the vendor's
credentials, and `OTEL_SERVICE_NAME`. There is no DDK-specific setting and no
code change between backends. Pick yours:
Grafana Cloud stores traces in **Tempo**. In the Grafana Cloud portal, open your
stack and choose **Configure** on the **OpenTelemetry** tile. It generates the
endpoint for your region and a token, already base64-encoded as
`instanceID:token`.
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp-gateway-prod-.grafana.net/otlp"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Basic%20"
export OTEL_SERVICE_NAME="support-triage"
```
Grafana's own note for Python: write the space after `Basic` as `%20`. Find the
spans in **Explore → Tempo** or in **Application Observability**. For
production, Grafana recommends sending through **Grafana Alloy** (its
OpenTelemetry Collector) rather than straight from the app.
Docs: [Send data to the Grafana Cloud OTLP endpoint](https://grafana.com/docs/grafana-cloud/send-data/otlp/send-data-otlp/)
On Elastic Cloud Hosted or Serverless, copy the **Managed OTLP endpoint** from
Kibana (**Add data → Applications → OpenTelemetry**), which also generates an
API key. Prefix the key with `ApiKey` yourself; Kibana gives you the bare key.
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="https://"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=ApiKey%20"
export OTEL_SERVICE_NAME="support-triage"
```
Spans appear in Kibana under **Observability → Applications**, grouped by
service. Self-managed Elasticsearch exposes the same protocol at
`/_otlp`; its API key needs `create_doc` and `auto_configure`
privileges.
Docs: [Send OTLP data to Elastic](https://www.elastic.co/docs/solutions/observability/get-started/quickstart-elastic-cloud-otel-endpoint) ·
[Elasticsearch OTLP/HTTP endpoint](https://www.elastic.co/docs/manage-data/ingest/otlp-endpoint)
Create an ingest API key in your Honeycomb environment settings.
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="https://api.honeycomb.io" # EU: https://api.eu1.honeycomb.io
export OTEL_EXPORTER_OTLP_HEADERS="x-honeycomb-team="
export OTEL_SERVICE_NAME="support-triage"
```
Honeycomb creates a dataset named after `OTEL_SERVICE_NAME`. Query
`donkey.policy.decision`, `donkey.cost.team` or `gen_ai.usage.output_tokens` like
any other field; for example, a `COUNT` where `donkey.policy.decision = refuse`,
grouped by `donkey.policy.type`, gives you refusals by type.
Docs: [Using the Honeycomb OpenTelemetry endpoint](https://docs.honeycomb.io/send-data/opentelemetry/)
Authenticate with your account's **license key** in an `api-key` header.
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp.nr-data.net" # EU: https://otlp.eu01.nr-data.net
export OTEL_EXPORTER_OTLP_HEADERS="api-key="
export OTEL_SERVICE_NAME="support-triage"
```
New Relic recommends `http/protobuf`, which is DDK's default. The service shows
up under **APM & Services**, and span attributes are queryable with NRQL, for
example `SELECT count(*) FROM Span WHERE donkey.policy.decision = 'refuse' FACET donkey.policy.type`.
Docs: [New Relic OTLP endpoint](https://docs.newrelic.com/docs/opentelemetry/best-practices/opentelemetry-otlp/)
The simplest route is the **Datadog Agent**, which accepts OTLP once its receiver
is enabled. Start the Agent with
`DD_OTLP_CONFIG_RECEIVER_PROTOCOLS_HTTP_ENDPOINT=0.0.0.0:4318` and expose port
`4318`, then point DDK at it. No credentials are needed in the app; the Agent
holds the API key.
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318" # the Agent's host
export OTEL_SERVICE_NAME="support-triage"
```
Spans appear in **APM → Traces**. Datadog also offers an agentless OTLP traces
intake that takes a `dd-api-key` header; its URL depends on your Datadog site.
Docs: [OTLP ingestion by the Datadog Agent](https://docs.datadoghq.com/opentelemetry/setup/otlp_ingest_in_the_agent/) ·
[OTLP traces intake endpoint](https://docs.datadoghq.com/opentelemetry/setup/otlp_ingest/traces/)
Create an access token with the `openTelemetryTrace.ingest` scope.
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="https://.live.dynatrace.com/api/v2/otlp"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Api-Token%20"
export OTEL_SERVICE_NAME="support-triage"
```
Use `live.dynatrace.com`, not `live.apps.dynatrace.com`. Through an ActiveGate the
base URL is `https://:9999/e//api/v2/otlp`. Spans
appear in **Distributed Tracing**.
Docs: [Dynatrace OTLP API endpoints](https://docs.dynatrace.com/docs/ingest-from/opentelemetry/otlp-api)
Langfuse is built for LLM traces, so the `gen_ai.*` model and token attributes
on DDK spans are what it reads. Authenticate with your project's public and
secret key, base64-encoded:
```bash
AUTH_STRING=$(echo -n "pk-lf-...:sk-lf-..." | base64)
export OTEL_EXPORTER_OTLP_ENDPOINT="https://cloud.langfuse.com/api/public/otel" # US: https://us.cloud.langfuse.com/api/public/otel
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Basic%20${AUTH_STRING},x-langfuse-ingestion-version=4"
export OTEL_SERVICE_NAME="support-triage"
```
The `x-langfuse-ingestion-version=4` header makes spans appear in real time
instead of after a delay of up to ten minutes. Self-hosted Langfuse (v3.22 or
later) takes the same path on your own host.
Docs: [Langfuse OpenTelemetry](https://langfuse.com/docs/opentelemetry/get-started)
Phoenix is open source and runs locally with no account, which makes it a
quick way to look at LLM spans during development:
```bash
docker run -p 6006:6006 -p 4317:4317 arizephoenix/phoenix:latest
```
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:6006"
export OTEL_SERVICE_NAME="support-triage"
```
Open `http://localhost:6006` to browse the spans. A Phoenix deployment with
authentication enabled also needs its API key as a bearer token:
`OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer%20"`.
Docs: [Phoenix configuration](https://arize.com/docs/phoenix/self-hosting/configuration) ·
[Get started with tracing](https://arize.com/docs/phoenix/get-started/get-started-tracing)
Jaeger's all-in-one container is the fastest way to see spans on your laptop,
with no account and no credentials:
```bash
docker run --rm --name jaeger -p 16686:16686 -p 4318:4318 \
cr.jaegertracing.io/jaegertracing/jaeger:latest
```
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318"
export OTEL_SERVICE_NAME="support-triage"
```
Open `http://localhost:16686`, pick the `support-triage` service and search.
Storage is in memory, so traces are gone when the container stops.
Docs: [Jaeger getting started](https://www.jaegertracing.io/docs/latest/getting-started/)
A few rules hold for every backend:
- **Base endpoint vs traces endpoint.** `OTEL_EXPORTER_OTLP_ENDPOINT` is a base
URL: the exporter appends `/v1/traces`. If you set
`OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` instead, give the full URL including
`/v1/traces`, or the backend answers `404`.
- **Encode spaces in headers.** Header values such as `Basic ` or
`ApiKey ` should be written with `%20` in place of the space; separate
several headers with commas.
- **Protocol.** DDK exports `http/protobuf`, which every backend above accepts.
For gRPC, set `OTEL_EXPORTER_OTLP_PROTOCOL=grpc` and install
`opentelemetry-exporter-otlp-proto-grpc`.
- **In production, consider a Collector.** Sending to an
[OpenTelemetry Collector](https://opentelemetry.io/docs/collector/) (or a
vendor distribution such as Grafana Alloy) lets you batch, retry, redact and
fan out to several backends without touching the app.
- **Prompts stay private by default.** None of these backends receives prompt or
completion text unless you opt in; see
[Message content stays off spans by default](#message-content-stays-off-spans-by-default).
### Semantic-convention stability and sinks
**The GenAI conventions are still `Development` status upstream**, so
attribute names can change. DDK **pins** the semconv version and
**dual-emits**: `gen_ai.*` at the pinned version, plus a stable `donkey.*`
namespace under this project's control. Your dashboards do not break when
upstream renames something.
DDK exports standard OTLP and leaves the sink to you. Whether Anypoint
Monitoring or Agent Visualizer ingests OTLP GenAI spans is not publicly
documented, so don't rely on spans appearing there.
### Message content stays off spans by default
Spans carry **metadata only** — model, token counts, policy decision, cost tags,
correlation id. They do **not** carry prompt or completion text.
**Spans are emitted upstream of the gateway's PII masking.** The Omni Gateway
masks sensitive content in *its* logs; a DDK span is created inside your
process, before the request reaches the gateway. Putting message text on the
span would re-export the very content the platform masks — straight to
whatever OTLP collector you have wired up. So capturing content is **opt-in,
and opting in is you assuming that obligation.**
Turn it on only when your collector is a trusted sink and you have accepted
responsibility for the content that lands there:
```bash
export DONKEY_TELEMETRY_CAPTURE_CONTENT=true
# or: telemetry_capture_content = true under [donkey] in .donkey-kit.toml
```
In code, layer it over the resolved config. It isn't a `Donkey.from_env()`
kwarg:
```python
from donkey_kit import Donkey, DonkeyConfig
donkey = Donkey(DonkeyConfig.from_env().with_overrides(telemetry_capture_content=True))
```
`telemetry_capture_content` resolves along the standard
[precedence](https://docs.donkey-kit.dev/reference/configuration.md#precedence) (set in code → env → config
files → default) and **defaults to `False`**. When
enabled, content is emitted under the pinned semconv attribute names —
`gen_ai.prompt` and `gen_ai.completion` — and no others. When off, those
attributes never reach a span, and the allowlist that builds every span drops
any content-shaped attribute a call site hands it, so there is no accidental
path for message text to leak.
## Routing & resilience
The gateway can fail over between providers when one degrades ("Enhanced
Resilience for Intelligent Routing"). It reports what it *did* with each request
on the response — which provider and model served it, how it routed, and whether
that was a **fallback**. DDK reads those signals off the shared transport, so
you get them with **no framework required** — the raw `donkey.llm.client()` path
benefits just as the deep adapters do.
Every governed call through the SDK's shared HTTP client exposes them on
`donkey.last_call`, beside the usage and identity fields:
```python
donkey = Donkey.from_env()
await donkey.openai().responses.create(model="gpt-5.1", input="…")
r = donkey.last_call
r.requested_model # "gpt-5.1" — what you asked for
r.served_model # "gpt-5.1" — or a substitute after failover
r.served_provider # "openai"
r.routing_type # "ModelBased" — or "Semantic" on a semantic-routing proxy
r.fallback # False — True if the gateway failed over
r.substituted # False — served_model != requested_model
r.matched_topic # None — the topic a Semantic proxy matched (else None)
r.routing_score # None — that match's similarity score (else None)
```
On a **semantic-routing** proxy (`routing_type == "Semantic"`), the gateway also
reports *why* it picked a provider: the topic your prompt matched and the
similarity score behind that match. Those land on `r.matched_topic` /
`r.routing_score`. On a model-based proxy both are `None` — the gateway emits no
semantic header, and DDK never fabricates a value for a signal it did not
observe.
They also land on the span (`gen_ai.response.model`, `donkey.routing.type`,
`donkey.routing.fallback`, and — on a semantic route —
`donkey.routing.matched_topic` / `donkey.routing.score`) — the most useful thing
to have on hand when latency spikes: it tells an operator whether a slow call was
routed normally or recovered from a degraded provider.
For the complete list of `last_call` fields — observability status, gateway
identity, routing, and token usage — see the [`last_call` field
reference](https://docs.donkey-kit.dev/reference/last-call.md).
The same record also carries the call's identity and usage: `request_id` (the
upstream provider's id), `api_instance_id` and `environment_id` (which gateway
instance served it), and the token counts including cached and reasoning
tokens. The [gateway identity example](https://docs.donkey-kit.dev/examples/general/gateway-identity.md) walks
through every field.
### When `last_call` is unavailable
`donkey.last_call` is populated only when the governed response passes through
the SDK's shared HTTP client, and only in the context that made the call.
CrewAI routes outside that client: its native OpenAI provider builds its own,
so no response reaches the record and the run's correlation ID isn't sent.
LlamaIndex, Microsoft Agent Framework and ADK's `model()` send through the
shared client, so a call populates the record in its own context and carries
the run's correlation ID. They are still listed as not observing calls, so a
cold read on a `Donkey` that resolved only these adapters or CrewAI reports
`UNAVAILABLE` rather than `UNOBSERVED`. Aligning that is tracked in
[#740](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/740).
ADK's `gemini()` uses the shared client and observes calls once it has been
called (see [Native Gemini](https://docs.donkey-kit.dev/frameworks/adk.md#native-gemini) for reading
`last_call` inside an ADK run).
For a `Donkey` that resolved only CrewAI:
```python
r = donkey.last_call
r.status # LastCallStatus.UNAVAILABLE
r.available # False
r.surface # "crewai"
```
This is different from `UNOBSERVED`, which means the current context has not
yet received a governed response. On an unavailable surface the SDK cannot
observe any response-derived `last_call` field, including gateway identity,
routing, fallback, and usage. If multiple non-observing adapters were resolved,
`surface` lists their names.
An agent runner that sends the model call from a task of its own leaves its
caller reading `UNOBSERVED`, even though the SDK observed the call: the record
lands in that task. The OpenAI Agents SDK's `Runner` sends a run's first turn
that way (and every turn of `run_streamed`), and ADK's `Runner` runs the whole
agent that way. Read `last_call` in the framework's model hook, which runs in the
call's own task: `RunHooks.on_llm_end` for the
[OpenAI Agents SDK](https://docs.donkey-kit.dev/frameworks/openai.md#donkeylast_call), `after_model_callback`
for [ADK](https://docs.donkey-kit.dev/frameworks/adk.md#native-gemini). Calling the model directly (Strands'
`Agent.invoke_async`, `AsyncAnthropic.messages.create`, a model's own async
method) records in the caller's context.
### Two behaviours worth knowing
**The SDK never double-retries a fallback.** DDK retries `502/503/504` with
backoff, but if the gateway already failed over internally, a `503` it marked as
a fallback is **not** retried again — a second recovery layer stacked on a
working first one just multiplies latency against an outage the gateway already
handled.
**Opt in to model determinism.** A silent substitution is surfaced passively on
`last_call.substituted` by default. When a substitution is not acceptable — your
evaluation, cost model and token assumptions are all pinned to one model — opt
into a hard error:
```python
donkey = Donkey.from_env(on_model_substitution="raise")
# raises ModelSubstituted when the served model differs from the requested one
```
`on_model_substitution` resolves along the standard precedence (kwarg → env
`DONKEY_ON_MODEL_SUBSTITUTION` → `.donkey-kit.toml` → default) and **defaults to
`"off"`**.
On a model-based routing proxy you address models as `provider/model` (for
example `openai/gpt-5-mini`), and the gateway reports the served model without
the prefix (`gpt-5-mini`), with the provider in its own header. A prefix that
names the served provider is not counted as a difference, so asking for
`openai/gpt-5-mini` and being served `gpt-5-mini` by `openai` is not a
substitution. The same model served by a different provider still is. On a
proxy with more than one provider, a bare name the gateway cannot pin to one
provider is rejected with a `400` before any upstream call, and DDK raises
[`ModelNotRoutable`](https://docs.donkey-kit.dev/errors.md).
### Semantic caching & semantic routing Live
Omni Gateway can answer a request from its **semantic cache** (a similar
prompt was answered before, so no provider round-trip and no token cost) and
can **route semantically** (pick the model by matching the prompt to a
topic). Both happen at the gateway. DDK caches nothing and computes no
embeddings itself; it lets you steer the gateway's cache per request — skip it,
not store a response, or override the TTL or similarity threshold via
[`donkey.cache(...)`](https://docs.donkey-kit.dev/budget.md#semantic-cache-steering) — and reports what
happened:
- the cache outcome (`hit`, `miss`, `bypass`, `no-store`) and similarity score
on `donkey.last_call.cache_status` / `.cache_score` and the span
(`donkey.cache.status` / `donkey.cache.score`);
- the matched routing topic and its score, beside the routing fields above.
A cache `hit` is a verbatim replay with no provider round-trip, so it never
advances the [token budget](https://docs.donkey-kit.dev/budget.md) — the replayed `usage` is not fresh spend.
### TypeScript parity Roadmap
The same signals will surface as `donkey.lastRouting.servedModel` / `.fallback`
in the TypeScript SDK.
## Correlation IDs
Set a per-**run** id once, and every call inside the block carries it — on the
wire, on every span, and on every exception — with nothing threaded through your
framework state.
```python
async with donkey.run(id=ticket.id):
await triage_agent.run(ticket)
```
`donkey.run(id=…)` binds **two ids**:
- a **run id** → the `X-Correlation-Id` request header → the span's
`donkey.correlation_id` → `DonkeyError.correlation_id`. Shared by every call in
the block, so a client-side log line **joins** to the gateway's own record.
- a fresh **per-call id** → the `X-Donkey-Request-Id` request header →
`DonkeyError.call_id`. Unique per logical request, stable across that request's
retries, so one call is pinpointable within a run.
The gateway reads the inbound `X-Correlation-Id` and echoes it verbatim on the
response, so the request and response `x-correlation-id` are the same value.
`X-Donkey-Request-Id` is a client-owned per-call id the gateway does not consume.
Both header names are overridable (`correlation_header` / `call_id_header`) for a
gateway that expects different ones.
Propagation is contextvar-based, so it reaches through framework nodes (every
LangGraph node, for instance) without threading an argument through every
function, and concurrent runs never leak into each other. It works with or
without OpenTelemetry installed. `donkey.run(...)` is a **dual sync/async**
context manager (plain `with` works too); nested blocks rebind then restore.
### The decorator form
When a whole function should be one run, `@donkey.governed` is the decorator
equivalent of wrapping its body in `donkey.run()`:
```python
@donkey.governed(team="support")
async def handle_ticket(ticket):
await triage_agent.run(ticket)
```
Each call opens its own run — a fresh run/correlation id — and binds the
optional per-run cost tags, the OTel span, and typed refusals, exactly the scope
`donkey.run()` establishes. It wraps **both sync and async** callables and is
usable bare (`@donkey.governed`) or parametrised. There is deliberately
**no `id=`**: pinning one id across every call would collapse unrelated runs into
a single correlation, so when you need a specific id, use `donkey.run(id=…)`
directly.
## Cost-attribution tags
A small, fixed set of tags (`team`, `project`, `env`, `enduser.id`) set once
and recorded on every model call's span:
```python
donkey = Donkey.from_env(team="support", project="triage-v2", env="prod")
async with donkey.run(id=ticket.id, enduser_id=agent_user.id):
await triage_agent.run(ticket)
```
The tags resolve in this order: `Donkey.from_env(team=…)` kwargs, then
`DONKEY_COST_*` env vars, then a `[donkey.cost]` table in the config files (see
[Configuration](https://docs.donkey-kit.dev/reference/configuration.md#cost-attribution-tags)). Per-run
overrides layer on top: `donkey.run(team=…, project=…, env=…, enduser_id=…)`
wins **per field** for its block, and the rest fall back to the configured
tags. The key set is **fixed**: an unknown dimension is a configuration error,
never a silently dropped tag. Values are **validated** (fixed keys, bounded
length), so nobody stuffs a JSON blob into a span attribute or header.
### Where the tags go
| Carrier | Names | When | Goes to |
|---|---|---|---|
| Span attributes | `donkey.cost.team`, `donkey.cost.project`, `donkey.cost.env`, `donkey.cost.enduser.id` | Always, for each tag that is set, on every `donkey.llm.chat` span. Needs `telemetry` on (the default) and OpenTelemetry installed. The SDK's shared client opens the span, so CrewAI, which has only header injection, records none. | Your OpenTelemetry pipeline: the OTLP exporter when `OTEL_EXPORTER_OTLP_ENDPOINT` is set, or any `TracerProvider` your process installed |
| Request headers | `X-Anypoint-Cost-Team`, `X-Anypoint-Cost-Project`, `X-Anypoint-Cost-Env`, `X-Anypoint-Cost-Enduser-Id` (renamable with `cost_*_header`) | Only with `send_cost_headers` / `DONKEY_SEND_COST_HEADERS=true`. Off by default. | Model requests to the LLM proxy only. See the note below. |
Each header carries the tag's value unchanged (for example
`X-Anypoint-Cost-Enduser-Id: user-42`). A tag that isn't set is sent under
neither carrier.
The Anypoint LLM Gateway doesn't read cost tags from request headers: it meters
cost from token usage per API instance and consuming client application. So the
span attributes are the carrier that matters. Turn the headers on only when
something of your own reads them.
**With `send_cost_headers` on, the headers go only on model requests to the
LLM proxy.** The connected-app token request and other Anypoint
control-plane requests never carry them, so the end-user ID doesn't reach
those hosts. Adapters with transport injection also send per-run
`donkey.run(...)` values; adapters that get only a `default_headers` snapshot
send the configured tags.
**The end-user ID reaches your trace backend regardless of
`send_cost_headers`.** If `enduser_id` is set, every model-call span carries
`donkey.cost.enduser.id`, and any exporter you've configured receives it. To
keep it out of traces, don't set `enduser_id` (`DONKEY_COST_ENDUSER_ID`,
`"enduser.id"` in `[donkey.cost]`, or `donkey.run(enduser_id=…)`), drop the
attribute in your collector, or turn spans off with `DONKEY_TELEMETRY=false`.
### The question this answers
Finance asks what the support agent cost last month versus the HR bot. Without
tags, both agents share one `client_id` and there is no way to split the bill.
With tags it is a group-by.
For compliance — *"prove the HR bot's answer to user X on date Y went through
the content-safety policy"* — the correlation ID on the log line joins to the
gateway record, and the span carries `donkey.cost.enduser.id` and `donkey.policy.type`. An
EU AI Act Article 12 log request becomes one query, not an investigation.
---
Source: https://docs.donkey-kit.dev/simulator.md
# Local simulator
Live
Nobody can make the production gateway emit a PII block on cue, so the branch
of your agent that handles `PIIDetected` usually runs for the first time on a
real customer's data.
`donkey mock` fixes that. It is a local HTTP server that behaves like the
governed LLM proxy *for the failure paths*: it replays captured rejection
fixtures on demand and a captured happy-path completion for everything else.
```bash
pip install "donkey-kit[cli,local]"
donkey mock --port 8080 # --host defaults to 127.0.0.1
```
Point **any** client at `http://localhost:8080` — the SDK, a stock OpenAI
client, or cURL. To trigger a specific rejection, set the request's **`model`**
to the sentinel `donkey-sim/`:
```bash
# no /v1 segment — the governed proxy (and the simulator) has none
curl -s http://localhost:8080/responses \
-H 'content-type: application/json' \
-d '{"model": "donkey-sim/pii-detected"}' -i | head -1
# HTTP/1.1 403 Forbidden (byte-identical to the captured PII block)
```
The selectable shapes are `token-rate-limit`, `pii-detected`,
`injection-protection`, `regex-prompt-guard`, `content-safety`,
`content-moderation`, `agent-killed`, `model-not-found`, `upstream-5xx`,
`client-id-missing`, and `model-not-routable` — the nine documented rejections
plus the consumer-auth `401` and the gateway's `400` for a bare model name on a
multi-provider proxy (`ModelNotRoutable`). One happy-path
variant is selectable the same way: `donkey-sim/success-semantic` replays the
captured **semantic-routing** `200` (`routing_type == "Semantic"`), so
`donkey.last_call.matched_topic` and `routing_score` light up offline (see
[Gateway identity](https://docs.donkey-kit.dev/examples/general/gateway-identity.md#semantic-routing-the-matched-topic-and-score)).
Any other `model` value gets the default model-based happy path. The
`donkey-sim/` prefix is a simulator-only control surface; the real gateway never
interprets it.
Because each body is the *same fixture* the SDK's error classifier is tested
against, a sentinel request surfaces in your agent as the typed exception — a
[`PIIDetected`](https://docs.donkey-kit.dev/errors.md), not a raw `403`.
## Choosing a port or host
`donkey mock` binds `127.0.0.1:8080` by default. Two flags move it:
| Flag | Default | Meaning |
|---|---|---|
| `--port` | `8080` | TCP port to bind. Pass another value if `8080` is in use. |
| `--host` | `127.0.0.1` | Interface to bind. `0.0.0.0` exposes it beyond localhost. |
```bash
donkey mock --port 9000 --host 0.0.0.0
```
Binding `0.0.0.0` makes the simulator reachable from other machines on your
network. It is a fixture replay that enforces no policy and ignores
credentials, so only expose it on a network you trust.
Point your client at the same address: set
`DONKEY_LLM_PROXY_URL=http://127.0.0.1:` (for the example above,
`http://127.0.0.1:9000`). The SDK accepts plain `http://` for loopback hosts
(`localhost`, `127.0.0.0/8`, `::1`); `http://0.0.0.0:…` is refused. To reach the
simulator from another machine or container by LAN address or hostname (for
example `http://simulator:9000`), set `DONKEY_ALLOW_HTTP=1` in that client's
environment; the SDK then emits a `ConfigWarning` naming the endpoint. See
[Endpoints must use `https://`](https://docs.donkey-kit.dev/reference/configuration.md#endpoints-must-use-https).
Behind an HTTP proxy (`HTTP_PROXY` or `ALL_PROXY` set, as on many corporate
machines and CI runners), the client sends loopback requests to the proxy
too, and the proxy usually cannot reach your simulator. Add the simulator's
host to `NO_PROXY`, for example `NO_PROXY=127.0.0.1,localhost`.
The simulator's port is **not** read from an environment variable or
`.donkey-kit.toml` — the `--port`/`--host` flags (or the `serve()` keyword
arguments below) are the only way to change it.
Set the simulator URL in the environment (or in code), as above. If you keep it
in `.donkey-kit.toml` instead, put the credentials in `.donkey-kit.local.toml`
next to it: a URL from the working directory's config files, loopback included,
only receives credentials from those files. See
[Which credentials a URL receives](https://docs.donkey-kit.dev/reference/configuration.md#which-credentials-a-url-receives).
If you run the simulator's ASGI app under `uvicorn` yourself instead of
through `donkey mock`, remember uvicorn's own default port is `8000`, not
`8080`. Using `donkey mock` (or `serve()`) gives you the documented `8080`
default.
### Running the simulator from Python
`donkey mock` is a thin wrapper over `serve()`. To boot the simulator from your
own code, call it directly (the `[local]` extra must be installed):
```python
from donkey_kit.simulator import serve
serve(host="127.0.0.1", port=9000)
```
`serve(*, host="127.0.0.1", port=8080, config=None)` is keyword-only and blocks
until the process is stopped. Pass a `SimulatorConfig` (also importable from
`donkey_kit.simulator`) as `config=` to script the same fault-injection
scenarios the `--scenario` flag builds. If the `[local]` extra is missing,
`serve()` raises `ImportError`.
## Budget windows on the happy path
Happy-path `200` responses carry a synthesised `x-llm-proxy-ratelimit` budget
window that decrements on every call, so the pacing logic from
[Budget & pacing](https://docs.donkey-kit.dev/budget.md) runs end-to-end against the simulator. For a real
windowed counter that resets and emits the `429` on exhaustion, use the
[`budget` scenario](#scenario-scripting).
## Scenario scripting
The sentinel lets *you* decide which call fails. A **scenario** lets the
simulator decide, on a rule you set once at boot, so failures fire on their own
and deterministically across a whole run. Pass `--scenario` once per rule; the
grammar is `:`.
```bash
donkey mock --port 8080 \
--scenario pii_block:every=5 \
--scenario budget:limit=20000,window=60s \
--scenario injection:on-pattern="ignore previous"
```
When more than one is set, they are evaluated `injection` → `pii_block` →
`budget`, and the first that fires wins. The `donkey-sim/` sentinel
takes precedence over all of them — it is an explicit "force this exact shape"
override.
The sentinel is for one-off failures; a scenario runs across a whole session.
Use both together to cover specific calls and background failure rates.
### `pii_block:every=N`
Serves the captured `pii-detected` **403** on every Nth `POST /responses`; the
other calls get the happy path.
```bash
donkey mock --scenario pii_block:every=5
```
Every fifth call raises a typed [`PIIDetected`](https://docs.donkey-kit.dev/errors.md) in your agent, so your
masking logic runs in your own terminal before the pull request is opened.
### `budget:limit=,window=`
Runs a **real, wall-clock-windowed token counter**. Shrink an hour-long window
to a minute and your [pacing and resume logic](https://docs.donkey-kit.dev/budget.md) is exercised in ninety
seconds.
```bash
donkey mock --scenario budget:limit=20000,window=60s
```
Each call deducts the happy-path completion's reported token usage (override
with `cost=`). While budget remains, the `200` carries the
`x-llm-proxy-ratelimit` prose header, decreasing as you spend. Once the window
is spent, calls receive the captured `token-rate-limit` **429**, with
`x-token-remaining` and `x-token-reset` recomputed from the counter and the
milliseconds left in the window, until the window rolls over. `duration`
accepts `ms`, `s`, or `m` (a bare number is seconds).
The numeric `x-token-*` headers appear only on the `429`, and the prose
`x-llm-proxy-ratelimit` header only on the `200` — matching the real gateway.
The simulator never emits a header shape the gateway does not.
### `injection:on-pattern=`
Serves the `injection-protection` **400** on any request whose `input`,
`messages`, or `instructions` text contains the substring (case-insensitive).
```bash
donkey mock --scenario injection:on-pattern="ignore previous"
```
The request surfaces as a typed [`PromptInjectionBlocked`](https://docs.donkey-kit.dev/errors.md) (via the
`x-injection-protection: blocked` discriminator), so the refusal branch of a
guardrailed bot is exercised before it faces a real attack.
## How the simulator works
**Every simulator response carries `x-donkey-simulator: true`** — including
framework-generated `405`/`500` responses — so a simulated response is never
confused with a real gateway in a log, a trace, or a screenshot.
- **It replays; it does not evaluate policy.** The simulator tests how your
agent handles a refusal, never *which* prompts get refused. You choose the
refusal (the `donkey-sim/` sentinel or a `--scenario` rule); the
simulator never inspects a prompt and decides it violates a policy, and it
ignores authentication. Testing whether a prompt would be blocked by your
deployed policy configuration requires the real gateway.
- **It works with any client.** A stock, non-SDK client — plain `httpx` or
`openai.OpenAI(base_url="http://localhost:8080", …)` — receives the
byte-identical rejection bodies and exact discriminator headers.
- **Its fixtures are the SDK's test fixtures.** The simulator serves the same
files the error classifier is tested against, so the simulator and the
[error taxonomy](https://docs.donkey-kit.dev/errors.md) cannot drift apart. The wheel ships a sha256
integrity manifest of every fixture, so a changed byte fails loudly rather
than silently altering what the simulator replays.
- **It is pure Python — not Omni/Flex Gateway running in Local Mode.**
`pip install "donkey-kit[local]"` adds Starlette and Uvicorn; no Docker
required. DDK does not support Local Mode as a deployment target (see the
[unsupported boundary](https://docs.donkey-kit.dev/reference/unsupported-boundary.md)); real-gateway
behavior is exercised against a hosted (Connected Mode) proxy instead.
## Related
- [Testing & conformance](https://docs.donkey-kit.dev/testing.md) — `simulate()` for in-process unit tests,
and the `gateway` pytest fixture that runs this simulator on an ephemeral port.
- [Error taxonomy](https://docs.donkey-kit.dev/errors.md) — the typed exceptions each shape maps to.
- [CLI](https://docs.donkey-kit.dev/cli.md) — the full `donkey` command reference.
---
Source: https://docs.donkey-kit.dev/testing.md
# Testing & conformance
Live
Three test-time tools, all serving the same captured gateway rejections:
- **`simulate()`** — make the next N calls through a `Donkey` return a chosen
refusal, in process, with no server.
- **The conformance plugin** — a pytest suite you point at *your own* agent to
find governance bugs you did not know you had.
- **The `gateway` fixture** — a real simulator on an ephemeral port, for
subjects that do not import the SDK, with a spy on what it received.
```bash
pip install "donkey-kit[test]" # simulate() + conformance plugin
pip install "donkey-kit[test,local]" # add the gateway fixture
```
## `simulate()` — in-process, no server
A context manager that makes the next N calls through a `Donkey` return a
chosen refusal. It injects the **same captured rejection fixture** the error
classifier and the [local simulator](https://docs.donkey-kit.dev/simulator.md) use, so your agent sees exactly
that typed refusal — not a hand-rolled stand-in. No network, no server, fast
enough for unit tests:
```python
from donkey_kit import Donkey, PIIDetected
async def test_agent_masks_pii():
donkey = Donkey.from_env()
with donkey.simulate(PIIDetected):
result = await triage_agent.run(ticket_with_card_number)
assert "****" in result.draft_reply
```
- `times=N` (default `1`) sets how many calls are refused before requests
proceed normally. Retries of a single logical call count once.
- The previous transport is restored on exit, even if the block raises, and
nested `simulate()` blocks compose.
- Every injected response carries `x-donkey-simulator: true`, so a simulated
refusal is never mistaken for a real gateway response in a log or trace.
- Injected calls stay in-process even when `HTTP_PROXY`, `HTTPS_PROXY` or
`ALL_PROXY` is set. Calls that proceed after `times` still go through the
configured proxy.
The selectable refusals are the ones produced from a captured fixture:
`TokenBudgetExceeded`, `PIIDetected`, `PromptInjectionBlocked`,
`ContentSafetyBlocked`, `AgentKilled`, `UpstreamRequestError`, `ModelNotRoutable`,
`UpstreamModelError`, `AuthError`, and the generic `PolicyViolation` (the
content-moderation shape).
Asking for a refusal no gateway response produces — for example a client-side
`ConfigError` — raises a `ValueError`.
`simulate()` injects the fixture verbatim. For scripted behaviour such as
every-Nth-call PII blocks or a shrinking budget window, use the
[`gateway` fixture](#the-gateway-fixture) with
`set_scenarios(...)`, or [`donkey mock --scenario`](https://docs.donkey-kit.dev/simulator.md#scenario-scripting).
## The conformance plugin — run it against your agent
Installing `donkey-kit[test]` registers a pytest plugin with a conformance suite
you point at your own agent factory:
```bash
pytest --donkey-conformance --agent=my_app.agent:build
```
It answers questions you probably cannot currently answer about your own code:
- Does your agent **retry** a `TokenBudgetExceeded`? *It must not* — a policy
refusal is terminal, and retrying it burns budget to earn the same refusal.
- Does it swallow `PIIDetected` as a generic exception?
- Does it propagate the correlation ID into its own logs?
- Does it still work when budget headers are absent entirely?
Each scenario becomes one pytest item, and the run prints a table of
scenario → pass / fail / **exempt**. A failure reads as a finding about your
agent — for example, *"your agent retried a budget refusal 3 times"* — found in
CI, before production.
The suite grades what it sees on the `Donkey`'s transport, so your agent's model
calls must go through the `donkey` it is given. A scenario in which no call
reaches that transport **fails** with *"no model call went through the Donkey
client"*, never passes by default. That covers an agent that makes no call, one
that builds its own `httpx` or OpenAI client, and an adapter that sends calls
outside the shared transport (CrewAI, LlamaIndex, Microsoft Agent Framework,
and ADK `model()`). The finding names the adapter. If the gap is expected,
record an asserted exemption for it. The run never reaches the network: the
harness replaces the configured proxy URL and every credential with
placeholders, and it refuses any request that would leave the process.
The factory is called once per scenario and receives the `donkey` fixture if it
declares one. With `--donkey-conformance`, the plugin runs the suite
exclusively in place of normal test collection; without the flag it is inert.
You can also run it through the CLI with `donkey test --agent my_app.agent:build`
(see [CLI](https://docs.donkey-kit.dev/cli.md#donkey-test)).
### Exemptions
**Exemptions are asserted, never silently skipped.** If your agent
legitimately cannot satisfy a scenario, record it in a `KNOWN_LIMITATIONS`
mapping of `{scenario: reason}` as an explicit, reviewable claim.
By default the plugin reads a `KNOWN_LIMITATIONS` attribute from the `--agent`
module; point elsewhere with `--donkey-known-limitations=module:NAME`. The
mapping is validated at collection time, so an unknown scenario key or an empty
reason fails the run before any scenario executes.
## The `gateway` fixture
`simulate()` and the conformance plugin run your subject **in process**. Some
subjects cannot be: a containerised agent, a Node service, an A2A client, a
manual `curl`. Those need a real listener on a real port — and often you need
to know *what the gateway actually received*. "Did my agent stop after the
refusal, or retry four more times?" is answerable only from the gateway's side.
The `gateway` fixture boots the [local simulator](https://docs.donkey-kit.dev/simulator.md) on an ephemeral
port, hands your test its `.url`, and records every request:
```python
async def test_agent_stops_after_a_refusal(gateway):
gateway.set_scenarios("pii_block:every=1")
app = deploy(env={"DONKEY_LLM_PROXY_URL": gateway.url})
await app.run(ticket_with_card_number)
# A policy refusal is terminal — the agent must not retry it.
assert gateway.requests_received == 1
```
- `gateway.url` is a real `http://127.0.0.1:` any process can point
`DONKEY_LLM_PROXY_URL` at.
- The port is bound to `0`, so parallel `pytest -n` runs never collide.
- `gateway.requests_received` counts requests; `gateway.requests` exposes each
one's method, path, and headers, with `client_secret` **redacted**.
- `gateway.set_scenarios(...)` arms [scenario scripting](https://docs.donkey-kit.dev/simulator.md#scenario-scripting)
(`pii_block`, `injection`, `budget`) **per test**.
- Every response carries `x-donkey-simulator: true`.
- The server is torn down when the test exits, including on failure.
The fixture needs the `[local]` extra (Starlette + Uvicorn); requesting it
without that extra raises an `ImportError` naming the exact `pip install`.
## Choosing a tool
| | [`donkey mock`](https://docs.donkey-kit.dev/simulator.md) | `gateway` fixture | `simulate()` |
|---|---|---|---|
| Shape | A **server** you run | A **server** a test starts | **In-process** context manager |
| Port | You pick it (`--port`) | Ephemeral (`0`), on `.url` | None |
| Use it for | Manual dev, demos, any client | Testing an out-of-process subject | Fast unit tests |
| Needs a network | Yes (localhost) | Yes (localhost) | No |
| Works with a stock OpenAI client | Yes | Yes | No — it hooks the `Donkey` transport |
| Asserts on what the gateway received | No | Yes (`requests_received`) | No |
All three read the **same rejection fixtures**, so they cannot drift apart.
---
Source: https://docs.donkey-kit.dev/cli.md
# CLI & decorators
Live
Two on-ramps to the SDK: decorators that govern a function in one line, and a
four-command CLI for setup, diagnosis, local simulation, and conformance
testing.
## Decorators
### `@donkey.governed`
Runs a function inside a [`donkey.run()`](https://docs.donkey-kit.dev/telemetry.md) scope:
```python
@donkey.governed(team="support")
async def handle_ticket(ticket): ...
```
One decorator gives the function a run/correlation ID, cost tags, an OTel span,
and typed refusals — the four things you would otherwise set up per call site.
- Wraps **both sync and async** callables.
- Works bare (`@donkey.governed`) or parametrised
(`@donkey.governed(team=…, project=…, env=…, enduser_id=…)`).
- Takes **no `id=`**: each call opens its own run, so unrelated calls are never
collapsed into one correlation. When you need a specific id, use
`donkey.run(id=…)` directly.
### `@donkey.tool`
```python
@donkey.tool
async def lookup_crm(customer_id: str) -> dict:
"""Look up a customer record by id."""
...
```
Marks a function as a governed tool **without changing how it's called**. It
returns the same function with a `__donkey_tool__` marker and records a
`ToolSpec` (name, qualname, signature, docstring, `is_async`) in a
process-global registry you read with `registered_tools()`. Both `ToolSpec` and
`registered_tools` are exported from `donkey_kit`.
A tool with **no docstring is rejected at decoration time** (`ValueError`) — an
undescribed tool is useless to a model and to a registry.
The same markers are what the planned [scanner](https://docs.donkey-kit.dev/publishing.md) and
[A2A](https://docs.donkey-kit.dev/a2a.md) agent-card generator read, so marking tools now carries forward.
## The CLI
```bash
pip install "donkey-kit[cli]"
```
```bash
donkey init # writes a commented .donkey-kit.toml, names every missing env var at once
donkey doctor # checks creds, reaches the gateway, reports budget state
donkey mock # the local simulator — --scenario scripts failures
donkey test # a thin front end to pytest --donkey-conformance
```
### Global flags
Three global flags precede the subcommand:
```bash
donkey --config ./cfg.toml init # write the generated config file to a non-default path
donkey --env Sandbox init # write this Anypoint environment into the generated file
donkey --json init # machine-readable output where a command supports it
```
`--config` and `--env` apply to `init` only. Every other command reads its
configuration from environment variables and the working directory's config
files, so it rejects them with exit `2` rather than run against a configuration
you didn't ask for. To point `doctor` at another environment, set
`ANYPOINT_ENV`, or run it from the directory that holds the config file.
### Exit codes
Every command exits non-zero on failure, so any of them drops into CI as a
preflight. A command that needs an optional extra (`donkey mock` → `[local]`,
`donkey test` → `[test]`, `donkey doctor` → `[llm]`) prints the exact
`pip install` line and exits `1` — never a stack trace.
### `donkey init`
Resolves your current configuration (env vars → config files → defaults; see
[Precedence](https://docs.donkey-kit.dev/reference/configuration.md#precedence)) and writes a **commented** `.donkey-kit.toml` with the values it
found.
- **Names every missing required field at once** — control plane *and* LLM
proxy — using the same validation the SDK runs at call time, so `init` and a
real request never disagree about what is required.
- **Never writes a secret.** `client_secret`, `llm_proxy_client_secret`, and
`llm_proxy_key` are emitted as commented pointers, not values. Put them in a
gitignored `.donkey-kit.local.toml` next to the file, or in environment
variables. A URL in `.donkey-kit.toml`, including a loopback one, only
receives credentials from those two files unless you opt in; see
[Which credentials a URL receives](https://docs.donkey-kit.dev/reference/configuration.md#which-credentials-a-url-receives).
The file's comments also say that URLs must use `https://`, that `localhost`
is accepted over plain `http://`, and that `DONKEY_ALLOW_HTTP=1` in the
environment (not in the file) allows plain `http://` to other hosts.
- **Idempotent.** An existing file is left untouched unless you pass `--force`.
```bash
donkey init --force
donkey --json init # {"path": "...", "written": true, "missing": [...]}
```
### `donkey doctor`
A governed call can fail for several reasons that look identical from the
outside. `doctor` makes one real governed call and tells them apart:
- **Wrong credentials** — the `client_id`/`client_secret` pair is rejected.
- **Wrong URL** — the credentials are fine but the base URL is not a proxy
instance. (A trailing `/v1` lands here; the governed proxy has no `/v1`
segment.)
- **Credentials fine, model not in the allow-list** — nothing is misconfigured;
your platform team has not granted that model.
Each verdict prints the same remediation string the matching typed exception
carries, so the fix is in the output rather than in a runbook. The budget line
states how old the reading is, since the proxy has no budget endpoint.
Before the call, `doctor` prints each endpoint's host and where it came from:
`env`, `project file`, `local overlay`, `user file` or `default`. If the LLM
proxy URL comes from the working directory's config files but its credentials
don't (a loopback URL such as the simulator's included), the `config` line
fails with the
[remediation](https://docs.donkey-kit.dev/reference/configuration.md#which-credentials-a-url-receives) and no
request is sent. A control-plane endpoint with the same problem shows its
remediation on the `control plane` line without failing the report, since
`doctor` only calls the LLM proxy. In `jwt` mode, an LLM proxy URL from those
files always fails the `config` line, because the JWT never comes from a file:
set the URL in the environment or opt in.
With `DONKEY_ALLOW_HTTP` on in the environment, a `plain http` line says so and
quotes the value as set, for example `DONKEY_ALLOW_HTTP=true` (in `--json`
output, the entry with `"name": "plain http"`):
```text
[i] plain http allowed to non-loopback hosts (DONKEY_ALLOW_HTTP=1 in env)
```
```bash
donkey doctor
```
```text
[ok] config env (3 fields)
[i] llm endpoint (env)
[i] control plane anypoint.mulesoft.com (default)
[ok] gateway reachable, responded
[ok] credentials client_id accepted
[ok] model accepted by the proxy
[i] budget 99,000 / 100,000 remaining, resets in 59s, observed 0s ago
```
```bash
donkey doctor --model gpt-4o # model to test against the allow-list (default gpt-4o)
donkey doctor --json # machine-readable checks
```
### `donkey mock`
Runs the [local simulator](https://docs.donkey-kit.dev/simulator.md), which replays captured gateway
rejections. Needs the `[local]` extra.
```bash
donkey mock --port 8080 --host 127.0.0.1 \
--scenario pii_block:every=5 \
--scenario budget:limit=20000,window=60s
```
| Flag | Default | Meaning |
|---|---|---|
| `--port` | `8080` | TCP port to bind. |
| `--host` | `127.0.0.1` | Host/interface to bind. |
| `--scenario` | none | Fault-injection rule, repeatable. See [Scenario scripting](https://docs.donkey-kit.dev/simulator.md#scenario-scripting). |
The port isn't read from an environment variable or `.donkey-kit.toml`; these
flags (or the `serve()` keyword arguments) are the only way to change it. See
[Choosing a port or host](https://docs.donkey-kit.dev/simulator.md#choosing-a-port-or-host) for the
programmatic entry point and the `0.0.0.0` caveat.
An invalid `--scenario` exits with code `2`.
### `donkey test`
A thin front end to `pytest --donkey-conformance` — it does not re-implement the
runner. Point it at your agent factory and pass any trailing pytest arguments
straight through; pytest's exit code becomes `donkey test`'s own. Needs the
`[test]` extra.
```bash
donkey test --agent my.pkg:make_agent -k governance -x
```
See [Testing & conformance](https://docs.donkey-kit.dev/testing.md) for what the suite checks.
## Planned commands Roadmap
These commands are part of the [Roadmap](https://docs.donkey-kit.dev/roadmap.md) and are not available in the
CLI yet:
- `donkey scan` and `donkey publish` — derive a manifest and agent card from
your code and register them with Exchange. See [Scan & publish](https://docs.donkey-kit.dev/publishing.md).
- `donkey serve`, `donkey expose`, and `donkey dev` — serve your agent over
A2A and expose it through the gateway. See [A2A agents](https://docs.donkey-kit.dev/a2a.md).
Run `donkey --help` to see the commands available in your installed version.
---
Source: https://docs.donkey-kit.dev/use-with-your-agent.md
# Use these docs with your agent
Live
You probably ask your coding assistant before you open a docs site. These docs
are published in the [llms.txt convention](https://llmstxt.org) so an assistant
can read them directly and write **correct, governed** code — the base URL with
no `/v1`, the `client_id` / `client_secret` header pair, the typed
refusal taxonomy — instead of guessing.
## What's published
Three machine-readable artifacts are generated from these pages on every docs
build, so they never drift from what you see here:
| Artifact | What it is | URL |
| --- | --- | --- |
| `llms.txt` | A curated **index** — one link per page, grouped by section. Best when your tool ingests a doc index and fetches pages on demand. | [`/llms.txt`](https://docs.donkey-kit.dev/llms.txt) |
| `llms-full.txt` | **Every page inlined** into one file. Best for pasting straight into an assistant — no fetching required. | [`/llms-full.txt`](https://docs.donkey-kit.dev/llms-full.txt) |
| Per-page `.md` | The raw markdown for any page, served next to its HTML. Append `.md` to any page URL. | e.g. [`/quickstart.md`](https://docs.donkey-kit.dev/quickstart.md) |
The site is served under a project sub-path today
(`/donkey-development-kit`), so the files live at
`https://docs.donkey-kit.dev/llms.txt`
rather than a bare domain-root `/llms.txt`. Use the full URLs above.
## Point your assistant at them
Ask Claude Code to read the full docs, then build:
```text
Read https://docs.donkey-kit.dev/llms-full.txt,
then write a governed LangGraph model call using the donkey-kit SDK.
```
In Cursor, add the docs as a source (**Settings → Features → Docs → Add**) with
the index URL, then `@Docs` it in chat:
```text
https://docs.donkey-kit.dev/llms.txt
```
For any assistant, paste the contents of `llms-full.txt` into the conversation,
then ask your question. Because every page is inlined, the assistant has the
full context without following links:
```text
Now write a governed OpenAI Agents SDK setup that handles a TokenBudgetExceeded
refusal.
```
If your tool follows a doc index, give it `llms.txt`. Each entry links to the
page's `.md`, so the tool fetches only the pages it needs:
```text
https://docs.donkey-kit.dev/llms.txt
```
These docs cover both live capabilities and ones on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md).
Roadmap pages show the planned API, which is not callable yet. When you ask an
assistant to write code, tell it to use only capabilities marked **Live**, and
review generated code before running it against a real gateway.
---
Source: https://docs.donkey-kit.dev/tool-access.md
# Tool access
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
Governed tool access lets an agent discover the MCP tools your organisation has
published and governed, filter them down to what it actually needs, and bind
them into any of the eight supported frameworks as that framework's **native
tool objects** — the same "no wrapper" approach as [model access](https://docs.donkey-kit.dev/frameworks.md).
## Two lines from catalog to agent
```python
tools = await donkey.tools.discover(domain="hr", tags=["approved"])
agent = create_react_agent(donkey.langgraph.chat_model("gpt-4o"), tools.langgraph())
```
Everything else in this section — search filters, session management,
per-framework binding, pinning, and A2A tool handles — makes those two lines
hold up in production.
## Discover and filter
`donkey.tools.discover(...)` is also the filter entry point. One call narrows
the catalog by governance, domain and tags, so an agent binds only the tools it
needs:
```python
tools = await donkey.tools.discover(
domain="hr",
tags=["approved"],
governed=True, # default criteria, or a GovernanceCriteria
)
```
Name search, asset type and environment filters live one layer down, on
`donkey.registry.search()`.
"Governed" is a computed, environment-scoped predicate rather than a flag in
Exchange — see [Discovery, search & filter](https://docs.donkey-kit.dev/tool-access/discovery.md).
## In this section
Narrow the catalog by governance, domain and tags, or search it by name and asset type.
Turn a `ToolSet` into each framework's own native tool objects.
Pin resolved versions and digests so a run is reproducible.
Treat a governed agent-to-agent endpoint as another bindable tool.
---
Source: https://docs.donkey-kit.dev/tool-access/discovery.md
# Discovery, search & filter
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
`donkey.tools.discover(...)` is the one entry point for finding governed tools
to bind. It narrows the catalog by **governance**, domain and tags, so an agent
binds only the tools it needs rather than the entire catalog. It returns a
[`ToolSet`](https://docs.donkey-kit.dev/tool-access/binding.md) whose per-framework methods hand back native
tool objects. Name search, asset type and environment are filters on the
lower-level [`donkey.registry.search()`](#registry-search).
## The two most common calls
```python
# 1. Everything governed in a domain:
tools = await donkey.tools.discover(domain="hr", governed=True)
# 2. Governed tools in a domain that carry a tag:
tools = await donkey.tools.discover(domain="hr", tags=["approved"], governed=True)
```
## Filter reference
Every argument is keyword-only and optional, and filters combine with `AND`
semantics:
```python
tools = await donkey.tools.discover(
domain="hr", # catalog domain
tags=["approved"], # all tags must be present
governed=True, # True = default criteria; or an explicit GovernanceCriteria
locked=False, # True = resolve only what donkey.lock pins
)
```
| Argument | Type | Meaning |
|---|---|---|
| `domain` | `str \| None` | Catalog domain. |
| `tags` | `list[str] \| None` | All listed tags must be present. |
| `governed` | `bool \| GovernanceCriteria \| None` | `True` applies the default criteria; pass a `GovernanceCriteria` (e.g. `STRICT`) for explicit rules; `None` = unfiltered (the default). |
| `governance` | `Any \| None` | Reserved in the signature; no behaviour is defined for it yet. |
| `locked` | `bool` | `True` resolves only the versions pinned in `donkey.lock` (see [Pinning & lockfile](https://docs.donkey-kit.dev/tool-access/lockfile.md)). Default `False`. |
## Registry search
`discover(...)` is the high-level facade over the registry. When you need name
search, asset types, an explicit environment or a result limit, call
`donkey.registry.search()` directly. It returns `AssetRef` handles rather than a
bindable `ToolSet`:
```python
refs = await donkey.registry.search(
query="*accounts*", # glob over asset name + description; None = no text filter
asset_types=["mcp"], # restrict to MCP servers, agents, etc.
tags=["approved"],
domain="hr",
environment="Production", # environment-scoped governance
governed=True,
limit=50, # max results (default 50)
)
```
## "Governed" is a computed predicate
Publication to Exchange says nothing about whether an asset is fronted by a
gateway, has policies applied, or passes the org's rulesets — there is no single
boolean to query. "Governed" is **computed** by joining state across systems,
and it is **environment-scoped**: an asset governed in Production may be
ungoverned in Sandbox. `GovernanceCriteria` makes every condition explicit:
```python
from donkey_kit.registry.governance import GovernanceCriteria, STRICT
@dataclass(frozen=True)
class GovernanceCriteria:
require_api_instance: bool = True # an API Manager instance exists in this env
require_deployed: bool = True # deployed to a gateway, not just configured
require_any_policy: bool = True # at least one policy applied
required_policies: list[str] = ... # e.g. ["client-id-enforcement"]
forbidden_policies: list[str] = ...
require_governance_pass: bool = False # passes org rulesets with no `error` findings
require_gateways: list[str] = ... # only assets behind these named gateways
require_tags: list[str] = ...
require_lifecycle: list[str] = ... # e.g. ["published", "approved"]
allow_unknown: bool = False # if a check can't be evaluated, does it pass?
# A ready-made strict preset:
STRICT = GovernanceCriteria(
require_governance_pass=True,
required_policies=["client-id-enforcement"],
allow_unknown=False,
)
```
Pass it straight through:
```python
tools = await donkey.tools.discover(domain="hr", governed=STRICT)
```
`allow_unknown` matters more than it looks. If the platform doesn't expose,
say, ruleset results, then `require_governance_pass=True` with
`allow_unknown=False` filters the whole catalog to zero — so every
filtered-out asset carries a **reason**, surfaced by `explain()`.
## `explain()` — why a tool was included or excluded
Without it, `governed=True` returning an empty list is indistinguishable
from a broken credential. `explain()` reports every check:
```python
report = await donkey.registry.explain(ref, criteria=STRICT)
# GovernanceReport(governed=False, checks=[
# Check("api_instance_exists", True, "instance 19283 in Sandbox"),
# Check("deployed", True, "gateway managed-omni-eu-1"),
# Check("required_policies", False, "missing: client-id-enforcement"),
# Check("governance_pass", None, "UNKNOWN: rulesets API returned 403"),
# ])
```
Each check is `True` (passed), `False` (failed, with the reason), or `None`
(couldn't be evaluated — resolved via `allow_unknown`). The empty-result warning
message points you to `explain()`.
## Defaults and performance
- **Unfiltered by default.** `governed` defaults to `None`, and a startup
log line states that discovery is unfiltered. Changing the default to `True`
is reserved for a future major version.
- **Warm the index.** The governance join is a bulk operation, not one API call
per asset. A long-running agent can build the index at startup with
`donkey.registry.warm(environment=...)` so the first discovery is fast.
## Related
- [Framework binding](https://docs.donkey-kit.dev/tool-access/binding.md) — turn a `ToolSet` into native tools.
- [Pinning & lockfile](https://docs.donkey-kit.dev/tool-access/lockfile.md) — pin resolved versions for production.
---
Source: https://docs.donkey-kit.dev/tool-access/binding.md
# Framework binding
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
Once [discovery](https://docs.donkey-kit.dev/tool-access/discovery.md) hands you a `ToolSet`, binding turns
governed MCP servers into a framework's own tool objects — nothing wrapped,
nothing re-implemented.
## MCP session management
MCP servers created by MCP Bridge are gateway endpoints speaking streamable
HTTP, protected by gateway policies. The SDK's session layer handles four
things for you:
- **Auth.** Client-credentials OAuth is the machine-to-machine case. Strands'
`MCPClient` already builds streamable HTTP with a `client_credentials` grant
internally; every other framework needs headers supplied explicitly.
`McpServerHandle.auth_headers()` returns a ready-to-use header dict, refreshed
automatically on a `401`.
- **Connection lifecycle.** MCP clients are stateful, and several frameworks
connect lazily. `donkey.tools.discover()` never opens a connection — it
returns handles, and the connection opens on first tool use.
- **Multi-server aggregation.** `ToolSet` wraps N `McpServerHandle`s. When two
servers expose a tool with the same name, the collision is resolved by
prefixing the server's short name — for example `hr__get_employee` — and the
mapping is available on `ToolSet.name_map`, so you can see exactly why the
model called that name.
- **Filtering.** Enterprise MCP servers can expose dozens of tools. Handing 60
tool descriptors to a model degrades it and inflates token cost, so filter
before you bind:
```python
tools = await donkey.tools.discover(domain="hr")
filtered = tools.filter(allow=["get_employee", "search_employees"])
# or:
filtered = tools.filter(deny=["delete_*"])
# or a predicate over the tool descriptor:
filtered = tools.filter(predicate=lambda t: t.name.startswith("get_"))
```
The SDK logs the descriptor token count for a `ToolSet` at debug level, so
you can see the cost of skipping `filter()` before a model does.
## Per-framework binding
`ToolSet` exposes one method per installed integration, each returning the
framework's **native** tool type:
```python
ts = await donkey.tools.discover(domain="hr")
ts.langgraph() # -> list[BaseTool]
ts.adk() # -> list[McpToolset]
ts.strands() # -> list[MCPClient]
ts.llamaindex() # -> list[FunctionTool]
# etc.
```
Note the shape difference: ADK and Strands take a toolset/provider object,
while LangGraph and LlamaIndex take a flat tool list. Each method matches its
framework's own idiom rather than forcing a uniform return type, and its
docstring calls out the difference.
| Framework | Binding |
|---|---|
| LangGraph | `langchain_mcp_adapters.client.MultiServerMCPClient({...}).get_tools()` — the SDK builds the connection dict from your handles, transport `"streamable_http"`, headers injected. |
| Google ADK | `McpToolset(connection_params=StreamableHTTPConnectionParams(url=..., headers=...), tool_filter=[...])`, passed straight into `LlmAgent(tools=[...])`. |
| MS Agent Framework | The framework's MCP client/tool class for streamable HTTP. |
| OpenAI Agents SDK | `agents.mcp.MCPServerStreamableHttp(params={"url": ..., "headers": ...})`, passed into `Agent(mcp_servers=[...])`. |
| Anthropic SDK | The `anthropic` SDK has no native client-side MCP binding; the SDK binds via the MCP Python SDK's streamable-HTTP client and passes the resulting tool schemas to `messages.create(tools=...)`. |
| CrewAI | The framework's MCP adapter for streamable-HTTP servers, yielding native `crewai` tool objects for a `Crew`/`Agent`. |
| LlamaIndex | `llama_index.tools.mcp.BasicMCPClient` + `McpToolSpec(...).to_tool_list_async()`. |
| Strands | `MCPClient(lambda: streamablehttp_client(url, headers=...))` — implements `ToolProvider`, so it can be passed directly into `Agent(tools=[...])` with automatic lifecycle management. |
A binding failure or a `401` surfaces as a typed exception from the
[error taxonomy](https://docs.donkey-kit.dev/errors.md).
## Related
- [Discovery, search & filter](https://docs.donkey-kit.dev/tool-access/discovery.md) — produce the `ToolSet`.
- [A2A agent tools](https://docs.donkey-kit.dev/tool-access/a2a.md) — bind a remote agent the same way.
- [Frameworks](https://docs.donkey-kit.dev/frameworks.md) — governed model access per framework.
---
Source: https://docs.donkey-kit.dev/tool-access/lockfile.md
# Pinning & lockfile
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
Governed tool catalogs change under you. A platform team edits an MCP server's
tool schema, bumps a policy, or republishes an asset — and if your agent
resolves `version="latest"` at startup, that change silently alters agent
behaviour in production. Pinning and the lockfile keep what your agent binds
under version control.
## Pin by default
`donkey.tools.discover()` requires an explicit version on every asset reference
by default. `version="latest"` is allowed, but logs a warning — it is an opt-in
escape hatch, not the default path:
```python
# explicit — no warning
tools = await donkey.tools.discover(domain="hr", version="1.2.0")
# opt-in — logs a warning every time it resolves
tools = await donkey.tools.discover(domain="hr", version="latest")
```
## The lockfile
`donkey.tools.lock()` resolves the current discovery call and writes a
`donkey.lock` file recording the resolved versions and content digests of every
asset it touched:
```bash
python -c "import asyncio; from donkey_kit import Donkey; asyncio.run(Donkey.from_env().tools.lock())"
```
```yaml
# donkey.lock (illustrative)
lockedAt: 2026-08-28T00:00:00Z
assets:
- ref: com.acme/hr-tools-mcp/1.2.0
digest: sha256:...
- ref: com.acme/vendor-shipment-mcp/1.0.0
digest: sha256:...
```
Once a `donkey.lock` exists, pass `locked=True` and discovery refuses to resolve
anything not already in the lockfile:
```python
tools = await donkey.tools.discover(domain="hr", locked=True)
# raises if discovery would resolve an asset/version not in donkey.lock
```
Commit `donkey.lock` alongside your agent code and treat it as a required step
before deploying. A version bump then becomes a reviewable diff in a pull
request instead of a runtime surprise.
## Registry caching
Registry lookups (`ExchangeRegistry.search()`, `resolve_mcp()`,
`resolve_agent()` — see [Discovery, search & filter](https://docs.donkey-kit.dev/tool-access/discovery.md))
are cached with a configurable TTL (default 300 seconds):
```python
from donkey_kit import Donkey, DonkeyConfig
donkey = Donkey(DonkeyConfig.from_env().with_overrides(registry_cache_ttl_s=60))
```
- `donkey.registry.refresh()` invalidates the cache and re-fetches on next use —
call it after a platform team publishes a change you need to see immediately.
- Set `DONKEY_REGISTRY_CACHE_TTL_S` to change the TTL from the environment.
- Set `DONKEY_NO_CACHE=1` to bypass the cache entirely, for debugging a
discovery result that looks stale.
Caching and pinning solve different problems: caching controls how often you
re-ask the registry the same question; pinning controls whether an answer is
allowed to change under a running agent at all.
---
Source: https://docs.donkey-kit.dev/tool-access/a2a.md
# A2A agent tools
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
Agent Broker is an A2A server, and A2A-compliant agents in the registry can be
consumed directly from Python. `AgentHandle.as_tool()` wraps a remote A2A agent
as a callable tool in your framework — the same "resolve, then bind" shape as
MCP tools, for a remote agent instead of a remote tool server.
```python
handle = await donkey.registry.resolve_agent("com.acme/claims-triage-agent/1.0.0")
tool = handle.as_tool() # callable in your framework's native tool shape
agent = create_react_agent(donkey.langgraph.chat_model("gpt-4o"), [tool])
```
This lets a Python agent **delegate** to an Agent Broker agent without you
learning the A2A protocol. Under the hood, `as_tool()` uses the official
`a2a-sdk` for protocol handling.
## Install
A2A support ships behind its own extra, so installs that only need MCP tools
carry no extra dependency:
```bash
pip install "donkey-kit[a2a]"
```
## Where this fits
- `AgentHandle` comes from `ExchangeRegistry.resolve_agent()` — the same
registry surface that resolves MCP servers into `McpServerHandle`s. See
[Discovery, search & filter](https://docs.donkey-kit.dev/tool-access/discovery.md).
- `as_tool()` returns a framework-native callable, following the same
conventions as [Framework binding](https://docs.donkey-kit.dev/tool-access/binding.md).
- To make *your* agent callable by others over A2A, see [A2A agents](https://docs.donkey-kit.dev/a2a.md).
---
Source: https://docs.donkey-kit.dev/a2a.md
# A2A agents
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
Three steps take an agent from "runs on my laptop" to "callable by other agents,
through the gateway": **serve** it, **expose** it, and **develop** against a
gateway locally. The protocol is the official `a2a-sdk`, wrapped; the SDK adds
the governance around it.
## Why the agent still runs a listener
An A2A agent *is* a server: it serves an agent card at a well-known path and
answers JSON-RPC task calls. Something must accept the socket, so the SDK hides
that listener behind one line rather than removing it.
What it can remove is everything painful around the listener — TLS, auth, rate
limits, public exposure, and registration. Those are the gateway's job. Omni
Gateway is an Envoy-based data plane whose policies compile to WASM; it runs as
a separate process in front of your agent, not as a library inside it.
**Enforcement stays outside the agent's process.** If the enforcement point
lived inside the agent, the agent's own code could bypass it. The gateway's
value depends on being outside the thing it governs.
## `donkey serve` — the listener, in one line
```python
@donkey.agent(name="support-triage", skills=["triage", "draft-reply"])
async def handle(task: A2ATask) -> A2AResult:
return await graph.ainvoke(task.input)
donkey.serve(handle) # A2A server on 127.0.0.1:8000, card auto-generated
```
- **The card is generated from your code** — the same `@donkey.tool` /
`@donkey.agent` markers the [scanner](https://docs.donkey-kit.dev/publishing.md) reads — so there is no
second description of your agent to keep in sync.
- **It binds to localhost by default**, so your agent is never the public face;
the gateway is.
- **Inbound calls are governed.** Every call arriving over A2A gets the same
treatment as an outgoing one: correlation ID, cost tags, an OTel span, and
typed refusals when the agent's own downstream calls are blocked.
## `donkey expose` — the ingress, registered from code
```
$ donkey expose --env prod
✓ A2A proxy https://gw.acme.internal/agents/support-triage
✓ policies token-budget, pii-detection, trusted-agent-identity
✓ registry support-triage v1.4.0
```
The gateway does the work — provisioning the A2A proxy, attaching the policy
set, and registering the card. `donkey expose` turns that console session into
one command that runs in CI.
## `donkey dev` — a gateway in front of your laptop
`donkey dev` puts a gateway in front of `donkey serve` on your machine, so you
can call your agent over A2A through governance before deploying. Two modes are
planned:
- **A real gateway.** Where a self-managed Omni Gateway image is available for
local development, `donkey dev` runs it alongside `donkey serve`, so A2A calls
pass through real policies.
- **Simulated ingress.** The [local simulator](https://docs.donkey-kit.dev/simulator.md) gains an A2A ingress
mode: a local fake gateway in front of `donkey serve`, replaying the same
rejection fixtures. A plain A2A client sees byte-identical responses, and
every response carries `x-donkey-simulator: true`.
## Related
- [A2A agent tools](https://docs.donkey-kit.dev/tool-access/a2a.md) — call *other* A2A agents as tools.
- [Scan & publish](https://docs.donkey-kit.dev/publishing.md) — register your agent card with Agent Registry.
- [Local simulator](https://docs.donkey-kit.dev/simulator.md) — the fixture-replay server behind simulated ingress.
---
Source: https://docs.donkey-kit.dev/publishing.md
# Scan & publish
Roadmap
This capability is on the [Roadmap](https://docs.donkey-kit.dev/roadmap.md); the API shown here is the planned design.
`donkey scan` walks your repository, finds everything marked
[`@donkey.tool`](https://docs.donkey-kit.dev/cli.md#donkeytool), plus MCP server definitions and agent entry
points, and produces a manifest and A2A agent card. `donkey publish` registers
them with Exchange / Agent Registry. A GitHub Action runs both on every merge to
`main`.
**The registry becomes a consequence of the code, not a chore.** A support agent
with six tools registered by hand has a stale tool list within a week; with the
Action, every merge updates it.
**This complements MuleSoft's Agent Scanners.** Agent Scanners discover agents
from Agentforce, Bedrock, Vertex AI, and Copilot Studio at **runtime**. Scan &
publish is the **design-time / CI** complement: it registers agents built in
plain Python that no cloud scanner can see.
## Scope: code-first assets
Publication is for assets that **originate in your code**:
- an MCP server written in Python or TypeScript,
- an agent exposed over [A2A](https://docs.donkey-kit.dev/a2a.md),
- an agent exposed as a tool without an A2A surface.
It is not for assets the platform already owns. An MCP server created by MCP
Bridge from an existing API is already in Exchange; publishing a second,
code-derived descriptor would create two catalog entries for one capability.
### Collision check
Before publishing, the SDK searches Exchange for an existing asset with the
same endpoint, name, or derived tool signature. On a probable match it
**refuses** and prints the existing asset's coordinates. Override with an
explicit `--allow-duplicate`, which logs at `WARNING`.
## The `Publication` object
```python
from donkey_kit import Contact, Publication, PublicationAssetType
pub = Publication(
asset_type=PublicationAssetType.MCP_SERVER, # MCP_SERVER | A2A_AGENT | AGENT | API
group_id="${ANYPOINT_ORG_ID}",
asset_id="hr-tools-mcp",
version="1.3.0", # semver
name="HR Tools",
description="Employee lookup and leave-balance tools for HR agents.",
# Discovery metadata
tags=["hr", "internal", "agent-tool"],
categories={"Domain": "People", "Lifecycle": "Production"},
contact=Contact(team="People Platform", email="people-plat@acme.com"),
# Type-specific descriptor — exactly one, matching asset_type
descriptor="auto", # introspect the live server
# Documentation pages, published alongside the asset
docs=[
("home", "docs/exchange/overview.md"),
("getting-started", "docs/exchange/quickstart.md"),
],
# Where it actually lives — metadata only
endpoint="https://hr-tools.internal.acme.com/mcp",
)
```
`asset_type` determines which descriptor is required and how it is generated:
| `asset_type` | Descriptor | Generated from |
|---|---|---|
| `MCP_SERVER` | MCP tool manifest — server info, tool names, descriptions, JSON Schema inputs | live `tools/list` against the running server |
| `A2A_AGENT` | A2A Agent Card | declared skills, endpoint, auth schemes, input/output modes |
| `AGENT` | agent descriptor (no A2A surface) | framework introspection, best-effort |
| `API` | OpenAPI / AsyncAPI | user-supplied file; no generation |
## `descriptor="auto"` — deriving the spec from code
Hand-maintained catalog descriptors go stale within a sprint, so generation is
the core of this feature.
Every supported framework already derives JSON Schema from function signatures,
type hints, and docstrings — `@tool` in LangChain and Strands, `FunctionTool` in
ADK and LlamaIndex, `@mcp.tool()` in the MCP Python SDK, and the equivalent
conventions in CrewAI, the OpenAI Agents SDK, and the Anthropic SDK. The SDK
**asks the framework for the schema it already computed** rather than
re-deriving it, so the catalog documents exactly the schema the model sees.
### Derivation modes
```python
descriptor="auto" # object introspection — the default
descriptor="auto:live" # live protocol introspection — highest fidelity
descriptor="auto:static" # AST only — lowest fidelity, no code execution
descriptor="auto:check" # generate, diff against committed file, fail on mismatch
```
- **`auto:live`** starts the server, performs the MCP initialize handshake, and
calls `tools/list` (plus `resources/list` and `prompts/list`). It is exactly
what a client sees, but the server must actually run, with whatever
credentials and network that needs.
- **`auto`** (the default) imports your module, locates the tool and agent
objects, and reads their already-computed schemas. No server, no network,
works in CI. Importing user code executes it, so tool definitions must be
```toml
[publication.entrypoints]
"hr-tools-mcp" = "acme.hr.server:mcp" # module:attribute
"hr-agent" = "acme.hr.agent:build_agent" # a zero-arg factory also works
```
- **`auto:static`** parses decorators, signatures, and docstrings via AST
without executing anything — for environments where importing user code is
unacceptable. It cannot see tools registered in a loop, from config or a
database, behind a feature flag, attached dynamically at startup, or built
from imported/generated pydantic models, so it emits a completeness warning
whenever it hits a pattern it cannot resolve. It is never the default.
- **`auto:check`** generates the descriptor, diffs it against the committed
file, and fails on mismatch — useful as a CI gate.
`donkey publish --cross-check` runs `auto` and `auto:live` and diffs them. A
disagreement means either dynamic registration the object graph does not
reflect, or a broken framework adapter.
Type hints give the **shape**, not the **meaning** — `department: str` becomes
`{"type": "string"}` and says nothing about which departments are valid, or
when to use this tool over a similar one. `auto` fails publication on a
description that is tautological or missing, and `preview()` reports
description quality.
## Related
- [CLI & decorators](https://docs.donkey-kit.dev/cli.md) — mark tools with `@donkey.tool` today.
- [A2A agents](https://docs.donkey-kit.dev/a2a.md) — serve and expose the agent whose card you publish.
- [Tool access](https://docs.donkey-kit.dev/tool-access.md) — discover and bind published tools from Exchange.
---
Source: https://docs.donkey-kit.dev/scenarios.md
# Scenarios
The feature pages tell you what each piece *is*. These pages show three of them
working together on a real job, start to finish — the same three scenarios the
build guide uses to justify the skeleton (`BG §1.8`). Each one is runnable
today against the [local simulator](https://docs.donkey-kit.dev/simulator.md), with no Anypoint credentials
and no real gateway, so you can watch the governance branch execute in your own
terminal before it ever runs on a customer's data.
| Scenario | The job | What it exercises |
| --- | --- | --- |
| [Support triage](https://docs.donkey-kit.dev/scenarios/support-triage.md) | Draft replies to a queue of support tickets; one carries PII | Typed refusals ([`PIIDetected`](https://docs.donkey-kit.dev/errors.md)), [correlation IDs](https://docs.donkey-kit.dev/telemetry.md), [OTel spans](https://docs.donkey-kit.dev/telemetry.md) — the shipped LangGraph demo, end to end |
| [Nightly batch](https://docs.donkey-kit.dev/scenarios/nightly-batch.md) | Enrich 50,000 records overnight against a windowed budget, unattended | [Budget pacing](https://docs.donkey-kit.dev/budget.md) (`pace()` / `wait_for_reset()`) and resume, driven by the [simulator's `budget` scenario](https://docs.donkey-kit.dev/simulator.md#scenario-scripting) |
| [Internal copilot](https://docs.donkey-kit.dev/scenarios/internal-copilot.md) | An internal assistant whose output must clear a content-safety guardrail | [`ContentSafetyBlocked`](https://docs.donkey-kit.dev/errors.md) via the `donkey-sim/content-safety` sentinel, plus per-run correlation |
**"Scenario" means two different things in these docs — don't conflate them.**
These pages are the three *product* scenarios (a job you'd actually run). The
[simulator's `--scenario` flag](https://docs.donkey-kit.dev/simulator.md#scenario-scripting) names three
*failure-injection rules* (`pii_block`, `budget`, `injection`). The pages
below use those rules as the engine, but the job is the story.
## What's runnable, and what's honest about being blocked
Support triage runs **completely** today — it is the Phase-1 acceptance
artefact (`BG §1.8`), timed in CI so its first-run experience can't rot. The
nightly-batch and internal-copilot walkthroughs run their **governed-call and
refusal-handling** paths against the simulator now; where a step depends on a
surface that is still blocked on verification — MCP
tool access, agent identity, the kill switch — the page says so in the reader's
terms and shows the shape without inventing an endpoint (verification discipline). A "blocked"
note here means *known, deliberate, and tracked*, never *guessed*.
---
Source: https://docs.donkey-kit.dev/scenarios/support-triage.md
# Support triage
A support agent drafts a one-line reply to each ticket in a queue. Most are
routine; one contains a customer's SSN. On a bare `base_url` that PII-laden call
sails straight through to the model. Through a governed proxy it comes back as a
**typed refusal** — and the branch that masks and re-routes it is a branch you
can watch execute *before* it runs on real data.
This is the shipped Phase-1 acceptance demo (`BG §1.8`): `pip install` to a
drafted reply against the [local simulator](https://docs.donkey-kit.dev/simulator.md) only — no Anypoint
credentials, no real gateway — exercising four governance pieces at once.
## What it demonstrates
- **Runs with no gateway.** The demo boots the [local simulator](https://docs.donkey-kit.dev/simulator.md)
in-process on an ephemeral port and points the SDK at it; the simulator
ignores auth, so the credentials are throwaway placeholders.
- **The PII-masking branch actually executes.** The simulator runs the
`pii_block:every=5` [scenario](https://docs.donkey-kit.dev/simulator.md#scenario-scripting) — every fifth
`POST /responses` is served the captured `pii-detected` **403**. The fifth
ticket (the one with the SSN) is blocked, and the refusal surfaces out of the
LangGraph run as a typed [`PIIDetected`](https://docs.donkey-kit.dev/errors.md), not a framework-wrapped
generic error.
- **Correlation reaches every node for free.** Each ticket runs inside a
[`donkey.run(id=...)`](https://docs.donkey-kit.dev/telemetry.md) block; the run id shows up in a graph
node's own logs without being threaded through graph state, because LangGraph
runs nodes on context-copying `asyncio` tasks.
- **Every governed call emits an OTel span** — refusals included — exported to a
bundled local OTLP collector with [zero config](https://docs.donkey-kit.dev/telemetry.md): setting the
standard `OTEL_EXPORTER_OTLP_ENDPOINT` is all it takes.
## Run it
### Install the extras
```bash
pip install "donkey-kit[langgraph,local,otel]"
```
`langgraph` brings the framework and the adapter, `local` the simulator, `otel`
the OpenTelemetry export path.
### Run the demo
```bash
python -m examples.langgraph.main
```
No environment setup: the demo boots its own simulator, sets its own throwaway
credentials, and stands up its own local OTLP collector. You'll see one drafted
reply per routine ticket, a `BLOCKED by policy — PIIDetected` line for the
PII-laden one, the remaining budget, and the count of spans exported.
## The governed core
The whole demo is ordinary framework code; the SDK touches it in exactly two
places. First, the model is built off a shared `Donkey` so it rides that
instance's governed transport:
```python
model = donkey.langgraph.chat_model("gpt-4o")
```
Second, the node that calls the model wraps the call in `typed_refusals()`, so a
proxy rejection comes back as a `DonkeyError` subclass instead of a
framework-wrapped generic error:
```python
from donkey_kit.integrations.langgraph import typed_refusals
async def _call_model(state):
with typed_refusals():
reply = await model.ainvoke(state["messages"])
return {"messages": [reply]}
```
The caller binds a run id per ticket and catches the typed refusal:
```python
async with donkey.run(id=f"ticket-{i}"):
try:
result = await agent.run(f"Draft a one-line support reply to: {ticket}")
except PIIDetected as refusal:
print(f" ticket {i}: BLOCKED by policy — PIIDetected ({refusal.policy})")
```
**Honest note on the block.** The simulator triggers on the request *count*
(`every=5`), not by scanning content — it replays a captured fixture, it is
not a PII detector. The real gateway does the detection; here the SSN in
ticket five just makes the blocked ticket read true. See
[It replays; it does not evaluate](https://docs.donkey-kit.dev/simulator.md#how-the-simulator-works).
## Verification status
The proxy *contract* the demo depends on — the base URL shape (no `/v1`), the
`client_id`/`client_secret` header pair, the attribution headers, and the four
live-verified rejection shapes including PII — is **live-verified**
(see the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md)). `ChatOpenAI` / `StateGraph` are
the frameworks' own classes and `.ainvoke` is their documented API:
construction via the SDK factory is the verified surface, and everything after
is the framework's own runtime.
## Where to go next
- [Nightly batch](https://docs.donkey-kit.dev/scenarios/nightly-batch.md) — the same governance, applied to an
unattended overnight job that paces itself against a budget window.
- [Typed refusals](https://docs.donkey-kit.dev/errors.md) — the full exception taxonomy the block above lands
in, with a per-exception "retryable?" cookbook.
- [Testing & conformance](https://docs.donkey-kit.dev/testing.md) — run this same agent factory through the
conformance suite (`pytest --donkey-conformance`).
---
Source: https://docs.donkey-kit.dev/scenarios/nightly-batch.md
# Nightly batch
50,000 product records, enriched overnight against a governed model, budget
window resetting every hour, no human awake. Without a budget object the script
runs flat out, takes a `429` partway through, crashes, and someone re-runs it
from record zero in the morning — spending the budget twice to do the same work.
The fix is to make the remaining budget a [first-class object](https://docs.donkey-kit.dev/budget.md) and pace
against it: slow down *before* the wall, wait for the window to reset, and carry
on. This page runs that loop end to end against the [local
simulator](https://docs.donkey-kit.dev/simulator.md) in about ninety seconds, instead of "we'll find out
tonight."
## What it demonstrates
- **`pace()` raises before the request that would cross your reserve**, not
after a `429` comes back — the distinction that is the whole feature.
- **The window actually resets and the job resumes**, driven by the simulator's
`budget` [scenario](https://docs.donkey-kit.dev/simulator.md#scenario-scripting), which runs a *real*
wall-clock-windowed token counter and serves the captured `token-rate-limit`
**429** on exhaustion.
- **The batch finishes unattended** — nobody re-runs anything.
## Run it
### Install the extras
```bash
pip install "donkey-kit[llm,local]"
```
### Boot the simulator with a one-minute budget window
In one terminal, shrink the hour-long window to a minute so the whole
pace-exhaust-reset cycle plays out in seconds:
```bash
donkey mock --port 8080 --scenario budget:limit=20000,window=60s
```
The happy-path `200` carries the live `x-llm-proxy-ratelimit` prose window
(decreasing as you spend); once the window's 20,000 tokens are gone, calls get
the `token-rate-limit` **429** with `x-token-remaining` / `x-token-reset`
recomputed from the real milliseconds left, until the window rolls over.
### Point the SDK at it and run the batch
In a second terminal:
```bash
export DONKEY_LLM_PROXY_URL=http://localhost:8080
export DONKEY_LLM_PROXY_CLIENT_ID=local # simulator ignores auth
export DONKEY_LLM_PROXY_CLIENT_SECRET=local
python enrich.py
```
## The batch loop
The pacing and resume logic is a handful of lines. `pace()` guards each batch;
on `BudgetReserveReached` you wait for the window and continue from where you
left off when `.reset_at` is known. If it is unknown, propagate the signal
instead of retrying at zero delay:
```python
import asyncio
from donkey_kit import Donkey, BudgetReserveReached
async def enrich_all(records, enrich):
async with Donkey.from_env() as donkey:
i = 0
while i < len(records):
batch = records[i : i + 200]
try:
async with donkey.budget.pace(reserve=0.05):
await enrich(donkey, batch)
except BudgetReserveReached as exc:
# We're within 5% of the window's limit — don't take the 429.
if exc.reset_at is None:
raise # waiting cannot make progress without a reset time
await donkey.budget.wait_for_reset() # sleeps until reset_at
continue # retry the same batch
checkpoint(batch) # only advance on success
i += 200
```
`pace(reserve=0.05)` raises `BudgetReserveReached` **before** issuing the
request that would cross the last 5% of the window — so you never spend the
request that earns the `429`. `wait_for_reset()` sleeps until
`donkey.budget.reset_at`, computed from the gateway's `x-token-reset` header
(milliseconds, converted for you). Once that time has elapsed, the old
observation is stale, so `pace()` no longer refuses. A response carrying a
budget signal updates the observed fields; a fresh future `reset_at` makes
the guard active again. A response that does not supply a fresh future
`reset_at` leaves the stale pass-through open. If a partial observation reaches
the reserve without a `reset_at`, the loop re-raises after one attempt instead
of spinning at zero delay. See [Budget & pacing](https://docs.donkey-kit.dev/budget.md).
## The honest limitation
**Budget is only visible in-band.** The gateway reports it on response
headers; there is **no endpoint that answers "what is my remaining budget?"**.
So `donkey.budget.remaining` is only as fresh as your last call, and a
brand-new process knows nothing until its first request completes — which is
why `donkey.budget.observed_at` is part of the public surface. A budget-query
endpoint is filed as an [upstream gap](https://docs.donkey-kit.dev/roadmap.md) against the gateway.
## Verification status
The budget object, `pace()`, and `wait_for_reset()` are **shipped** (Phase 1).
The windowed-counter behaviour you're pacing against here is the
[simulator's](https://docs.donkey-kit.dev/simulator.md) — a faithful replay of the observed live contract
(prose window on the `200`, numeric `x-token-*` trio on the `429`), never a
header shape the gateway does not emit. The end-to-end assertion against the simulator
is exactly what this scenario runs.
## Where to go next
- [Budget & pacing](https://docs.donkey-kit.dev/budget.md) — the full `Budget` object and its two helpers.
- [Local simulator](https://docs.donkey-kit.dev/simulator.md) — the `budget` scenario and how the windowed
counter is computed.
- [Internal copilot](https://docs.donkey-kit.dev/scenarios/internal-copilot.md) — a content-safety guardrail
and per-run correlation for an internal assistant.
---
Source: https://docs.donkey-kit.dev/scenarios/internal-copilot.md
# Internal copilot
An internal copilot answers employee questions against company systems. Its
output has to clear a **content-safety guardrail** before it reaches a person,
and when the guardrail fires you need the refusal to (a) surface as something
you can branch on and (b) carry an id that joins the block back to the run in
your own logs. This page runs that content-safety branch against the [local
simulator](https://docs.donkey-kit.dev/simulator.md), and is honest about the parts of a full internal
copilot that are still blocked on verification.
## What it demonstrates
- **The content-safety branch executes** as a typed
[`ContentSafetyBlocked`](https://docs.donkey-kit.dev/errors.md), driven by the `donkey-sim/content-safety`
sentinel — no need to craft a prompt that a real guardrail would reject.
- **Per-run correlation joins the refusal to your logs.** Inside a
[`donkey.run(id=...)`](https://docs.donkey-kit.dev/telemetry.md) block the bound id becomes the
`X-Correlation-Id` on every request *and* lands on
`ContentSafetyBlocked.correlation_id`, so a log line for the block joins to
the gateway's own record with no extra wiring.
## Run it
### Install the extras
```bash
pip install "donkey-kit[llm,local]"
```
### Boot the simulator
```bash
donkey mock --port 8080
```
No `--scenario` needed: the [`donkey-sim/` sentinel](https://docs.donkey-kit.dev/simulator.md) lets
*you* pick which call fails by setting the request's `model`. The selectable
shapes include `content-safety`.
### Force a content-safety block and catch it typed
```python
import asyncio
import openai
from donkey_kit import Donkey, ContentSafetyBlocked
from donkey_kit.core.errors import classify
async def ask(donkey, question, *, run_id):
client = donkey.llm.client() # the native AsyncOpenAI, governed transport
async with donkey.run(id=run_id):
try:
return await client.chat.completions.create(
# the sentinel forces the captured content-safety 403 from the simulator
model="donkey-sim/content-safety",
messages=[{"role": "user", "content": question}],
)
except openai.APIStatusError as e:
governed = classify(e.response) # -> a DonkeyError subclass
if isinstance(governed, ContentSafetyBlocked):
print(f"[{run_id}] blocked by guardrail:", governed.remediation)
print(f"[{run_id}] correlation id:", governed.correlation_id)
raise governed from e
asyncio.run(...) # DONKEY_LLM_PROXY_URL=http://localhost:8080, throwaway creds
```
**The raw client raises `openai.APIStatusError`, not a `DonkeyError`.**
`donkey.llm.client()` is the real OpenAI SDK, so you bridge into the taxonomy
with `classify(e.response)` — see [Bridging from the raw
client](https://docs.donkey-kit.dev/errors.md#bridging-from-the-raw-client). An adapter that wraps calls in
`typed_refusals()` (as the [support-triage](https://docs.donkey-kit.dev/scenarios/support-triage.md) demo
does) surfaces the typed refusal directly instead.
## What a full internal copilot also needs — and what's blocked
A production internal copilot wants more than a content-safety branch. Three of
those pieces are **not yet buildable** because they depend on surfaces still
blocked on verification — this page shows the *shape* without inventing
an endpoint:
**Governed access to internal tools** — reaching company systems through
governed MCP tools rather than ad-hoc HTTP — is a [Phase 2 surface](https://docs.donkey-kit.dev/roadmap.md).
Exchange→MCP tool discovery is still blocked on verification; the SDK raises
`NotImplementedError("blocked on verification: …")` at call time rather than
guessing an endpoint. See [Tool access](https://docs.donkey-kit.dev/tool-access.md).
**Agent identity and a kill switch** — a verifiable identity for the copilot,
and the ability to disable it centrally — are platform capabilities the SDK's
job is to make *reachable and typed*, not to reimplement. The kill switch's
refusal is already typed: a blocked agent surfaces as
[`AgentKilled`](https://docs.donkey-kit.dev/errors.md), never retried. Stopping a run cleanly on it, and
identity itself, are [Phase 2](https://docs.donkey-kit.dev/roadmap.md); see [Identity](https://docs.donkey-kit.dev/identity.md). Until
then the SDK does not fabricate a stand-in.
## Verification status
The content-safety **discriminator** (the vendor `…-action: reject` header) is
typed by `donkey_kit.core.errors.classify()`, but its exact body is
**documented-but-not-live-captured** — pinned from the policy pages and pending
a live sandbox round-trip. The
simulator replays the captured fixture so the branch runs today; no verification
row flips to `verified` until a live capture confirms the shape. The correlation
mechanism (`donkey.run()` → `X-Correlation-Id` → `.correlation_id`) is shipped
and framework-agnostic.
## Where to go next
- [Typed refusals](https://docs.donkey-kit.dev/errors.md) — the `ContentSafetyBlocked` shape and the
documented-but-not-captured caveat, plus the retryable cookbook.
- [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md) — how the run id ties spans, logs, and the
gateway record together.
- [Support triage](https://docs.donkey-kit.dev/scenarios/support-triage.md) — the same governance surfacing a
refusal directly through a framework adapter.
---
Source: https://docs.donkey-kit.dev/examples.md
# Examples
The [DDK demos repo](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos)
holds runnable examples for every piece of the SDK: the governed client, typed
refusals, budget pacing, simulation, the conformance suite, telemetry,
framework objects and `last_call`. The section has two parts:
- **[General](#general)** — one page per SDK capability, pairing the examples
that show it with the command to run them.
- **[By framework](#by-framework)** — one page per framework, walking through
every script in its folder: what it shows, what it needs, and what it
prints.
The repo ships **two suites on purpose**. They cover the same SDK, but they are
not interchangeable:
| | Narrative demos | Framework scripts |
| --- | --- | --- |
| **Where** | [`demos/claude-made/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made) | [`demos/human-made/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made), one folder per framework |
| **Best for** | A room, a recording, or CI that must stay offline | A terminal you type in, or paste into your own project |
| **Shape** | Numbered demos (`01`–`10`), each a `demo.py` + README told in acts | Short, top-to-bottom scripts, one file each, across ten frameworks |
| **Runner** | `make demo N=03`, `make offline` | `python "demos/human-made//.py"` |
| **Network** | Nine of ten run offline against the local simulator; demo 09 needs a live gateway | Most need gateway credentials; the `*-simulated`, `start-gateway` and `gateway-unavailable` scripts need none |
| **Output** | Masked by default | Not masked — use on a private terminal |
Every narrative demo takes `--target mock` (the default, pointing at the local
simulator) or `--target live` (your real credentials). The framework scripts
have no harness: they use the same environment you already use for the SDK.
Frameworks pin conflicting dependencies, so one virtual environment per
framework is the safe default; each framework page has its install line.
## Setup
### Create a virtual environment
```bash
git clone https://github.com/Donkey-Development-Kit/donkey-development-kit-demos.git
cd donkey-development-kit-demos
python3 -m venv .venv # .venv/ is git-ignored
source .venv/bin/activate # once per terminal
```
Homebrew's `python3` and most Linux distro Pythons are marked *externally
managed* (PEP 668), so a global `pip3 install` fails with
`externally-managed-environment`. Install into `.venv` instead of reaching for
`--break-system-packages`.
### Install the demo harness
```bash
python -m pip install -e .
```
This adds the shared harness to the path. It does not pin an SDK.
### Install the SDK
Pick one:
```bash
# from git, with the extras every offline demo needs
python -m pip install -e ".[sdk]"
# from git, with everything including OpenTelemetry and LangGraph
python -m pip install -e ".[full]"
# from your own SDK checkout
python -m pip install -e "../donkey-development-kit/python[llm,local,test,otel,langgraph,cli]"
# a published dev build from TestPyPI (its dependencies come from PyPI)
python -m pip install -i https://test.pypi.org/simple/ \
--extra-index-url https://pypi.org/simple/ "donkey-kit[llm,local,test,otel,langgraph,cli]"
```
**Quote the `[...]` extras.** zsh, the macOS default shell, treats unquoted
brackets as a glob and fails with `no matches found`.
`uv` users can replace the venv step with `uv venv` and `pip install` with
`uv pip install`.
## Live credentials
The offline narrative demos need none of this. Live runs, the LangGraph agent
example and most framework scripts need three variables for the governed LLM
proxy, plus the model to ask for:
| Variable | What it is |
| --- | --- |
| `DONKEY_LLM_PROXY_URL` | The proxy's base URL, **with a trailing `/`** and no `/v1` — the OpenAI SDK appends `/responses` itself |
| `DONKEY_LLM_PROXY_CLIENT_ID` | The client id the proxy authenticates on (a header, not a bearer token) |
| `DONKEY_LLM_PROXY_CLIENT_SECRET` | The matching client secret |
| `DEMO_MODEL` | The model id the narrative demos ask for; it must be one your proxy routes |
Copy the template and fill it in:
```bash
cp .env.example .env.local # .env.local is git-ignored
```
A shell `export` always wins over a file value, so you can skip the file and
the framework scripts do not (the SDK never reads dotenv files on its own), so
```bash
set -a; source .env.local; set +a
python "demos/human-made/openai/02 - basic-responses-gw.py"
```
Three things trip people up with the framework scripts. **The model id:**
they hardcode `gpt-4o`, while the provisioned DDK proxies route
`gpt-5-mini` — change the string, or expect a routing refusal or a
`ModelSubstituted`. **Filenames have spaces:** always quote the path.
**Policies decide refusals:** a `typed-refusals-live` script prints `NO
REFUSAL` when the proxy does not have that policy applied — that is the
proxy telling the truth, not the script failing.
## Commands
```bash
make list # the narrative demos and what each one needs
make demo N=03 # one narrative demo
make demo N=01 ARGS="--target live" # the same demo against your gateway
make offline # every narrative demo that needs no credentials
make doctor # what is installed, and what will therefore run
make mock # the local simulator in the foreground, for a second pane
DEMO_PAUSE=1 make demo N=03 # pause between acts — use this when presenting
```
Flags for the demo go in `ARGS`, not on the end of the `make` line. `python
run.py 03` works too, and each narrative demo is a plain script
(`python demos/claude-made/03_budget_and_pacing/demo.py`) once the repo is
installed. `make doctor` reports what it found without printing any values.
## Credential safety
The narrative demos assume they will be screen-shared and recorded:
- **Output is masked by default.** Every value a narrative demo prints is
scrubbed: gateway hostnames, the client id and secret, and the tenant
identifiers that ride along in captured responses. Header *names* are shown;
their values are not.
- **Turning masking off announces itself.** `DEMO_REDACT=0` prints a warning
banner in every narrative demo's run context. Use it only when debugging
privately.
- **No captured traffic is vendored.** Fixtures load from the installed SDK,
not from copies in the demos repo.
- **`make scan` reads content, not filenames.** It fails on assigned credential
values, bearer tokens, instance ids, UUIDs and non-allowlisted hostnames.
`make hooks` installs it as a pre-commit hook.
```bash
make scan # everything tracked
make hooks # then it runs on every commit
```
**The framework scripts do not mask anything.** They print completions and
error strings exactly as the SDK returned them. Run them on a private
terminal, not on a shared screen or recording.
## General
Stock client versus governed client, then `@donkey.governed` and `@donkey.tool`.
Every captured rejection shape through `classify()`, plus `GatewayUnavailable`.
The token window as an object; `pace(reserve=)` and `wait_for_reset()`.
Run your refusal branch with `donkey.simulate()` and `donkey mock --scenario`.
`pytest --donkey-conformance` grading a naive agent, then the fixed one.
GenAI spans, `donkey.run(id=…)` correlation, and zero-config OTLP.
Native framework objects, `connection_kwargs()`, and `resolve()`.
A real tool-calling loop, governed end to end.
Who served the call, what it routed to, what it cost — and `ModelSubstituted`.
## By framework
Each framework page covers one folder of
[`demos/human-made/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made).
How much DDK can do depends on who owns the HTTP transport: where the SDK does
(OpenAI, LangGraph, OpenAI Agents SDK, Strands, Anthropic), you get `last_call`,
run ids and `simulate()`; where the framework does, only the credential
headers go on the wire.
The stock client: `last_call`, refusals, pacing, spans, streaming, guardrails, JWT wallets.
A real `ChatOpenAI`, `typed_refusals()` out of `create_agent`, the run id in tools.
`OpenAIResponsesModel` on the governed client, `Runner.run` with a tool.
`OpenAIChatClient` with the governed headers; refusals from `ChatClientException`.
`OpenAIModel(client=donkey.openai())`, tools, and a 429 Strands retries itself.
`donkey.crewai.llm()` for a direct call and a crew; typed refusals.
An `OpenAILike` for `complete()` and `chat()`; typed refusals.
A `LiteLlm` model in an `InMemoryRunner` — and why its refusals are not typed.
The native client on `/v1/messages`, `last_call` and simulated refusals.
Plain `httpx` against a `Format=Gemini` proxy, errors typed by `classify()`.
---
Source: https://docs.donkey-kit.dev/examples/general/governed-client.md
# Governed client
You do not need the SDK to reach the gateway — a stock OpenAI client with a
`base_url` and two headers gets there. These examples show that first, then
show what the stock client leaves you holding. Both clients hit the same
gateway and get the same PII refusal back; the stock one gives you an
`openai.APIStatusError` and a JSON body to parse, the governed one gives you a
`PIIDetected` with the flagged entities and a remediation string. And because
every request leaves through one client, that same call also updated the
budget window, stamped a correlation id and opened a GenAI span, with no wiring
from you.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 01 | `donkey.openai()` and `sync=True`, the injected headers, one call and what it updated, stock vs governed on the same PII refusal, `@donkey.governed` / `@donkey.tool` | Nothing (simulator) |
| OpenAI script 01 | Stock OpenAI with no gateway — the baseline | `OPENAI_API_KEY` |
| OpenAI script 02 | `donkey.openai()` plus `last_call` on a successful call | Proxy credentials |
| OpenAI script 09 | `@donkey.governed` and `@donkey.tool` on their own | Nothing |
## Run it
```bash
make demo N=01 # offline, against the local simulator
make demo N=01 ARGS="--target live" # against your gateway
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 01 — the governed client
Reaching the gateway is easy. Everything that hangs off the client is the product.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target mock
proxy base_url http://127.0.0.1:8080/
output masking on
credentials fake — the simulator enforces no auth
[1] donkey.openai() returns a real OpenAI client, already governed
donkey = Donkey.from_env()
client = donkey.openai() # -> openai.AsyncOpenAI
blocking = donkey.openai(sync=True) # -> openai.OpenAI, same governance
type openai.AsyncOpenAI
sync=True openai.OpenAI
base_url http://127.0.0.1:8080/
headers the SDK injects:
Accept application/json
Content-Type application/json
User-Agent AsyncOpenAI/Python 3.19.2
client_id (41 chars)
client_secret (45 chars)
OpenAI-Organization
OpenAI-Project
Note the base URL has no /v1 — the ingress is https://// and the
OpenAI SDK appends /responses itself. Auth is the client_id / client_secret header
pair, not a bearer token. sync=True is the same client without asyncio — useful for a
straight-line script.
[2] One call. Nothing new to learn — it is the OpenAI SDK.
response = await client.responses.create(
model='gpt-4o',
input='Say hello in exactly three words.',
)
WARN The simulator replays a captured success response, so the reply below answers the prompt that was recorded, not the one just sent. Run with --target live for a real completion.
reply A sleepy unicorn named Luma painted soft silver stars across the night sky with her glowing horn, then curled up on a moonbeam so all the children below could fall asleep beneath her gentle, sparkling light.
input tokens 17
output tokens 51
That call also did these things nobody asked for, because every request leaves through
one client:
budget.remaining 98000
budget.limit 100000
budget.fraction_used 2.0%
budget.observed_at 2026-09-24 09:27:22.710265+00:00
last_call.status observed
last_call.served_model gpt-5.1
last_call.total_tokens 68
• a correlation id went out on the request
• a gen_ai.* span opened and closed around it (demo 06)
donkey.last_call is the success-path counterpart to a typed refusal: who served this,
what they actually routed to, and what the call cost. Demo 10 walks the whole record —
routing, fallback, cached/reasoning tokens, and the opt-in ModelSubstituted error.
[3] The same refusal, through both clients
The simulator serves a real captured PII rejection when the model id is the sentinel
below. This is the byte-identical body a live gateway sent.
# A: stock OpenAI client, no SDK — just base_url + headers
raw = openai.AsyncOpenAI(base_url=..., api_key=..., default_headers=...)
await raw.responses.create(model='donkey-sim/pii-detected', input='My email is a@b.com')
A: raised openai.PermissionDeniedError
A: status 403
A: you get a JSON body to parse, and a status code to guess from
# B: the same request, through the governed client
try:
await client.responses.create(model=..., input="...")
except openai.APIStatusError as exc:
raise classify(exc.response) from exc
B: raised PIIDetected
B: policy pii-detection
B: entities ['Email']
B: remediation The PII-detection policy blocked this request because the prompt (or completion) contained personally identifiable information. Remove or redact the flagged values, or relax the policy's entity list / action in API Manager.
PASS A 403 that is a policy refusal, not an auth failure — and it says so.
classify() is the bridge, because the raw client raises openai.* errors and the SDK
does not silently re-map them. Demo 02 walks the full taxonomy.
[4] @donkey.governed and @donkey.tool — the one-line on-ramps
@donkey.governed(team="support")
async def handle_ticket(ticket):
... # every model call inside shares one run id
There is deliberately no id= on the decorator: a fixed id pinned across every call
would collapse unrelated tickets into one correlation. When you need to pin a business
id, use donkey.run(id=...) directly (demo 06).
run id inside 7512d15e0ea94fe4b6f321aae903a781
cost tags team=support project=triage
run id inside 491f95bfc2cc4d88b7b0ae1f200de6b5
cost tags team=support project=triage
PASS each invocation opened a fresh run
run id after 235d9e98804b480092a2738f58fd7997
PASS restored to the enclosing context — nested run() rebinds, then restores
@donkey.tool
def lookup_sku(sku: str) -> str:
"""Return stock for a product SKU."""
...
registered lookup_sku
signature (sku: 'str') -> 'str'
docstring Return stock for a product SKU.
PASS same function object — the decorator records it, it does not wrap it
PASS ValueError — an undescribed tool is rejected at decoration time
The same marker is what a Phase 2 scanner and an A2A agent-card generator will both
read. Neither consumer is built yet; this is the annotation they will look for, not a
wrapper around the tool.
The point
─────────
The wrapper is not sold as a way to reach the gateway. It is the one place every
request enters and every response leaves — which is why budget, last_call, typed
refusals, correlation ids, spans and simulation can all attach without the developer
wiring each one. @donkey.governed is that attachment as a function decorator;
@donkey.tool is the marker a scanner can find without executing it.
────────────────────────────────────────────────────────────────────────────────────────
```
```bash
python "demos/human-made/openai/01 - basic-responses-no-gw.py" # needs OPENAI_API_KEY
python "demos/human-made/openai/02 - basic-responses-gw.py" # needs proxy credentials
python "demos/human-made/openai/09 - governed-and-tool.py" # no gateway
```
Against the simulator the reply in act 2 is a captured response, so it
answers the recorded prompt rather than the one just sent. The demo prints a
warning saying so. The PII act against `--target live` needs the PII
detection policy applied with `Email` among its entities and its action set
to `Reject` — the default action, `Log`, does not block.
## Key code
The governed client is a real OpenAI client, already pointed at the proxy
(OpenAI script 02):
```python
async with Donkey.from_env() as donkey:
client = donkey.openai() # THIS is returning the native openai
response = await client.responses.create(
model="gpt-4o",
input="Say hello in exactly three words.",
)
print(response.output_text)
last = donkey.last_call
print("last_call.status ", last.status.value)
print("last_call.served_model", last.served_model)
```
The same refusal through the governed client, typed with `classify()`
(narrative demo 01, act 3):
```python
client = donkey.openai()
try:
await client.responses.create(model=pii_model, input=PII_PROMPT)
except openai.APIStatusError as exc:
governed = classify(exc.response)
say.field("B: raised", type(governed).__name__)
say.field("B: policy", getattr(governed, "policy", "—"))
say.field("B: entities", getattr(governed, "entities", []), raw=True)
say.field("B: remediation", getattr(governed, "remediation", "—"))
```
The one-line on-ramps (OpenAI script 09). `@donkey.governed` opens a fresh run
per invocation; `@donkey.tool` records the function without wrapping it, and
rejects a tool with no docstring at decoration time:
```python
@donkey.governed(team="support", project="triage")
def handle_ticket(ticket: str) -> str:
print("run id inside", current_correlation_id())
print("cost tags ", current_cost_tags())
return ticket
@donkey.tool
def lookup_sku(sku: str) -> str:
"""Return stock for a product SKU."""
return "42"
print("same function object ", lookup_sku is registered_tools()[-1].func)
```
Note the base URL has no `/v1`, and auth is a `client_id` / `client_secret`
header pair rather than a bearer token. There is deliberately no `id=` on
`@donkey.governed`: a fixed id would collapse unrelated calls into one
correlation. Use `donkey.run(id=...)` when you need to pin a business id (see
[Telemetry](https://docs.donkey-kit.dev/examples/general/telemetry.md)).
**Learn more:** [Introduction](https://docs.donkey-kit.dev/) · [Model access](https://docs.donkey-kit.dev/frameworks.md) · [CLI & decorators](https://docs.donkey-kit.dev/cli.md)
**Source:**
[narrative demo 01](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/01_governed_client) ·
[script 01](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/01%20-%20basic-responses-no-gw.py) ·
[script 02](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/02%20-%20basic-responses-gw.py) ·
[script 09](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/09%20-%20governed-and-tool.py)
---
Source: https://docs.donkey-kit.dev/examples/general/typed-refusals.md
# Typed refusals
Governance outcomes should be something you branch on, not something you
parse. These examples run the gateway's rejection shapes through `classify()`
and show the exception hierarchy you write `except` clauses against. The
discriminator is deliberately not the status code: a PII block is a 403 but is
not an auth failure, and an injection block is identified by a header. They
also show the one failure `classify()` cannot produce — `GatewayUnavailable`,
raised when there is no HTTP response at all.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 02 | Nine captured shapes through `classify()`, what the hierarchy buys, the handler you write, and a dead origin raising `GatewayUnavailable` | Nothing — no gateway, no simulator |
| OpenAI script 04 | Live `UpstreamRequestError`, `PIIDetected`, `TokenBudgetExceeded` and `AuthError` on a blocking client, no `async` | Proxy credentials, plus policies for the PII and budget cases |
| OpenAI script 11 | A dead origin surfacing `GatewayUnavailable` as `__cause__` | Nothing |
## Run it
```bash
make demo N=02
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 02 — typed refusals
The gateway's rejection shapes, mapped to exceptions you can branch on.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target offline — no gateway, no simulator, no credentials
output masking on
Nine captured responses through classify()
──────────────────────────────────────────
from donkey_kit.core.errors import classify
governed = classify(response) # -> a typed DonkeyError subclass
client-id-missing consumer auth — a genuinely missing/wrong client id
HTTP 401
classified as AuthError
pii-detected PII policy — a 403 that is NOT an auth failure
HTTP 403
classified as PIIDetected
.policy pii-detection
.entities ['Email']
token-rate-limit token budget — a 429 with an EMPTY body; state is header-only
HTTP 429
classified as TokenBudgetExceeded
.policy token-rate-limit
.retry_after 41.728
injection-protection prompt injection — identified by a header, not a status
HTTP 400
classified as PromptInjectionBlocked
.policy prompt-injection-protection
regex-prompt-guard regex prompt guard — 403 keyed on matched_patterns, not auth
HTTP 403
classified as PromptInjectionBlocked
.policy regex-prompt-guard
content-safety content safety / guardrails — 403 keyed on a vendor reject header
HTTP 403
classified as ContentSafetyBlocked
.policy content-safety
.categories ['severity_hate', 'severity_violence']
content-moderation undiscriminated moderation — no live capture, left unnamed
HTTP 400
classified as PolicyViolation
.policy unknown
model-not-found upstream passthrough — the provider's own error, not a policy
HTTP 400
classified as UpstreamRequestError
.code model_not_found
.error_type invalid_request_error
.param model
upstream-5xx provider failure — retryable, unlike every refusal above
HTTP 503
classified as UpstreamModelError
Why the hierarchy is shaped this way
────────────────────────────────────
PASS a PII block is a policy refusal, not an auth error
PASS a token-budget 429 is also a policy refusal — so one `except PolicyViolation` catches both
PASS content-safety is ContentSafetyBlocked, still a PolicyViolation
PASS regex-prompt-guard is PromptInjectionBlocked with its own policy name
PASS an upstream 400 is NOT a policy refusal — it is your request that is wrong, not the gateway saying no
PASS the budget refusal carries retry_after, parsed from x-token-reset (milliseconds, not an epoch)
PASS GatewayUnavailable is not a PolicyViolation — nothing was refused, the request never arrived
PASS a transport failure has no request_id — there was no response to read the gateway's id from
What that looks like in your agent
──────────────────────────────────
def ask(client: Any, prompt: str) -> str:
try:
try:
response = client.responses.create(model="gpt-4o", input=prompt)
except openai.APIStatusError as exc: # the gateway answered: type it
raise classify(exc.response) from exc
except openai.APIConnectionError as exc: # raised on the transport, wrapped
cause = exc.__cause__ # GatewayUnavailable, ModelSubstituted
if isinstance(cause, DonkeyError):
raise cause from cause.__cause__ # keep its own cause chain
raise
return response.output_text
except PIIDetected as e: # 403, and e.entities says what tripped
return f"redact {e.entities} and retry"
except ContentSafetyBlocked as e: # 403, e.categories is the moderation analog
return f"revise for {e.categories}"
except TokenBudgetExceeded as e: # 429, terminal — never retry it
return f"wait {e.retry_after}s for the budget to reset"
except PolicyViolation as e: # any other gateway refusal
return f"escalate the {e.policy} refusal"
except GatewayUnavailable as e: # NO response — not a refusal
return f"checkpoint and shed: {e.base_url} is unreachable"
except ModelSubstituted as e: # NOT classify() — you opted in (demo 10)
return f"pin or accept {e.served_model}"
except UpstreamRequestError as e: # your request was wrong (e.code)
return f"fix the request: {e.code}"
except UpstreamModelError: # provider 5xx — this one IS retryable
return "retry with backoff"
The bridge is an INNER try. An exception raised inside one except clause is never
handed to a sibling clause of the same try, so the typed handlers have to sit one
level out. The bridge has two arms because the OpenAI client reports two ways: a
refusal is an APIStatusError carrying the gateway's response, and an error the
transport raises itself is an APIConnectionError with the typed DonkeyError on
__cause__.
PIIDetected redact ['Email'] and retry
ContentSafetyBlocked revise for ['severity_hate', 'severity_violence']
TokenBudgetExceeded wait 41.728s for the budget to reset
PromptInjectionBlocked escalate the prompt-injection-protection refusal
UpstreamRequestError fix the request: model_not_found
UpstreamModelError retry with backoff
GatewayUnavailable checkpoint and shed: http://127.0.0.1:9 is unreachable
PASS the flat version — typed handlers as siblings of the classify() clause — lets PIIDetected escape: its handler never runs
What is typed from docs, and what is still unnamed
──────────────────────────────────────────────────
Four of these shapes are live-verified against a real proxy: consumer auth, PII, token
rate limit, and upstream passthrough. Injection, regex prompt guard, and content-
safety are typed from the documented wire shapes — classify() produces
PromptInjectionBlocked / ContentSafetyBlocked — and are pending a live sandbox
capture. That is the same posture as header-based injection: named because the shape
is specified, not because a capture has landed yet.
content-moderation PolicyViolation
remediation This refusal matched no documented rejection shape, so its contract is unconfirmed (#184, #253). It is terminal and was NOT retried. Please file an issue on the donkey-development-kit repo with the response status, headers and body (all carried on this exception's .response) so the shape can be typed.
An undiscriminated content-moderation 4xx still falls through to a generic
PolicyViolation. That leftover shape has never been captured from a live gateway, so
it is left unnamed rather than given a class that would imply more certainty than
exists.
ModelSubstituted is not in the table above because it is not a gateway refusal and
classify() never produces it. It is raised by the transport when you opt into
on_model_substitution='raise' and the gateway serves a different model than you asked
for. Demo 10.
GatewayUnavailable is the other type classify() never produces: there is no HTTP
response to classify. DNS, connection refused, TLS, timeout — the transport wraps
those as a typed DonkeyError so a long-running agent can tell 'lost the gateway' from
a policy refusal without matching raw httpx exceptions. It is not retried. Act 5
actually raises it.
A refused connection, typed — not a raw httpx error
───────────────────────────────────────────────────
donkey = Donkey(DonkeyConfig(llm_proxy_url="http://127.0.0.1:9/", ...))
client.responses.create(...) # nothing is listening
# -> GatewayUnavailable, not ConnectError
PASS GatewayUnavailable — the request never left the building
base_url http://127.0.0.1:9
cause ConnectError
request_id None
call_id c46d490e7aef47b18a391a2764087952
The gateway could not be reached and no HTTP response came back. The three usual
causes: (1) the host is unreachable — DNS failure or the gateway is down; (2) the
configured base URL is wrong; or (3) network egress to the gateway is blocked — a
firewall or air-gapped environment. Run `donkey doctor` to diagnose connectivity, and
check `base_url` on this error against your gateway's address.
PASS not a PolicyViolation — nothing was refused, because nothing arrived
────────────────────────────────────────────────────────────────────────────────────────
```
```bash
python "demos/human-made/openai/04 - typed-refusals-live.py" # needs proxy credentials
python "demos/human-made/openai/11 - gateway-unavailable.py" # no gateway
```
Narrative demo 02 loads its fixtures from the installed SDK
(`donkey_kit.simulator.fixtures`) — the same bytes `classify()` is tested
against and `donkey mock` serves. OpenAI script 04 provokes the upstream and
auth cases with nothing extra; the PII case needs the PII detection policy
with `Email` and action `Reject`, and the budget case needs the token rate
limit policy with a small `maximumTokens`.
## Key code
The handler shape the hierarchy is designed for (narrative demo 02, act 3 —
the demo runs this exact function against a simulated refusal of each type):
```python
def ask(client: Any, prompt: str) -> str:
try:
try:
response = client.responses.create(model="gpt-4o", input=prompt)
except openai.APIStatusError as exc: # the gateway answered: type it
raise classify(exc.response) from exc
except openai.APIConnectionError as exc: # raised on the transport, wrapped
cause = exc.__cause__ # GatewayUnavailable, ModelSubstituted
if isinstance(cause, DonkeyError):
raise cause from cause.__cause__ # keep its own cause chain
raise
return response.output_text
except PIIDetected as e: # 403, and e.entities says what tripped
return f"redact {e.entities} and retry"
except ContentSafetyBlocked as e: # 403, e.categories is the moderation analog
return f"revise for {e.categories}"
except TokenBudgetExceeded as e: # 429, terminal — never retry it
return f"wait {e.retry_after}s for the budget to reset"
except PolicyViolation as e: # any other gateway refusal
return f"escalate the {e.policy} refusal"
except GatewayUnavailable as e: # NO response — not a refusal
return f"checkpoint and shed: {e.base_url} is unreachable"
except ModelSubstituted as e: # NOT classify() — you opted in (demo 10)
return f"pin or accept {e.served_model}"
except UpstreamRequestError as e: # your request was wrong (e.code)
return f"fix the request: {e.code}"
except UpstreamModelError: # provider 5xx — this one IS retryable
return "retry with backoff"
```
The bridge is an **inner** `try`. An exception raised inside one `except`
clause is never handed to a sibling clause of the same `try`, so typed handlers
written next to the `classify()` clause would never run. The bridge has two
arms because the OpenAI client reports failures two ways: a gateway refusal is
an `openai.APIStatusError` carrying the response, and an error the transport
raises itself (`GatewayUnavailable`, `ModelSubstituted`) arrives as an
`openai.APIConnectionError` with the typed error on `__cause__`.
A live refusal on a blocking client (OpenAI script 04):
```python
cfg = DonkeyConfig.from_env()
donkey = Donkey(cfg)
client = donkey.openai(sync=True)
with donkey.run(id="live-refusals-PIIDetected"):
try:
raw = client.responses.with_raw_response.create(model=MODEL, input=PII_PROMPT)
except openai.APIStatusError as err:
error = classify(err.response)
print(f" REFUSED {type(error).__name__} (HTTP {err.response.status_code})")
print(f" entities {getattr(error, 'entities', None)}")
print(f" remediation {getattr(error, 'remediation', None)}")
print(f" correlation_id {getattr(error, 'correlation_id', None)}")
```
When nothing is listening, the OpenAI client wraps the transport error and the
typed `GatewayUnavailable` sits on `__cause__` (OpenAI script 11):
```python
donkey = Donkey(
DonkeyConfig(
llm_proxy_url="http://127.0.0.1:9/",
llm_proxy_client_id="demo-client-id-not-a-real-credential",
llm_proxy_client_secret="demo-client-secret-not-a-real-credential",
timeout_s=2.0,
max_retries=0,
)
)
client = donkey.openai(sync=True)
try:
client.responses.create(model="gpt-4o", input="hello")
except Exception as err:
hit = err if isinstance(err, GatewayUnavailable) else err.__cause__
if isinstance(hit, GatewayUnavailable):
print("base_url ", hit.base_url)
print(hit.remediation)
```
A token-budget 429 and a PII 403 are both `PolicyViolation`s, so one `except
PolicyViolation` catches either. An upstream 400 is not — your request was
wrong, the gateway did not say no. `GatewayUnavailable` is not a
`PolicyViolation` either: nothing was refused, because nothing arrived. An
undiscriminated `content-moderation` 4xx falls through to a generic
`PolicyViolation`.
**Learn more:** [Typed refusals](https://docs.donkey-kit.dev/errors.md)
**Source:**
[narrative demo 02](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/02_typed_refusals) ·
[script 04](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/04%20-%20typed-refusals-live.py) ·
[script 11](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/11%20-%20gateway-unavailable.py)
---
Source: https://docs.donkey-kit.dev/examples/general/budget-and-pacing.md
# Budget & pacing
The gateway reports your token budget only in-band, on response headers —
there is no endpoint to ask how much is left. `donkey.budget` reads those
headers for you, so the window is an object rather than a header you parse.
The useful half is pacing: `pace(reserve=…)` refuses locally *before* issuing a
request that would cross your reserve, turning a 429 you would have to recover
from into an exception you chose to raise, and `wait_for_reset()` sleeps once
until the window rolls over.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 03 | A cold process knowing nothing, the window updating per response, `pace(reserve=0.05)` raising `BudgetReserveReached`, `wait_for_reset()`, and the terminal 429 | Nothing (simulator) |
| OpenAI script 05 | `pace(reserve=)` letting the first call through, stopping the second, then `wait_for_reset()` | Proxy credentials |
## Run it
```bash
make demo N=03
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 03 — budget and pacing
The token window as an object, and refusing to cross it before the gateway does.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target mock
proxy base_url http://127.0.0.1:8080/
output masking on
credentials fake — the simulator enforces no auth
[1] A new process knows nothing until its first call returns
before any call remaining=None limit=None used=unobserved
Every field is None rather than zero. An unobserved budget reports 'I don't know',
because reporting 0 remaining would be a lie that stops an agent that could have run.
[2] Each response updates the window, with no code from you
await client.responses.create(model=..., input=...)
donkey.budget.remaining # already up to date
after call 1 remaining=97500 limit=100000 used=2.5%
after call 2 remaining=97000 limit=100000 used=3.0%
after call 3 remaining=96500 limit=100000 used=3.5%
observed_at 2026-09-24 09:27:23.806252+00:00
reset_at 2026-09-24 09:28:23.806252+00:00
A live 200 carries this window as the prose header x-llm-proxy-ratelimit — that
sentence is live-verified. The numeric x-token-* trio is verified on the 429. The
simulator synthesises a decreasing window in the same prose shape, so the numbers
above are illustrative; the parse path is not.
[3] pace(reserve=…) refuses before the request goes out
Rather than issue the 200 calls it would take to drain the simulator's window, we let
the budget observe a response that says we are already at 96% — the same code path a
real near-exhausted window takes.
async with donkey.budget.pace(reserve=0.05):
await enrich(batch) # never runs if the reserve is crossed
observed remaining=4000 limit=100000 used=96.0%
reset_at 2026-09-24 09:27:24.806704+00:00
PASS BudgetReserveReached — the request was never issued
fraction_used 96.0%
reserve 5.0%
reset_at 2026-09-24 09:27:24.806704+00:00
BudgetReserveReached is deliberately NOT a PolicyViolation. A refusal is the gateway
saying no and is terminal; this is your own client-side signal, raised locally, that
you are expected to recover from.
try:
async with donkey.budget.pace(reserve=0.05):
await enrich(batch)
except BudgetReserveReached:
await donkey.budget.wait_for_reset() # one sleep, never a spin loop
after wait_for_reset remaining=4000 limit=100000 used=96.0%
PASS wait_for_reset() slept until reset_at — one sleep, never a spin loop
The local object is still the last observation. Waiting does not invent a fresh
window; the next call is what refreshes remaining / limit / reset_at. That is the same
in-band rule as act 1.
[4] And if you do cross it, the 429 is terminal
classified as TokenBudgetExceeded
retry_after 41.728
PASS The transport never retried it — retrying only burns the same window.
This is the scenario the conformance suite checks other people's agents for, because
retrying a budget refusal is the single most common way an agent turns one refusal
into a rate-limit spiral. See demo 05.
────────────────────────────────────────────────────────────────────────────────────────
```
```bash
python "demos/human-made/openai/05 - budget_and_pacing.py" # needs proxy credentials
```
On a live gateway the window arrives only when the token rate limit policy
is applied to the proxy. The simulator synthesises a decreasing window so
pacing can run locally; its happy-path numbers are illustrative, the parse
path is not.
## Key code
Pacing, and recovering from it (narrative demo 03, act 3):
```python
try:
async with donkey.budget.pace(reserve=0.05):
await enrich(batch)
except BudgetReserveReached:
await donkey.budget.wait_for_reset() # one sleep, never a spin loop
```
Against a live proxy, the first call is unobserved so `pace()` lets it through
and the window arrives in-band; the second trips the reserve (OpenAI script 05):
```python
async with donkey.budget.pace(reserve=0.99999):
response = await client.responses.create(
model="gpt-4o",
input="Say hello in exactly three words.",
)
budget = donkey.budget
print("after request 1, budget remaining is", budget.remaining)
print("after request 1, budget fraction_used is", budget.fraction_used)
try:
async with donkey.budget.pace(reserve=0.99999):
response = await client.responses.create(
model="gpt-4o",
input="Say hello in exactly three words.",
)
except BudgetReserveReached as exc:
print("stopped locally [in-script]", exc.fraction_used, exc.reserve)
```
An unobserved budget reports `None` for every field, never `0` — reporting `0`
remaining would stop an agent that could have run. `BudgetReserveReached` is
deliberately not a `PolicyViolation`: a refusal is the gateway saying no and is
terminal, while this is a local signal you are expected to recover from. If you
do cross the window, the resulting `TokenBudgetExceeded` is not retried by the
transport — retrying only burns the same window.
**Learn more:** [Budget & pacing](https://docs.donkey-kit.dev/budget.md)
**Source:**
[narrative demo 03](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/03_budget_and_pacing) ·
[script 05](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/05%20-%20budget_and_pacing.py)
---
Source: https://docs.donkey-kit.dev/examples/general/simulating-refusals.md
# Simulating refusals
Every agent has an `except PIIDetected:` branch that has never executed.
Getting a real gateway to refuse on demand means finding a prompt that trips a
live policy — slow, flaky, and not something you can put in CI.
`donkey.simulate()` swaps a fixture-returning transport onto the client for the
next N calls, so the branch runs against the same captured refusal a real
gateway sent, with no network. For a client that does not use the SDK at all,
`donkey mock --scenario` scripts the running simulator the same way.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 04 | `simulate()` for each refusal, `times=`, what it refuses to fake, the same injection through LangChain's `ChatOpenAI`, and `donkey mock --scenario` parsing | Nothing; the LangChain act runs only if `[langgraph]` is installed |
| OpenAI script 03 | A `simulate()` loop over five refusal types, including `ContentSafetyBlocked` | Proxy credentials in the environment (no network calls) |
## Run it
```bash
make demo N=04
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 04 — simulating refusals in-process
Run the except branch that has never executed. No network, no server, no credentials.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target mock
proxy base_url http://127.0.0.1:8080/
output masking on
credentials fake — the simulator enforces no auth
[1] Normally, the happy path is all you ever exercise
agent returned A sleepy unicorn named Luma painted soft silver stars across the night sky with her glowing horn, then curled up on a moonbeam so all the children below could fall asleep beneath her gentle, sparkling light.
refusal branch ran False
[2] One context manager, and the branch runs
with donkey.simulate(PIIDetected):
await agent.run("...") # fails as a real PIIDetected
agent returned redacted ['Email'] and asked the user to rephrase
handled ['PIIDetected']
PASS The branch executed against the real captured 403 body.
[3] Every refusal you need to handle, one line each
TokenBudgetExceeded queued for retry after 42s — did NOT retry now
PIIDetected redacted ['Email'] and asked the user to rephrase
ContentSafetyBlocked revised ['severity_hate', 'severity_violence'] and did not retry
[4] `times` counts calls, and normal service resumes after
with donkey.simulate(TokenBudgetExceeded, times=2):
await agent.run("a") # refused
await agent.run("b") # refused
await agent.run("c") # succeeds — the injection is spent
call 1 queued for retry after 42s — did NOT retry now
call 2 queued for retry after 42s — did NOT retry now
call 3 A sleepy unicorn named Luma painted soft silver stars across the night sky with her glowing horn, then curled up on a moonbeam so all the children below could fall asleep beneath her gentle, sparkling light.
[5] It will inject a documented shape, and refuse to invent the rest
with donkey.simulate(ContentSafetyBlocked):
await agent.run("...") # a real ContentSafetyBlocked
ContentSafetyBlocked revised ['severity_hate', 'severity_violence'] and did not retry
PASS The documented content-safety fixture classifies and injects.
with donkey.simulate(ToolInvocationError):
...
PASS ValueError — no captured fixture maps back to it
message simulate() cannot inject ToolInvocationError: no captured fixture maps back to it via classify(). Supported: AgentKilled, AuthError, ContentSafetyBlocked, ModelNotRoutable, PIIDetected, PolicyViolation, PromptInjectionBlocked, TokenBudgetExceeded, UpstreamModelError, UpstreamRequestError.
with donkey.simulate(GatewayUnavailable):
...
PASS ValueError — a transport failure has no captured body to inject
message simulate() cannot inject GatewayUnavailable: no captured fixture maps back to it via classify(). Supported: AgentKilled, AuthError, ContentSafetyBlocked, ModelNotRoutable, PIIDetected, PolicyViolation, PromptInjectionBlocked, TokenBudgetExceeded, UpstreamModelError, UpstreamRequestError.
Tool invocation, registry, and provisioning errors are not gateway refusals, and they
have no captured wire shape. GatewayUnavailable is the same kind of gap for a
different reason: there is no HTTP response at all, so there is nothing to replay.
Injecting a plausible body would let you write a handler against a body that does not
exist — so simulate() refuses instead. Provoke it by pointing at a dead origin (demo
02 act 5).
[6] It works through a framework too, because it is on the transport
model = donkey.langgraph.chat_model("gpt-4o") # a real ChatOpenAI
with donkey.simulate(PIIDetected):
await model.ainvoke("...")
LangChain raised OpenAIPermissionDeniedError
classifies as PIIDetected
PASS Same fixture, same taxonomy, through the framework's own object.
[7] The same idea, as a running simulator a stock client can hit
simulate() swaps the transport on a Donkey you already own — that is the unit-test
form. When the client is a stock OpenAI SDK pointed at donkey mock, you script the
server instead:
donkey mock --scenario pii_block:every=2 \
--scenario 'injection:on-pattern=ignore previous' \
--scenario budget:limit=200,window=5s,cost=80
pii_block:every=2 ['pass', 'pii-detected', 'pass', 'pii-detected']
injection('hello') pass
injection('ignore previous') injection-protection
budget spec BudgetScenario
PASS Three specs, three stateful rules — the same captured fixtures classify() is tested against.
pii_block fails every Nth call. injection matches request text. budget is a real wall-
clock window: passing 200s carry the prose x-llm-proxy-ratelimit header; exhaustion
serves the token-rate-limit 429 with live x-token-* until the window rolls over. A
stock client pointed at that mock sees the refusal with no SDK in the process — which
is how you test an agent that does not use this SDK at all.
The point
─────────
This needs no gateway, so it belongs in your unit tests. Demo 05 is the same idea
turned into a suite someone else can run against your agent without reading your code.
────────────────────────────────────────────────────────────────────────────────────────
```
```bash
python "demos/human-made/openai/03 - typed-refusals-simulated.py"
```
## Key code
One context manager, and the branch runs; `times=` counts calls and normal
service resumes after (narrative demo 04):
```python
with donkey.simulate(PIIDetected):
await agent.run("...") # fails as a real PIIDetected
with donkey.simulate(TokenBudgetExceeded, times=2):
await agent.run("a") # refused
await agent.run("b") # refused
await agent.run("c") # succeeds — the injection is spent
```
Looping over refusal types on a plain OpenAI client (OpenAI script 03):
```python
async with Donkey.from_env() as donkey:
client = donkey.openai()
for refusal in REFUSALS:
async with donkey.run(id=f"typed-refusals-{refusal.__name__}"):
# simulate() replays the captured gateway fixture in-process, so
# the refusal branch runs with no network and nothing to provoke.
with donkey.simulate(refusal):
try:
await client.responses.create(
model="gpt-4o",
input="Say hello in exactly three words.",
)
except openai.APIStatusError as err:
report(classify(err.response))
```
Scripting the running simulator instead, for a stock client (narrative demo 04,
act 7):
```bash
donkey mock --scenario pii_block:every=2 \
--scenario 'injection:on-pattern=ignore previous' \
--scenario budget:limit=200,window=5s,cost=80
```
`pii_block` fails every Nth call, `injection` matches request text, and
`budget` is a real wall-clock window that serves the token-rate-limit 429 on
exhaustion until the window rolls over.
`simulate()` only injects shapes that have a captured wire body.
`simulate(ToolInvocationError)` and `simulate(GatewayUnavailable)` raise
`ValueError` instead of inventing one — a transport failure has no HTTP
response to replay. To provoke `GatewayUnavailable`, point at a dead origin
(see [Typed refusals](https://docs.donkey-kit.dev/examples/general/typed-refusals.md)).
Because the injection sits on the transport, it also reaches framework objects
the SDK does not wrap — demo 04 drives `donkey.langgraph.chat_model("gpt-4o")`
through `simulate(PIIDetected)` and gets the same taxonomy back.
**Learn more:** [Local simulator](https://docs.donkey-kit.dev/simulator.md) · [Testing & conformance](https://docs.donkey-kit.dev/testing.md)
**Source:**
[narrative demo 04](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/04_simulate_refusals) ·
[script 03](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/03%20-%20typed-refusals-simulated.py)
---
Source: https://docs.donkey-kit.dev/examples/general/conformance.md
# Conformance suite
Four questions a team usually cannot answer about its own agent: does it retry
a budget refusal (it must not)? Does a typed refusal survive its error
handling? Does the run's correlation id reach its logs? Does it still work when
the gateway sends no budget headers? The conformance suite answers them without
reading your code — it swaps a fixture-serving transport underneath, calls
`agent.run(...)`, and watches the wire and the logs. So it grades behaviour, in
any framework, and it runs in your CI as a pytest plugin with no gateway.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 05 | The suite failing a naive agent, what each finding means, the fixed agent passing, a correct exemption, and two broken exemptions failing at collection time | Nothing (`[test]` + `[llm]`) |
## Run it
```bash
make demo N=05
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 05 — the conformance suite
Four questions about your agent that you cannot currently answer.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target offline — no gateway, no simulator, no credentials
output masking on
[1] An agent written the way people actually write them
for attempt in range(3): # retry on failure
try:
return await client.responses.create(...)
except openai.APIStatusError as exc:
if attempt == 2:
raise RuntimeError(...) from exc # friendly error
Nothing there is obviously wrong. Retrying is a sane default, wrapping errors keeps
stack traces out of the caller's face, and it logs what it is doing. Run the suite
against it.
pytest --donkey-conformance --agent=…:build_naive
FFF. [100%]
Donkey conformance
==================
FAIL Retries a budget refusal — retried a TokenBudgetExceeded 3× — a budget refusal is terminal; retrying only burns the same exhausted window
FAIL Swallows PII as a generic error — raised a bare RuntimeError — the PII refusal was swallowed as a generic error (bridge it with classify())
FAIL Propagates the correlation id — did not emit the correlation id in any log record — read it from current_correlation_id() and include it when you log
PASS Works without budget headers — completed a run when the gateway returned no budget headers
Donkey conformance: 1 passed, 3 failed
=========================== short test summary info ============================
FAILED ::donkey-conformance::retries_token_budget - Retries a budget refusal:...
FAILED ::donkey-conformance::swallows_pii_as_generic - Swallows PII as a gene...
FAILED ::donkey-conformance::correlation_id_propagated - Propagates the corre...
3 failed, 1 passed in 0.39s
exit status 1
What each failure actually means
────────────────────────────────
• Retried a TokenBudgetExceeded 3× — the window is already exhausted, so the retries
cannot succeed and the extra calls make the rate-limit situation worse for everyone
else on the same budget.
• Raised a bare RuntimeError for a PII block — the caller wanted to catch PIIDetected
and redact the flagged entities. It cannot, because the type was thrown away in the
name of a friendlier message.
• Never logged the correlation id — so when the platform team asks which gateway
request corresponds to this run, there is no answer.
[2] The same agent, after the findings
try:
return await client.responses.create(...)
except openai.APIStatusError as exc:
error = classify(exc.response)
log.warning("governed refusal", extra={
"correlation_id": current_correlation_id(),
"refusal": type(error).__name__,
})
raise error from exc # typed, and not retried
.... [100%]
Donkey conformance
==================
PASS Retries a budget refusal — issued one call and did not retry the budget refusal
PASS Swallows PII as a generic error — surfaced the refusal as a typed PIIDetected
PASS Propagates the correlation id — emitted the run's correlation id in its own logs
PASS Works without budget headers — completed a run when the gateway returned no budget headers
Donkey conformance: 4 passed, 0 failed
4 passed in 0.36s
exit status 0
[3] When an agent genuinely cannot pass, it says so out loud
Suppose the correlation finding is not fixable: your framework owns the HTTP transport
and gives you no per-request hook. That is a real limitation, so you assert it — with
a reason — and it becomes an `exempt` row rather than a failure. It is never a silent
skip, and the reason is meant to be published.
# exemptions.py
FRAMEWORK_LIMITS = {
"correlation_id_propagated": "This agent's framework owns the HTTP …",
}
pytest --donkey-conformance --agent=shipping_agent:build_naive \
--donkey-known-limitations=exemptions:FRAMEWORK_LIMITS
FF.. [100%]
Donkey conformance
==================
FAIL Retries a budget refusal — retried a TokenBudgetExceeded 3× — a budget refusal is terminal; retrying only burns the same exhausted window
FAIL Swallows PII as a generic error — raised a bare RuntimeError — the PII refusal was swallowed as a generic error (bridge it with classify())
EXEMPT Propagates the correlation id — This agent's framework owns the HTTP transport and offers no per-request context hook, so a run-scoped correlation id cannot reach the agent's logs.
PASS Works without budget headers — completed a run when the gateway returned no budget headers
Donkey conformance: 1 passed, 2 failed, 1 exempt
=========================== short test summary info ============================
FAILED ::donkey-conformance::retries_token_budget - Retries a budget refusal:...
FAILED ::donkey-conformance::swallows_pii_as_generic - Swallows PII as a gene...
2 failed, 2 passed in 0.35s
exit status 1
One row moved to EXEMPT with its reason attached. The other two findings are untouched
— an exemption excuses exactly what it names.
[4] And an exemption you get wrong fails the run, loudly
The mapping is validated at collection time, before any scenario runs. A typo'd
scenario name would otherwise exempt nothing while looking like it exempted something,
and an empty reason is a skip wearing a costume.
case a misspelled scenario name
ERROR: KNOWN_LIMITATIONS names unknown scenario 'retries_tokn_budget'; valid scenarios are ['correlation_id_propagated', 'retries_token_budget', 'swallows_pii_as_generic', 'works_without_budget_headers']
exit status 4
case an exemption with an empty reason
ERROR: KNOWN_LIMITATIONS['retries_token_budget'] must be a non-empty reason string — an asserted exemption, never a silent skip
exit status 4
The point
─────────
This is the deliverable, not our internal adapter matrix. It ships as a pytest plugin
so it runs in your CI, against your agent, in whatever framework you chose — and it
needs no gateway to do it.
────────────────────────────────────────────────────────────────────────────────────────
```
Against your own agent, the suite is a plain pytest invocation:
```bash
pip install "donkey-kit[test]"
pytest --donkey-conformance --agent=my_app.agent:build
```
## Key code
The agent under test is a factory the suite calls; it never reads the agent's
source. This is the fixed agent from `shipping_agent.py` — a refusal is
classified and escapes as its own type, it is never retried, and the
correlation id goes into the logs:
```python
class GovernedAgent:
def __init__(self, donkey: Donkey) -> None:
self._donkey = donkey
self._client = donkey.openai()
async def run(self, user_input: str) -> str:
log.info(
"handling request",
extra={"correlation_id": current_correlation_id(), "input": user_input[:40]},
)
try:
response = await self._client.responses.create(model="gpt-4o", input=user_input)
except openai.APIStatusError as exc:
error = classify(exc.response)
log.warning(
"governed refusal",
extra={
"correlation_id": current_correlation_id(),
"refusal": type(error).__name__,
},
)
raise error from exc
return getattr(response, "output_text", "")
def build_governed(donkey: Donkey) -> GovernedAgent:
return GovernedAgent(donkey)
```
When an agent genuinely cannot pass a scenario, you assert an exemption with a
reason (`exemptions.py`) and pass it on the command line:
```python
FRAMEWORK_LIMITS = {
"correlation_id_propagated": (
"This agent's framework owns the HTTP transport and offers no per-request "
"context hook, so a run-scoped correlation id cannot reach the agent's logs."
),
}
```
```bash
pytest --donkey-conformance --agent=shipping_agent:build_naive \
--donkey-known-limitations=exemptions:FRAMEWORK_LIMITS
```
An exemption becomes an `EXEMPT` row with its reason attached — never a
silent skip — and it excuses exactly the scenario it names. The mapping is
validated at collection time, so a misspelled scenario name or an empty
reason fails the run before any scenario executes.
The naive agent in the same file retries three times and re-raises a bare
`RuntimeError`. The suite flags the retried `TokenBudgetExceeded`, the lost
`PIIDetected` type, and the missing correlation id in its logs.
**Learn more:** [Testing & conformance](https://docs.donkey-kit.dev/testing.md)
**Source:**
[narrative demo 05](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/05_conformance)
---
Source: https://docs.donkey-kit.dev/examples/general/telemetry.md
# Telemetry
Platform teams ask for two things agent teams rarely deliver: a trace that
follows one logical run across every model call it fans out into, and spans in
the standard GenAI vocabulary so they land in existing dashboards. Both come
from the same client every request leaves through. Each governed call emits a
span carrying `gen_ai.*` and `donkey.*` attributes, `donkey.run(id=…)` ties a
whole run to one correlation id and one set of cost tags, and setting
`OTEL_EXPORTER_OTLP_ENDPOINT` is enough for `Donkey.from_env()` to install an
exporter.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 06 | One span per call with both namespaces and routing/usage, three calls under one run id and cost tags, a refused call as an `ERROR` span, and zero-config OTLP | `[otel]` (simulator) |
| OpenAI script 06 | A host-owned `TracerProvider` exporting over OTLP; Donkey rides it | Proxy credentials + an OTLP endpoint |
| OpenAI script 07 | Several `donkey.run(team=…, project=…)` blocks, including a refusal span | Proxy credentials + an OTLP endpoint |
| OpenAI script 10 | `Donkey.from_env()` installing OTLP itself when the env var is set | Proxy credentials |
## Run it
```bash
pip install "donkey-kit[otel]"
make demo N=06
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 06 — OTel GenAI spans and correlation ids
Standard GenAI telemetry and one trace per run, without instrumentation code.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target mock
proxy base_url http://127.0.0.1:8080/
output masking on
credentials fake — the simulator enforces no auth
[1] One governed call, one GenAI span — with no instrumentation code
# your usual OTel setup, then:
await client.responses.create(model=..., input=...)
span donkey.llm.chat
status UNSET
gen_ai.*
gen_ai.request.model gpt-4o
gen_ai.system openai
gen_ai.response.model gpt-5.1
gen_ai.usage.input_tokens 17
gen_ai.usage.output_tokens 51
donkey.*
donkey.routing.type ModelBased
donkey.routing.fallback False
donkey.usage.cached_tokens 0
donkey.usage.cache_write_tokens 0
donkey.usage.reasoning_tokens 0
donkey.policy.decision allow
donkey.budget.remaining 95000
donkey.cost.team platform
donkey.cost.env dev
donkey.correlation_id a8fdd39a06f145f3b20520cd9ee129fc
PASS gen_ai.prompt / gen_ai.completion are absent — capture is opt-in
The gen_ai.* keys are pinned to semantic-convention version 1.30.0. They are
transcribed in the SDK rather than imported from the semconv package, whose default
version drifts release to release — so what lands on your span is decided by a
reviewable edit, not by a transitive upgrade.
Prompt and completion stay off the span unless you set telemetry_capture_content=True
(or DONKEY_TELEMETRY_CAPTURE_CONTENT=1). The gateway masks PII in its logs; spans are
emitted upstream of that, so defaulting capture on would re-export the content the
platform just masked.
Routing and usage, on the same span
───────────────────────────────────
routing
gen_ai.request.model gpt-4o
gen_ai.response.model gpt-5.1
donkey.routing.type ModelBased
donkey.routing.fallback False
donkey.usage.*
donkey.usage.cached_tokens 0
donkey.usage.cache_write_tokens 0
donkey.usage.reasoning_tokens 0
gen_ai.response.model is what the gateway actually served. When it differs from
gen_ai.request.model, a failover happened — the fastest read on a latency spike.
donkey.routing.fallback is emitted even when False: 'we routed normally' is a signal,
not the absence of one. The same facts live on donkey.last_call without a span backend
(demo 10).
[2] One correlation id for a whole run, however many calls it makes
async with donkey.run(id=ticket.id, team="support", project="triage"):
await client.responses.create(...) # all three calls share
await client.responses.create(...) # one id, and the cost
await client.responses.create(...) # tags, on wire and spans
run id ticket-4417
spans emitted 3
distinct correlation ids 1
PASS all 3 spans carry the one run id
PASS run() overrode team/project; env inherited from from_env()
Nothing was threaded through the agent. The id is bound to the async context, and
tasks the framework spawns copy that context — so a LangGraph node running the model
on a child task is inside the same run without knowing the run exists. Concurrent runs
do not leak into each other, and nested run() blocks rebind then restore. Cost tags
ride the same context: run(team=..., project=...) overrides those dimensions for the
block and inherits the rest from from_env().
Two ids, two questions
──────────────────────
run header X-Correlation-Id
call header X-Donkey-Request-Id
The run id answers 'show me everything this ticket did'. The per-call id answers
'which one of those calls was this'. Both go out on every request, which is what lets
a line in your log join to the gateway's own record of the same call.
[3] A refusal is a failed span, not a successful-looking one
A span that ends OK on a request the gateway refused is worse than no span: it makes a
dashboard say everything is fine. So a refusal sets the span status to ERROR and
records what refused it.
span donkey.llm.chat
status ERROR
gen_ai.*
gen_ai.request.model donkey-sim/pii-detected
gen_ai.system openai
gen_ai.response.model gpt-5.1
donkey.*
donkey.routing.type ModelBased
donkey.routing.fallback False
donkey.policy.decision refuse
donkey.policy.type pii_detected
donkey.budget.remaining 1
donkey.cost.team platform
donkey.cost.env dev
donkey.correlation_id ticket-4417
donkey.policy.decision is 'refuse' and donkey.policy.type names the specific policy —
so a dashboard can separate 'the model failed' from 'governance said no', which are
very different operational stories.
The same two ids, on the exception
──────────────────────────────────
type PIIDetected
.correlation_id ticket-4417
.call_id c37c6199f07c43fe9f0555fd88cc4cdd
classify() read those back off the request the failed response came from, so the
exception you catch already carries the ids without you passing them in. Put
.correlation_id in the alert and the platform team can pull the gateway's record of
the same refusal.
[4] Zero-config OTLP: set the standard env var, or stay silent
# no Donkey-specific variable
export OTEL_EXPORTER_OTLP_ENDPOINT=https://…
donkey = Donkey.from_env() # installs OTLP behind a BatchSpanProcessor
# no endpoint → inert, silent, nothing connects
# DONKEY_TELEMETRY=false → opt out even if an endpoint is set
PASS no OTEL_EXPORTER_OTLP_ENDPOINT — export stayed inert and silent
Donkey.from_env() installs OTLP only when that standard env var is set. It will not
clobber a TracerProvider the host already installed — which is why this demo's in-
memory table still works. Opt out with DONKEY_TELEMETRY=false (or telemetry = false in
.donkey-kit.toml). Cost tags on donkey.run(team=..., project=...) are what let a
backend slice refusals, budget and latency by agent without another attribute
convention.
The point
─────────
Span name is 'donkey.llm.chat'. Nothing in the agent code above mentions OpenTelemetry
— the instrumentation hangs off the same transport hooks as the budget and the typed
refusals, which is why they all landed in one milestone rather than three.
Cost tags are the fixed four — team / project / env / enduser.id — set on from_env()
and overridable per donkey.run(). They land on donkey.cost.* whether or not the
gateway-side header names are verified yet. Routing (donkey.routing.*) and
cached/reasoning usage (donkey.usage.*) land on the same span. Zero-config OTLP is
shipped: OTEL_EXPORTER_OTLP_ENDPOINT, otherwise silent.
────────────────────────────────────────────────────────────────────────────────────────
```
```bash
python "demos/human-made/openai/06 - otel exporter simple.py" # proxy + OTEL_EXPORTER_OTLP_ENDPOINT / _HEADERS
python "demos/human-made/openai/07 - otel exporter advanced.py" # proxy + OTEL_EXPORTER_OTLP_ENDPOINT / _HEADERS
python "demos/human-made/openai/10 - zero-config-otlp.py" # proxy; set OTEL_EXPORTER_OTLP_ENDPOINT to export
```
Without `[otel]` installed, narrative demo 06 prints the install command and
exits cleanly. It installs an in-memory exporter so it can print the spans as a
table.
## Key code
One correlation id and one set of cost tags for a whole run (narrative demo 06,
act 2):
```python
async with donkey.run(id=ticket.id, team="support", project="triage"):
await client.responses.create(...) # all three calls share
await client.responses.create(...) # one id, and the cost
await client.responses.create(...) # tags, on wire and spans
```
Your own `TracerProvider`, with several runs and a refusal (OpenAI script 07):
```python
provider = TracerProvider(resource=Resource.create({"service.name": "donkey-dev-kit"}))
provider.add_span_processor(SimpleSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(provider)
donkey = Donkey.from_env()
client = donkey.openai(sync=True)
with donkey.run(id="agent-greeter", team="cx", project="welcome"):
reply = client.responses.create(model=MODEL, input="Say hello in exactly three words.")
print("greeter:", reply.output_text)
with donkey.run(id="agent-support", team="cx", project="tickets"):
try:
client.responses.create(model=MODEL, input=PII_PROMPT)
print("support: no refusal")
except openai.APIStatusError as err:
error = classify(err.response)
print("support:", type(error).__name__, getattr(error, "entities", None))
provider.force_flush()
donkey.close()
```
Zero-config export (OpenAI script 10) needs no provider setup at all:
```python
donkey = Donkey.from_env()
client = donkey.openai(sync=True)
with donkey.run(id="otel-zero-config", team="cx", project="welcome"):
reply = client.responses.create(model="gpt-4o", input="Say hello in exactly three words.")
print(reply.output_text)
```
`gen_ai.prompt` and `gen_ai.completion` stay off the span unless you set
`telemetry_capture_content=True` (or `DONKEY_TELEMETRY_CAPTURE_CONTENT=1`).
Spans are emitted upstream of the gateway's PII mask, so capturing by default
would re-export content the platform just masked.
- **Two ids per request.** The run id (`X-Correlation-Id`) answers "everything
this ticket did"; the per-call id (`X-Donkey-Request-Id`) answers "which call
was this". A caught refusal carries both as `.correlation_id` and `.call_id`.
- **Refusals are failed spans.** A refused call sets the span status to
`ERROR` and records `donkey.policy.decision=refuse` with the specific
`donkey.policy.type`.
- **Cost tags are a fixed four** — `team`, `project`, `env`, `enduser.id` —
set on `from_env()` and overridable per `run()`, landing on `donkey.cost.*`.
- **Export is opt-in.** No endpoint means inert and silent;
`DONKEY_TELEMETRY=false` opts out even when one is set, and a host
`TracerProvider` is never replaced.
**Learn more:** [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md)
**Source:**
[narrative demo 06](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/06_telemetry) ·
[script 06](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/06%20-%20otel%20exporter%20simple.py) ·
[script 07](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/07%20-%20otel%20exporter%20advanced.py) ·
[script 10](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/10%20-%20zero-config-otlp.py)
---
Source: https://docs.donkey-kit.dev/examples/general/framework-objects.md
# Framework objects & model handles
No adapter returns a wrapper. `donkey.langgraph.chat_model(...)` hands back a
real `langchain_openai.ChatOpenAI`, so everything your framework can do with a
model still works and nothing new appears in your stack traces. LangGraph is
the one deep, conformance-gated adapter; the other seven are supported at
`connection_kwargs()` — the SDK gives you the base URL, headers and client
configuration, and you pass them to the framework's own constructor. The
companion example covers model handles: `resolve()` gives a local capability
handle, and `list_models(live=True)` raises a `ConfigError` explaining that the
proxy has no catalog endpoint rather than guessing one.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 07 | `resolve()` capability handles, `list_models(live=True)` raising `ConfigError`, and config validation listing every missing field at once | Nothing |
| Narrative demo 08 | One factory call per framework and what came back, then `connection_kwargs()` for the shallow adapters | Nothing — objects are constructed, no network calls |
## Run it
```bash
make demo N=07
make demo N=08
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 07 — model handles and honest gaps
What the SDK does when the platform has no endpoint for what you asked.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target offline — no gateway, no simulator, no credentials
output masking on
[1] resolve() — a local capability handle for a known model id
handle = donkey.llm.resolve("gpt-4o")
handle.capabilities
gpt-4o ModelCapabilities(function_calling=True, vision=True, json_output=True, is_heuristic=True)
gpt-4o-mini ModelCapabilities(function_calling=True, vision=True, json_output=True, is_heuristic=True)
o3 ModelCapabilities(function_calling=True, vision=False, json_output=False, is_heuristic=True)
claude-3-5-sonnet ModelCapabilities(function_calling=True, vision=False, json_output=False, is_heuristic=True)
something-unknown-9 ModelCapabilities(function_calling=True, vision=False, json_output=False, is_heuristic=True)
These are heuristics derived from the model id, and the SDK says so rather than
implying it asked the gateway. They are useful for routing decisions in your own code;
they are not a governed catalog.
[2] list_models(live=True) — the honest failure
await donkey.llm.list_models(live=True)
raised ConfigError
The governed LLM proxy exposes no /models endpoint (GET /models → 404, verified
docs/verified-apis.md §2): it only routes requests carrying `model` in the body. Live
model listing is not available from the proxy. Use resolve(model_id) or source the
catalog from Exchange/provider config.
PASS It names the verified absence and points at the alternative, instead of guessing a /models path that would 404 in your sandbox.
[3] The same discipline applied to configuration
The most common reason someone abandons an SDK in the first five minutes is the one-
missing-variable-per-run loop: fix a variable, re-run, discover the next one. So
validation reports everything at once.
DonkeyConfig(llm_proxy_url="https://…").validated(need="llm")
Configuration for 'llm' is incomplete. Missing:
- llm_proxy_client_id (env DONKEY_LLM_PROXY_CLIENT_ID)
- llm_proxy_client_secret (env DONKEY_LLM_PROXY_CLIENT_SECRET)
Set them via kwargs, environment variables, or .donkey-kit.toml (secrets in .donkey-kit.local.toml).
Two missing fields, one error, each naming the environment variable that sets it. And
note the LLM proxy credential is validated separately from the Anypoint control-plane
one — a developer may legitimately have proxy access and no Exchange access.
When the failure is live rather than a missing variable — wrong URL, wrong
credentials, or a model the allow-list does not include — `donkey doctor` is the CLI
that distinguishes those three. It reuses the same remediation strings the typed
errors carry (demo 02).
────────────────────────────────────────────────────────────────────────────────────────
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 08 — native framework objects
One deep adapter, seven at connection_kwargs(), and no wrappers anywhere.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target offline — no gateway, no simulator, no credentials
output masking on
[1] One call per framework, and what came back
langgraph deep — the conformance-gated adapter
donkey.langgraph.chat_model(…)
PASS returned langchain_openai.chat_models.base.ChatOpenAI
adk connection_kwargs()
donkey.adk.model(…)
not installed: pip install "donkey-kit[adk]"
strands connection_kwargs()
donkey.strands.model(…)
not installed: pip install "donkey-kit[strands]"
agent_framework connection_kwargs()
donkey.agent_framework.chat_client(…)
not installed: pip install "donkey-kit[agent_framework]"
openai_agents no connection_kwargs() — builds its own client
donkey.openai_agents.model(…)
not installed: pip install "donkey-kit[openai-agents]"
anthropic connection_kwargs()
donkey.anthropic.client()
not installed: pip install "donkey-kit[anthropic]"
crewai connection_kwargs()
donkey.crewai.llm(…)
not installed: pip install "donkey-kit[crewai]"
llamaindex connection_kwargs()
donkey.llamaindex.llm(…)
not installed: pip install "donkey-kit[llamaindex]"
[2] connection_kwargs() — the surface that actually carries the roster
kwargs = donkey.strands.connection_kwargs()
SomeFrameworkModel(model="gpt-4o", **kwargs)
langgraph.connection_kwargs()
base_url https://demo-gateway.example.invalid/openai-sdk/
api_key client-id-enforced
default_headers.client_id (36 chars)
default_headers.client_secret (40 chars)
http_async_client
http_client
max_retries 0
use_responses_api True
Same base URL, same verified client_id / client_secret pair, handed to the framework's
own constructor. Bringing a framework up to the deep bar is demand-driven and happens
one at a time, so this is not a stepping stone that everything is queued behind — it
is the supported surface.
LangGraph is the only adapter held to the conformance bar. It sets
use_responses_api=True so ChatOpenAI calls /responses rather than its
/chat/completions default: /responses is the raw client's route and the only one the
local simulator serves.
What is and is not verified here
────────────────────────────────
The proxy contract these objects are configured against is live-verified: the base URL
shape, the credential header pair, the rejection shapes. The adapters themselves are
held to three bars (docs/verified-apis.md §8):
• Conformance-tested against the simulator: the raw client and LangGraph.
• Signature-confirmed offline: every other adapter, ADK's model() included. The SDK's
scripts/verify_frameworks.py builds each native object against the installed
framework; Agent Framework's model= kwarg (not model_id) is confirmed that way
against 1.19.0.
• Live-verified: ADK's gemini(), through a Format=Gemini proxy.
The adapters build the framework's native object directly. They refuse with 'blocked
on verification' only when the installed framework version lacks the class or field
the adapter depends on: an Agent Framework class rename, or ADK's gemini() before
google-adk 2.4.
────────────────────────────────────────────────────────────────────────────────────────
```
Narrative demo 08 uses obviously-fake config, so it needs no credentials.
Frameworks that are not installed are reported with their exact `pip install`
line.
## Key code
The roster narrative demo 08 walks — attribute on `Donkey`, factory method, and
depth:
```python
ROSTER = [
("langgraph", "chat_model", True, "deep — the conformance-gated adapter"),
("adk", "model", True, "connection_kwargs()"),
("strands", "model", True, "connection_kwargs()"),
("agent_framework", "chat_client", True, "connection_kwargs()"),
("openai_agents", "model", True, "no connection_kwargs() — builds its own client"),
("anthropic", "client", False, "connection_kwargs()"),
("crewai", "llm", True, "connection_kwargs()"),
("llamaindex", "llm", True, "connection_kwargs()"),
]
```
For the shallow adapters, `connection_kwargs()` is the whole supported surface:
```python
kwargs = donkey.strands.connection_kwargs()
SomeFrameworkModel(model="gpt-4o", **kwargs)
```
Model handles and the missing catalog (narrative demo 07):
```python
handle = donkey.llm.resolve("gpt-4o")
handle.capabilities
await donkey.llm.list_models(live=True) # raises ConfigError
```
And config validation reports every missing field in one error, each naming
the environment variable that sets it:
```python
DonkeyConfig(llm_proxy_url="https://…").validated(need="llm")
```
`donkey.openai_agents` is the OpenAI Agents SDK adapter; `donkey.openai()` is
the raw OpenAI client factory. The LangGraph adapter sets
`use_responses_api=True`, so `ChatOpenAI` calls the `/responses` route. The
adapters build the native object directly. They refuse with "blocked on
verification" only when the installed framework version lacks the class or
field the adapter depends on (for example, `gemini()` on google-adk older than
2.4). See [Model access](https://docs.donkey-kit.dev/frameworks.md#supported-frameworks) for each adapter's
verification status.
`resolve()` capabilities are heuristics derived from the model id, not a
governed catalog. The gateway returns 404 for `GET /models` because
model-based routing only routes requests that already carry `model` in the
body. When a live call fails — wrong URL, wrong credentials, or a model the
allow-list does not include — [`donkey doctor`](https://docs.donkey-kit.dev/cli.md) tells those apart.
**Learn more:** [Model access](https://docs.donkey-kit.dev/frameworks.md) · [LangGraph](https://docs.donkey-kit.dev/frameworks/langgraph.md) · [CLI & decorators](https://docs.donkey-kit.dev/cli.md)
**Source:**
[narrative demo 07](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/07_model_handles) ·
[narrative demo 08](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/08_framework_objects)
---
Source: https://docs.donkey-kit.dev/examples/general/langgraph-agent.md
# LangGraph agent
A real multi-step agent: the model decides to call two tools, the tools
return, and the model composes an answer. Every model call in that loop goes
through the governed proxy, and the object driving it is LangChain's own
`ChatOpenAI`, not a wrapper. The only DDK lines are the one that builds the
model, `donkey.run(id=…)` around the loop, `typed_refusals()` so a proxy 403
comes out of `astream` as `PIIDetected` rather than a framework-wrapped error,
and `@donkey.tool` on the two functions. Governance sits at the boundary, not
in the agent's control flow.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 09 | `donkey.langgraph.chat_model()`, a `create_agent` loop calling two tools, then the proxy's token window, `last_call` from the run scope and the registered tools | Live credentials + `[langgraph]` |
## Run it
```bash
make demo N=09 # needs live credentials
```
This example needs a live gateway. The local simulator replays a captured
`/responses` completion and will not decide to call tools, so there is no
offline version. Without credentials it exits cleanly with setup guidance.
The refusal path *can* run offline: [Simulating refusals](https://docs.donkey-kit.dev/examples/general/simulating-refusals.md)
drives the same `ChatOpenAI` through `donkey.simulate()`.
## Key code
The tools are plain LangChain tools, marked for the SDK's registry:
```python
@tool
@Donkey.tool
def check_inventory(sku: str) -> str:
"""Return the units in stock and warehouse for a product SKU."""
return INVENTORY.get(sku, "unknown SKU")
@tool
@Donkey.tool
def get_price(sku: str) -> str:
"""Return the list price for a product SKU."""
return PRICES.get(sku, "unknown SKU")
```
The model and the governed loop:
```python
async with Donkey.from_env() as donkey:
model = donkey.langgraph.chat_model(MODEL)
agent = create_agent(model, tools=[check_inventory, get_price])
async with donkey.run(id="sku-lookup"):
with donkey.langgraph.typed_refusals():
async for chunk in agent.astream(
{"messages": [("user", QUESTION)]}, stream_mode="updates"
):
...
budget = donkey.budget
last = donkey.last_call
```
After the loop, `donkey.budget` is the proxy's token window for this client
id, read in-band from every model call in the loop. It is shared, not a
per-run total, and it stays unobserved on a proxy without a token rate-limit
policy.
`donkey.last_call` reads `UNOBSERVED` here, by design. It is scoped per
`asyncio` task, so parallel calls never overwrite each other's record, and
LangGraph makes each model call on its own task, so the record never reaches
the caller's scope. On a direct call it is populated (see
[Gateway identity](https://docs.donkey-kit.dev/examples/general/gateway-identity.md)); a run-level record of every call is tracked
in [#613](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/613).
The adapter targets the
`/responses` route (`use_responses_api=True`), the same one `donkey.openai()`
uses. `DEMO_MODEL` defaults to `gpt-4o-mini` in this example; set it to a model
your proxy routes.
If the gateway is unavailable, [Framework objects](https://docs.donkey-kit.dev/examples/general/framework-objects.md)
constructs the same real framework objects with no network.
**Learn more:** [LangGraph](https://docs.donkey-kit.dev/frameworks/langgraph.md) · [Model access](https://docs.donkey-kit.dev/frameworks.md)
**Source:**
[narrative demo 09](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/09_langgraph_agent)
---
Source: https://docs.donkey-kit.dev/examples/general/gateway-identity.md
# Gateway identity (last_call)
A refusal already tells you which gateway said no. `donkey.last_call` is the
success-path counterpart: after every governed call it records which gateway
served the request, what it actually routed to, and what the call cost —
without parsing headers or running a span backend. That matters because a
silent model substitution is otherwise invisible to your cost model and your
evals, and reading only `total_tokens` misses cached and reasoning tokens. If
your assumptions are pinned to one model, `on_model_substitution="raise"`
turns a swap into a hard `ModelSubstituted` error.
| Example | Shows | Needs |
| --- | --- | --- |
| Narrative demo 10 | A cold `UNOBSERVED` record, one call populating identity, routing and usage, the substitution flag, and `on_model_substitution="raise"` | Nothing (simulator) |
| OpenAI script 08 | The full `last_call` record on a live call, then `on_model_substitution="raise"` | Proxy credentials |
## Run it
```bash
make demo N=10
```
```text
════════════════════════════════════════════════════════════════════════════════════════
Demo 10 — last_call, routing, and per-call usage
The success-path counterpart to a typed refusal: who served this, what they served, and
what it cost.
════════════════════════════════════════════════════════════════════════════════════════
Run context
───────────
target mock
proxy base_url http://127.0.0.1:8080/
output masking on
credentials fake — the simulator enforces no auth
[1] A new process has not observed a call yet — and it says so
status unobserved
observed False
available True
request_id None
PASS UNOBSERVED — not None, not 0, not 'unknown'. A cold read is a named state.
A bare None would be a lie of omission: you could not tell 'the gateway sent no id'
from 'we never saw a response'. Budget uses the same honesty rule for an unobserved
window (demo 03). UNAVAILABLE is the third state, for adapters that never route
through our transport — ADK via LiteLLM, CrewAI via its native OpenAI provider,
default_headers-only LlamaIndex, or Agent Framework (observes_last_call = False).
Those surfaces report UNAVAILABLE by name rather than looking like a cold read.
[2] One governed call, and the record is the success-path counterpart
await client.responses.create(model=..., input=...)
donkey.last_call.served_model
donkey.last_call.total_tokens
donkey.last_call.substituted
status observed
observed True
request_id req_f85003861d5348c9a1d152c276082b07
requested_model gpt-4o
served_model gpt-5.1
served_provider openai
routing_type ModelBased
fallback False
substituted True
input_tokens 17
output_tokens 51
total_tokens 68
cached_tokens 0
cache_write_tokens 0
reasoning_tokens 0
PASS OBSERVED — the SDK saw the response, even if some fields stayed None
request_id is the gateway's own id (x-request-id) — quote it in a ticket. It is the
same field classify() puts on a DonkeyError after a refusal, now present on the 200 as
well. api_instance_id and environment_id are parsed from x-envoy-decorator-operation;
they are masked in this output.
[3] Routing, fallback, and the cost-relevant token counts
What the gateway did with the request
─────────────────────────────────────
requested gpt-4o
served openai/gpt-5.1
routing_type ModelBased
fallback False
substituted True
PASS substituted — asked for gpt-4o, gateway served gpt-5.1
Against the simulator this is the captured happy-path fixture talking: it was recorded
against gpt-5.1, and we asked for a different id. That is not a live failover — and it
is exactly the mismatch last_call is for. A silent substitution is otherwise invisible
to your cost model, your eval, and your latency dashboard.
What this call cost
───────────────────
input / output / total 17 / 51 / 68
cached_tokens 0
cache_write_tokens 0
reasoning_tokens 0
cached_tokens are billed at the cached rate; reasoning_tokens are output the developer
never sees. Reading only total_tokens draws the wrong conclusion about both cost and
latency. An absent count is None, never 0 — 0 here means the gateway reported zero,
which is a different statement. These are per-call; donkey.budget is the shared window
(demo 03).
The SDK never double-retries a fallback. It retries 502/503/504 with backoff, but a
503 the gateway already marked as a failover is left alone — a second recovery layer
stacked on a working first one just multiplies latency against an outage the gateway
already handled.
[4] Opt in, and a substitution is a hard error instead of a flag
donkey = Donkey.from_env(on_model_substitution="raise")
# raises ModelSubstituted when served_model != requested_model
Off by default: the call succeeds and last_call.substituted is True. Raise is for
callers whose eval, cost model and token assumptions are pinned to one model.
ModelSubstituted is deliberately not a PolicyViolation — the request was neither
refused nor failed, it succeeded against a model you did not choose. Same shape as
BudgetReserveReached: a client-side signal you opted into.
PASS ModelSubstituted — the 200 never reached the caller
requested_model gpt-4o
served_model gpt-5.1
served_provider openai
request_id req_f85003861d5348c9a1d152c276082b07
last_call.substituted True
The record still populated — observe happens before the raise — so a handler that
decides to accept the served completion can read last_call the same way. The exception
also carries the response.
The point
─────────
The refusal path already told you which gateway said no. The success path now tells
you which gateway said yes, what it actually served, and what that call cost — without
a span backend, without parsing headers, and without a second accessor for routing or
usage.
────────────────────────────────────────────────────────────────────────────────────────
```
```bash
python "demos/human-made/openai/08 - last-call.py" # needs proxy credentials
```
## Key code
Reading the record after a call (OpenAI script 08):
```python
donkey = Donkey.from_env()
client = donkey.openai(sync=True)
print("before any call ", donkey.last_call.status.value)
reply = client.responses.create(model="gpt-4o", input="Say hello in exactly three words.")
last = donkey.last_call
print("status ", last.status.value)
print("request_id ", last.request_id)
print("requested_model ", last.requested_model)
print("served_model ", last.served_model)
print("served_provider ", last.served_provider)
print("routing_type ", last.routing_type)
print("fallback ", last.fallback)
print("substituted ", last.substituted)
print("total_tokens ", last.total_tokens)
print("cached_tokens ", last.cached_tokens)
print("reasoning_tokens", last.reasoning_tokens)
```
Opting in to a hard error. The OpenAI client wraps the transport error, so the
typed `ModelSubstituted` is on `__cause__`:
```python
strict = Donkey.from_env(on_model_substitution="raise")
strict_client = strict.openai(sync=True)
try:
strict_client.responses.create(model="gpt-4o", input="Say hello in exactly three words.")
print("NO RAISE substituted", strict.last_call.substituted)
except Exception as err:
hit = err if isinstance(err, ModelSubstituted) else err.__cause__
if not isinstance(hit, ModelSubstituted):
raise
print("requested_model ", hit.requested_model)
print("served_model ", hit.served_model)
```
- **Three named states, never a bare `None`.** `UNOBSERVED` is a cold read;
`OBSERVED` means the SDK saw a response (fields may still be `None` if the
gateway said nothing); `UNAVAILABLE` marks adapter surfaces that never route
through the SDK's transport.
- **`request_id` is the upstream provider's id, passed through by the
gateway** — `x-request-id` for OpenAI, `x-amzn-requestid` for Amazon Bedrock,
`apim-request-id` for Azure OpenAI, `request-id` for a native Anthropic
proxy. Quote it to the provider's support team.
It is the same field `classify()` puts on a refusal, now present on a 200 too,
and `None` on a route whose provider forwards no id.
- **Absent counts are `None`, never `0`.** `cached_tokens` are billed at the
cached rate; `reasoning_tokens` are output you never see.
- **`ModelSubstituted` is not a `PolicyViolation`.** The request succeeded,
against a model you did not choose. `last_call` is still populated before the
raise.
- **A `provider/` prefix is not a substitution.** On a model-based routing
proxy the gateway reports the served model without the `provider/` prefix;
`substituted` ignores a prefix that names the served provider (see
[Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md#two-behaviours-worth-knowing)).
The simulator's captured success response was recorded against `gpt-5.1`, so
asking for any other model id shows up as a substitution in narrative demo
10. That is the fixture, not a live failover — and it is exactly the
mismatch `last_call` exists to surface.
## Semantic routing: the matched topic and score
The call above went through a **model-based** proxy, where `routing_type` reads
`ModelBased`. A **semantic-routing** proxy instead classifies each prompt by
meaning and routes it to the matched topic's provider and model — there
`routing_type` reads `Semantic`, and `last_call` carries two more fields the
model-based path leaves `None`:
- `matched_topic` — which topic the prompt matched (e.g. `Finance`).
- `routing_score` — how close that match was, a bare `0.xx` similarity score.
Both come from the live-verified, semantic-only
`x-llm-proxy-semantic-routing-success` response header. The four routing fields
you already read (`routing_type`, `fallback`, `served_provider`, `served_model`)
are emitted identically to the model-based case, so the rest of the record reads
the same way — only these two are added.
You can exercise this branch offline: point the simulator at the captured
`Semantic` response by requesting the `donkey-sim/success-semantic` model id
(the same sentinel mechanism the [simulator](https://docs.donkey-kit.dev/simulator.md) uses to force a
refusal shape, here forcing a happy-path variant).
```python
client = donkey.openai(sync=True)
client.responses.create(
model="donkey-sim/success-semantic",
input="How does compound interest work?",
)
last = donkey.last_call
print("routing_type ", last.routing_type) # Semantic
print("matched_topic", last.matched_topic) # Finance
print("routing_score", last.routing_score) # 0.62
```
`matched_topic` and `routing_score` are `None` on a model-based proxy — the
`x-llm-proxy-semantic-routing-success` header is semantic-only. An unparseable
message leaves each field `None` rather than guessing a value.
**Learn more:** [Feature overview](https://docs.donkey-kit.dev/feature-overview.md) · [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md)
**Source:**
[narrative demo 10](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/claude-made/10_last_call) ·
[script 08](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/blob/main/demos/human-made/openai/08%20-%20last-call.py)
---
Source: https://docs.donkey-kit.dev/examples/openai.md
# OpenAI
`donkey.openai()` returns the stock `openai.AsyncOpenAI` (or `OpenAI` with
`sync=True`) on the SDK's governed transport. Because the SDK owns every
request, everything works here: `last_call`, the budget, spans, correlation
ids, typed refusals and `simulate()`. The route is `/responses`, which is
live-verified on the DDK proxies. This is the reference suite — the other
framework pages are measured against it.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-responses-no-gw.py` | Stock OpenAI, no SDK — the baseline | `OPENAI_API_KEY` |
| 02 | `basic-responses-gw.py` | The same call governed, then `last_call` | Proxy credentials |
| 03 | `typed-refusals-simulated.py` | Five captured refusals replayed in-process | Nothing — any placeholder values |
| 04 | `typed-refusals-live.py` | Four real refusals on a blocking client | Proxy + PII and token-rate policies |
| 05 | `budget_and_pacing.py` | `pace(reserve=)` stopping a call locally, `wait_for_reset()` | Proxy + token-rate policy |
| 06 | `otel exporter simple.py` | Donkey riding a host `TracerProvider` | Proxy + `[otel]` + OTLP endpoint |
| 07 | `otel exporter advanced.py` | Three runs with different cost tags, one refused | Proxy + `[otel]` + OTLP + PII policy |
| 08 | `last-call.py` | The full `last_call` record and `ModelSubstituted` | Proxy credentials |
| 09 | `governed-and-tool.py` | `@donkey.governed` and `@donkey.tool` | Nothing |
| 10 | `zero-config-otlp.py` | `Donkey.from_env()` installing OTLP itself | Proxy, optional OTLP endpoint |
| 11 | `gateway-unavailable.py` | A dead origin raising `GatewayUnavailable` | Nothing |
| 12 | `streaming.py` | `stream=True`, usage after the terminal event | Proxy credentials (live only) |
| 13 | `regex-and-content-safety.py` | Regex prompt guard and Azure Content Safety | Proxy + both policies |
| 14 | `jwt-wallet.py` | `llm_proxy_auth="jwt"` with a rotating token | Wallet-backed proxy + IdP JWT |
| 15 | `start-gateway.py` | The simulator on a real port, hit with `httpx` | `[local]` |
| 16 | `bedrock-guardrails.py` | Bedrock Guardrails as `ContentSafetyBlocked` | Proxy + Bedrock Guardrails policy |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,local]"
python -m pip install -e "../donkey-development-kit/python[otel]" # 06, 07, 10 only
set -a; source .env.local; set +a # proxy credentials
```
The scripts ask for `gpt-4o`. The provisioned DDK proxies route
`gpt-5-mini` — change the model string to one your proxy routes, or expect
a routing refusal or `ModelSubstituted`.
## 01 — Stock OpenAI, no gateway
```bash
export OPENAI_API_KEY=…
python "demos/human-made/openai/01 - basic-responses-no-gw.py"
```
The baseline every other script is compared to: `openai.AsyncOpenAI()` with no
`base_url`, talking to api.openai.com. It prints a three-word greeting and
nothing else — there is no SDK here.
## 02 — The same call, governed
```bash
python "demos/human-made/openai/02 - basic-responses-gw.py"
```
One line changes: the client comes from `donkey.openai()`. The call is
identical, and afterwards `donkey.last_call` says what actually happened.
```python
async with Donkey.from_env() as donkey:
client = donkey.openai() # the native AsyncOpenAI, on the governed transport
response = await client.responses.create(
model="gpt-4o",
input="Say hello in exactly three words.",
)
print(response.output_text)
last = donkey.last_call
print("last_call.status ", last.status.value)
print("last_call.served_model", last.served_model)
print("last_call.total_tokens", last.total_tokens)
print("last_call.substituted ", last.substituted)
```
**You should see:** the reply, then `last_call.status observed`, the served
model, total tokens and `substituted False` — or `True` if the proxy served a
different model than you asked for.
## 03 — Typed refusals, simulated
```bash
python "demos/human-made/openai/03 - typed-refusals-simulated.py"
```
`donkey.simulate(...)` replays the captured gateway fixture for each refusal
in-process, so the refusal branch runs with no network and no policy to
provoke. Each error is run through `classify()` and printed with its policy,
remediation, correlation id (the `donkey.run(id=…)` around it), call id and
request id. It ends by asking for `simulate(GatewayUnavailable)`, which
refuses: there is no response to replay for a request that never arrived.
```python
for refusal in REFUSALS:
async with donkey.run(id=f"typed-refusals-{refusal.__name__}"):
with donkey.simulate(refusal):
try:
await client.responses.create(model="gpt-4o", input="...")
except openai.APIStatusError as err:
report(classify(err.response))
```
```text
PIIDetected (policy: pii-detection)
Request contains PII data: [
{
"pii_type": "Email",
"value": "john.doe@example.com",
"start": 12,
"end": 32
}
]
remediation The PII-detection policy blocked this request because the prompt (or completion) contained personally identifiable information. Remove or redact the flagged values, or relax the policy's entity list / action in API Manager.
correlation_id typed-refusals-PIIDetected
call_id 60dfc226ece343f3b266ac13cab99edf
request_id None
entities Email
PromptInjectionBlocked (policy: prompt-injection-protection)
Request blocked by the injection-protection policy (400).
remediation The prompt-injection-protection policy flagged this request as a prompt-injection attempt. Review and sanitise the untrusted input in the prompt, or adjust the policy's sensitivity in API Manager.
correlation_id typed-refusals-PromptInjectionBlocked
call_id b317e2e1535a429da48541292612ba00
request_id None
ContentSafetyBlocked (policy: content-safety)
Request blocked by Azure Content Safety (severity_hate, severity_violence) (403).
remediation The Azure Content Safety content-moderation policy blocked this request. Revise the flagged content, or adjust the policy's categories / severity thresholds in API Manager.
correlation_id typed-refusals-ContentSafetyBlocked
call_id df36871604c9408788a1aa37874365a7
request_id None
categories severity_hate, severity_violence
TokenBudgetExceeded (policy: token-rate-limit)
Token rate limit or budget exceeded (429).
remediation A token-rate-limit policy exhausted the budget window. Wait for it to reset (see retry_after / x-token-reset) or request an increase in API Manager.
correlation_id typed-refusals-TokenBudgetExceeded
call_id e7580506e64c44cd8f54ef00a5134188
request_id None
retry_after 41.728
PolicyViolation (policy: unknown)
Request refused by a gateway policy; shape unconfirmed (status 400). It matched no documented rejection contract.
remediation This refusal matched no documented rejection shape, so its contract is unconfirmed (#184, #253). It is terminal and was NOT retried. Please file an issue on the donkey-development-kit repo with the response status, headers and body (all carried on this exception's .response) so the shape can be typed.
correlation_id typed-refusals-PolicyViolation
call_id 6910b17935bf4820a651d7fc7794eac5
request_id None
simulate(GatewayUnavailable)
simulate() cannot inject GatewayUnavailable: no captured fixture maps back to it via classify(). Supported: AgentKilled, AuthError, ContentSafetyBlocked, ModelNotRoutable, PIIDetected, PolicyViolation, PromptInjectionBlocked, TokenBudgetExceeded, UpstreamModelError, UpstreamRequestError.
```
`request_id` is `None` throughout because a simulated response carries no
gateway request id; `call_id` is the SDK's own id and is always set.
## 04 — Typed refusals, live
```bash
python "demos/human-made/openai/04 - typed-refusals-live.py"
```
Four real refusals on the blocking client (`donkey.openai(sync=True)`, no
`async`): `UpstreamRequestError` from a model that does not exist,
`PIIDetected` from a contact record, `TokenBudgetExceeded` from a long answer,
and `AuthError` from a second `Donkey` built with deliberately wrong
credentials via `cfg.with_overrides(...)`. This script has its own
`MODEL = "gpt-4-turbo-2024-04-09"` constant at the top.
```python
with donkey.run(id="live-refusals-PIIDetected"):
try:
raw = client.responses.with_raw_response.create(model=MODEL, input=PII_PROMPT)
except openai.APIStatusError as err:
error = classify(err.response)
print(f" REFUSED {type(error).__name__} (HTTP {err.response.status_code})")
print(f" entities {getattr(error, 'entities', None)}")
except openai.APIConnectionError as err:
print(f" UNREACHABLE {type(err.__cause__ or err).__name__}")
else:
print(f" NO REFUSAL (HTTP {raw.status_code})")
```
**Needs:** `llm-pii-detection-policy` with `Email` among its entities and
action `Reject` (the default, `Log`, lets it through), and
`llm-token-rate-limit` with a tiny `maximumTokens` for the budget case. The
upstream and auth cases need nothing extra.
**You should see:** a `REFUSED (HTTP n)` block per case. A case whose
policy is not applied prints `NO REFUSAL` with the served model or the budget
window instead — that is the proxy telling the truth, not the script failing.
## 05 — Budget & pacing
```bash
python "demos/human-made/openai/05 - budget_and_pacing.py"
```
`donkey.budget.pace(reserve=0.99999)` means "keep almost the whole window".
The first call is let through because nothing has been observed yet; its
response carries the token window in-band. The second `pace()` sees that any
usage breaches the reserve and raises `BudgetReserveReached` locally — the
request never leaves the process. The script then observes a crafted
one-second window and `wait_for_reset()` sleeps until it resets.
```python
async with donkey.budget.pace(reserve=0.99999):
response = await client.responses.create(model="gpt-4o", input="...")
print("after request 1, budget remaining is", donkey.budget.remaining)
try:
async with donkey.budget.pace(reserve=0.99999):
await client.responses.create(model="gpt-4o", input="...")
except BudgetReserveReached as exc:
print("stopped locally [in-script]", exc.fraction_used, exc.reserve)
await donkey.budget.wait_for_reset()
```
**Needs:** a proxy that sends the token-window header (for example
`ddk-token-rate-limit`). **You should see:** the budget fields after request 1,
then `stopped locally [in-script]` and `Ended script after reset`. Without the
header every field is `None` and the second call simply goes through.
## 06 — OpenTelemetry, host-owned provider
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT=https:// OTEL_EXPORTER_OTLP_HEADERS=…
python "demos/human-made/openai/06 - otel exporter simple.py"
```
The process installs its own `TracerProvider` with an OTLP exporter before
building `Donkey`, so the SDK rides that provider instead of installing one.
Two calls run inside `donkey.run(id="otel-demo")`.
```python
provider = TracerProvider(resource=Resource.create({"service.name": "donkey-dev-kit"}))
provider.add_span_processor(SimpleSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(provider)
donkey = Donkey.from_env()
client = donkey.openai(sync=True)
with donkey.run(id="otel-demo"):
client.responses.create(model="gpt-4o", input="Say hello in exactly three words.")
client.responses.create(model="gpt-4o", input="Say goodbye in exactly three words.")
```
**You should see:** two replies in the terminal, and in your collector two
spans sharing the correlation id `otel-demo` under `service.name
donkey-dev-kit`. Nothing about spans is printed locally.
## 07 — OpenTelemetry with cost tags
```bash
python "demos/human-made/openai/07 - otel exporter advanced.py"
```
Three `donkey.run(...)` scopes, each tagged with a different `team` /
`project`: a greeter that succeeds, a support agent whose prompt trips the PII
policy, and a researcher on a second model. Budget and `last_call` are printed
between them.
```python
with donkey.run(id="agent-greeter", team="cx", project="welcome"):
reply = client.responses.create(model=MODEL, input="Say hello in exactly three words.")
with donkey.run(id="agent-support", team="cx", project="tickets"):
try:
client.responses.create(model=MODEL, input=PII_PROMPT)
except openai.APIStatusError as err:
print("support:", type(classify(err.response)).__name__)
```
**You should see:** `greeter:` with its reply, budget and `last_call`;
`support: PIIDetected ['Email']` (or `no refusal` without the policy);
`researcher:`; then the final budget. In the collector: three traces, one of
them an `ERROR` span naming the PII policy.
## 08 — The full `last_call` record
```bash
python "demos/human-made/openai/08 - last-call.py"
```
Before any call `last_call` is `unobserved`. After one, it carries status,
request id, requested versus served model, provider, routing type, fallback,
substitution, and total, cached and reasoning tokens. The second half builds
a client with `on_model_substitution="raise"`, which turns a silent model swap
into an exception.
```python
strict = Donkey.from_env(on_model_substitution="raise")
try:
strict.openai(sync=True).responses.create(model="gpt-4o", input="...")
print("NO RAISE substituted", strict.last_call.substituted)
except Exception as err:
hit = err if isinstance(err, ModelSubstituted) else err.__cause__
print("requested_model ", hit.requested_model)
print("served_model ", hit.served_model)
```
**You should see:** `before any call unobserved`, then the populated record.
The second half prints `NO RAISE` when the proxy served the model you asked
for, or `RAISED … ModelSubstituted` with both model ids.
## 09 — `@donkey.governed` and `@donkey.tool`
```bash
python "demos/human-made/openai/09 - governed-and-tool.py"
```
No gateway, decorators only. `@donkey.governed(team=…, project=…)` gives each
call of the handler its own run id and cost tags, and clears them afterwards.
`@donkey.tool` records a callable in the tool registry without wrapping it —
the function object is unchanged — and rejects a tool with no docstring,
because an undescribed tool is useless to a model.
```python
@donkey.governed(team="support", project="triage")
def handle_ticket(ticket: str) -> str:
print("run id inside", current_correlation_id())
print("cost tags ", current_cost_tags())
return ticket
@donkey.tool
def lookup_sku(sku: str) -> str:
"""Return stock for a product SKU."""
return "42"
```
```text
run id before None
run id inside 4a2ff59b241a493ca27886058e97e52e
cost tags CostTags(team='support', project='triage', env=None, enduser_id=None)
run id inside 13e6ef2bf82e4b83b4cb96c387699484
cost tags CostTags(team='support', project='triage', env=None, enduser_id=None)
run id after None
lookup_sku('AF-1001') 42
same function object True
registered lookup_sku (sku: str) -> str
undescribed tool @donkey.tool requires a docstring on 'undescribed': an undescribed tool is useless to a model and to the registry (#200). Add a short description of what the tool does.
```
## 10 — Zero-config OTLP
```bash
python "demos/human-made/openai/10 - zero-config-otlp.py" # inert
OTEL_EXPORTER_OTLP_ENDPOINT=https:// python "demos/human-made/openai/10 - zero-config-otlp.py"
```
With no `TracerProvider` in the process, `Donkey.from_env()` installs an OTLP
exporter itself when `OTEL_EXPORTER_OTLP_ENDPOINT` is set, and stays silent
when it is not. `DONKEY_TELEMETRY=false` opts out even with an endpoint set.
**You should see:** the endpoint (or `(unset — Donkey.from_env() will not
install an exporter)`), the reply, and `last_call observed `.
## 11 — Gateway unavailable
```bash
python "demos/human-made/openai/11 - gateway-unavailable.py"
```
A governed client aimed at `127.0.0.1:9`, where nothing listens, with
`max_retries=0` and a two-second timeout. There is no HTTP response to
classify, so the transport raises the typed `GatewayUnavailable` — a long-running
agent can tell "lost the gateway" from a policy refusal without matching raw
`httpx` exceptions. The script reads it from the exception or from
`__cause__`, whichever holds it.
```python
try:
client.responses.create(model="gpt-4o", input="hello")
except Exception as err:
hit = err if isinstance(err, GatewayUnavailable) else err.__cause__
if isinstance(hit, GatewayUnavailable):
print("base_url ", hit.base_url)
print(hit.remediation)
```
```text
raised GatewayUnavailable
cause GatewayUnavailable
base_url http://127.0.0.1:9
request_id None
call_id 787cba4a82ed421bb49c786049401575
The gateway could not be reached and no HTTP response came back. The three usual causes: (1) the host is unreachable — DNS failure or the gateway is down; (2) the configured base URL is wrong; or (3) network egress to the gateway is blocked — a firewall or air-gapped environment. Run `donkey doctor` to diagnose connectivity, and check `base_url` on this error against your gateway's address.
```
`request_id` is `None` because no response ever came back; `call_id` is set
because the SDK assigned it before sending.
## 12 — Streaming
```bash
python "demos/human-made/openai/12 - streaming.py"
```
`responses.create(stream=True)`; the `response.output_text.delta` events are
joined, and `last_call` usage is read after the terminal event.
**Live only** — the simulator's stream is truncated SSE, so terminal usage
would be a lie there. **You should see:** the reply, `observed`, the served
model and total tokens.
## 13 — Regex prompt guard and Azure Content Safety
```bash
python "demos/human-made/openai/13 - regex-and-content-safety.py"
```
Two live guardrails. A "reveal your system prompt" prompt trips
`regex-prompt-guard` and comes back as `PromptInjectionBlocked`; a hateful
prompt trips Azure Content Safety and comes back as `ContentSafetyBlocked`
with its categories.
**Needs:** a proxy with both policies (for example `ddk-injection-guard` and
`ddk-azure-content-safety`; point `DONKEY_LLM_PROXY_URL` at each in turn if
they are separate proxies). **You should see:** the type and policy per case,
or `NO REFUSAL`.
## 14 — JWT wallet auth
```bash
export DONKEY_LLM_PROXY_WALLET_CLIENT_ID=… DONKEY_LLM_JWT=…
python "demos/human-made/openai/14 - jwt-wallet.py"
```
`llm_proxy_auth="jwt"` sends `X-Client-Id` plus a JWT and no client secret.
The JWT is never a config field: it is passed as an auth provider, so it can
rotate. JWT mode is async-only.
```python
cfg = DonkeyConfig.from_env().with_overrides(
llm_proxy_auth="jwt",
llm_proxy_wallet_client_id=os.environ["DONKEY_LLM_PROXY_WALLET_CLIENT_ID"],
)
async with Donkey(cfg, llm_auth=StaticToken(os.environ["DONKEY_LLM_JWT"])) as donkey:
client = donkey.openai()
```
**Needs:** a wallet-backed proxy URL in `DONKEY_LLM_PROXY_URL` and a JWT from
your IdP. **You should see:** the reply, `observed` and the served model. An
invalid JWT is an `AuthError`.
## 15 — The simulator on a real port
```bash
python "demos/human-made/openai/15 - start-gateway.py"
```
`start_gateway()` boots the SDK's local simulator on an ephemeral port — the
same fixtures as 03, but served over HTTP, so any client can hit it.
`set_scenarios("pii_block:every=1")` makes every request the captured PII 403,
and a stock `httpx.post` gets it.
```python
gw = start_gateway()
gw.set_scenarios("pii_block:every=1")
response = httpx.post(f"{gw.url}/responses", json={"model": "gpt-4o", "input": "hello"}, headers=...)
print(type(classify(response)).__name__, response.status_code)
print("requests", gw.requests_received)
gw.close()
```
```text
PIIDetected 403
requests 1
```
## 16 — Amazon Bedrock Guardrails
```bash
python "demos/human-made/openai/16 - bedrock-guardrails.py"
```
Bedrock Guardrails rejecting a prompt. It is the same `ContentSafetyBlocked`
class as Azure in 13, with the vendor in the message. Bedrock's own id rides
`x-amzn-requestid`, not `x-request-id`, and the SDK still fills `request_id`.
**Needs:** a proxy with Bedrock Guardrails applied (for example
`ddk-bedrock-guardrails`). **You should see:** `ContentSafetyBlocked
content-safety ['content_filter']`, the message, `x-request-id None` and a
populated `request_id`.
**Learn more:** [Typed refusals](https://docs.donkey-kit.dev/errors.md) · [Budget & pacing](https://docs.donkey-kit.dev/budget.md) ·
[Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md) · [Local simulator](https://docs.donkey-kit.dev/simulator.md)
**Source:**
[`demos/human-made/openai/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/openai)
---
Source: https://docs.donkey-kit.dev/examples/langgraph.md
# LangGraph
`donkey.langgraph("…")` returns a real `langchain_openai.ChatOpenAI` with
`use_responses_api=True`, so calls go to the live-verified `/responses` route
on the governed transport. It is the one deep, conformance-gated adapter:
`donkey.langgraph.typed_refusals()` turns the openai error a graph node raises
back into the SDK's typed refusal.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-no-gw.py` | Stock `ChatOpenAI`, no SDK — the baseline | `OPENAI_API_KEY` |
| 02 | `basic-gw.py` | The governed model, LangChain's own usage | Proxy credentials |
| 03 | `typed-refusals-simulated.py` | Five refusals out of `create_agent`, typed | Nothing — any placeholder values |
| 04 | `typed-refusals-live.py` | `PIIDetected`, `UpstreamRequestError`, `AuthError` live | Proxy + PII policy |
| 05 | `agent-and-tool.py` | The run id reaching a tool inside the graph | Proxy credentials (live only) |
| 06 | `otel exporter simple.py` | Donkey riding a host `TracerProvider` | Proxy + `[otel]` + OTLP endpoint |
| 07 | `streaming.py` | `astream` chunks and terminal usage | Proxy credentials (live only) |
| 08 | `gateway-unavailable.py` | `GatewayUnavailable` two causes down | Nothing |
| 09 | `start-gateway.py` | An agent loop over the local simulator | `[local]` |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,local,langgraph]" "langchain>=1.0"
python -m pip install -e "../donkey-development-kit/python[otel]" # 06 only
set -a; source .env.local; set +a
```
`langchain>=1.0` provides `create_agent` (03, 04, 05, 09); it is not part of
any `donkey-kit` extra.
The scripts use `ainvoke` / `astream`. Sync `invoke()` / `stream()` are
governed the same way, through `ChatOpenAI`'s `http_client`, except in `jwt`
mode, where they raise `ConfigError`.
## 01 — Stock ChatOpenAI, no gateway
```bash
export OPENAI_API_KEY=…
python "demos/human-made/langgraph/01 - basic-no-gw.py"
```
`ChatOpenAI(model="gpt-4o", use_responses_api=True)` straight to
api.openai.com. It prints a three-word greeting.
## 02 — The governed model
```bash
python "demos/human-made/langgraph/02 - basic-gw.py"
```
```python
async with Donkey.from_env() as donkey:
model = donkey.langgraph("gpt-4o") # native ChatOpenAI on /responses
reply = await model.ainvoke("Say hello in exactly three words.")
print(reply.text)
print("model_name ", reply.response_metadata.get("model_name"))
print("usage ", reply.usage_metadata)
print("last_call ", donkey.last_call.status.value)
```
**You should see:** the reply, `model_name`, LangChain's `usage` dict, and
`last_call unobserved`. That last line is expected: LangChain drives the call
on its own task, and `last_call` is scoped per task, so it never reaches the
caller. Use `usage_metadata` instead.
## 03 — Typed refusals, simulated
```bash
python "demos/human-made/langgraph/03 - typed-refusals-simulated.py"
```
`donkey.simulate(...)` replays five captured refusals in-process while a
`create_agent` graph runs. Without help, the error would surface wrapped by
the graph; `typed_refusals()` re-raises it as the SDK type, so a plain `except
PolicyViolation` catches all five.
```python
agent = create_agent(donkey.langgraph("gpt-4o"), tools=[])
for refusal in REFUSALS:
async with donkey.run(id=f"lg-simulated-{refusal.__name__}"):
with donkey.simulate(refusal):
try:
with donkey.langgraph.typed_refusals():
await agent.ainvoke({"messages": [{"role": "user", "content": "hello"}]})
except PolicyViolation as error:
print(type(error).__name__, error.policy, error.correlation_id)
```
```text
PIIDetected pii-detection lg-simulated-PIIDetected
PromptInjectionBlocked prompt-injection-protection lg-simulated-PromptInjectionBlocked
ContentSafetyBlocked content-safety lg-simulated-ContentSafetyBlocked
TokenBudgetExceeded token-rate-limit lg-simulated-TokenBudgetExceeded
PolicyViolation unknown lg-simulated-PolicyViolation
```
The correlation id on each error is the `donkey.run(id=…)` around the graph —
it survived LangGraph's own task scheduling.
## 04 — Typed refusals, live
```bash
python "demos/human-made/langgraph/04 - typed-refusals-live.py"
```
Three cases, each with its own `Donkey`: `PIIDetected` from a contact record,
`UpstreamRequestError` from a model that does not exist, and `AuthError` from
deliberately wrong credentials. All three are caught as `DonkeyError` out of
`typed_refusals()`.
**Needs:** `llm-pii-detection-policy` with `Email` and action `Reject` for the
first case; the other two need nothing extra. **You should see:** ` ->
` per case, or ` NO REFUSAL` when the policy
is not applied.
## 05 — Agent and tool
```bash
python "demos/human-made/langgraph/05 - agent-and-tool.py"
```
`create_agent` with one tool (`lookup_sku`) inside `donkey.run(...)`. LangGraph
copies the context into every node, so the tool reads the same run id the
caller set — every model call and tool call in the loop is correlated.
```python
@tool
def lookup_sku(sku: str) -> str:
"""Return stock for a product SKU."""
print("tool sees run id", current_correlation_id())
return "42"
async with donkey.run(id="lg-ticket-4417", team="support", project="triage"):
out = await agent.ainvoke({"messages": [{"role": "user", "content": "How many AF-1001 are in stock?"}]})
```
**Live only** — the simulator replays one completion and never decides to call
a tool. **You should see:** `run id lg-ticket-4417`, then `tool sees run id
lg-ticket-4417`, the final answer and the message count.
## 06 — OpenTelemetry
```bash
export OTEL_EXPORTER_OTLP_ENDPOINT=https:// OTEL_EXPORTER_OTLP_HEADERS=…
python "demos/human-made/langgraph/06 - otel exporter simple.py"
```
The host owns the `TracerProvider` and Donkey rides it, as in
[OpenAI 06](https://docs.donkey-kit.dev/examples/openai.md#06--opentelemetry-host-owned-provider). Two
`ainvoke` calls run in one tagged `donkey.run(...)`. **You should see:** two
replies locally, and two spans with one correlation id in the collector.
## 07 — Streaming
```bash
python "demos/human-made/langgraph/07 - streaming.py"
```
`astream` chunks are added together; usage lands on the terminal event.
```python
final = None
async for chunk in model.astream("Say hello in exactly three words."):
final = chunk if final is None else final + chunk
print(final.text)
print("usage", final.usage_metadata)
```
**Live only** — `ChatOpenAI` rejects the simulator's truncated SSE. **You
should see:** the reply and a `usage` dict.
## 08 — Gateway unavailable
```bash
python "demos/human-made/langgraph/08 - gateway-unavailable.py"
```
A governed model aimed at `127.0.0.1:9`, where nothing listens.
`typed_refusals()` passes transport failures through untouched, so the script
walks `__cause__` until it finds `GatewayUnavailable`.
```python
try:
with donkey.langgraph.typed_refusals():
await donkey.langgraph("gpt-4o").ainvoke("hello")
except Exception as err:
hit = err
while hit is not None and not isinstance(hit, GatewayUnavailable):
hit = hit.__cause__
```
```text
raised GatewayUnavailable
cause GatewayUnavailable
base_url http://127.0.0.1:9
The gateway could not be reached and no HTTP response came back. The three usual causes: (1) the host is unreachable — DNS failure or the gateway is down; (2) the configured base URL is wrong; or (3) network egress to the gateway is blocked — a firewall or air-gapped environment. Run `donkey doctor` to diagnose connectivity, and check `base_url` on this error against your gateway's address.
```
## 09 — An agent loop over the local simulator
```bash
python "demos/human-made/langgraph/09 - start-gateway.py"
```
No gateway and no credentials. `start_gateway()` boots the simulator with
`pii_block:every=2`, so every second request is the captured PII 403. A
`create_agent` loop handles two support tickets: the first gets the replayed
completion, the second is refused and caught as a typed `PolicyViolation`.
```python
gw = start_gateway()
gw.set_scenarios("pii_block:every=2")
cfg = DonkeyConfig(llm_proxy_url=gw.url, llm_proxy_client_id=..., llm_proxy_client_secret=...)
async with Donkey(cfg) as donkey:
agent = create_agent(donkey.langgraph("gpt-4o"), tools=[], system_prompt="Reply in one sentence.")
for ticket in TICKETS:
try:
with donkey.langgraph.typed_refusals():
out = await agent.ainvoke({"messages": [{"role": "user", "content": ticket}]})
print("ok ", out["messages"][-1].text[:60])
except PolicyViolation as error:
print("refused", type(error).__name__, error.entities)
```
```text
ok A sleepy unicorn named Luma painted soft silver stars across
refused PIIDetected ['Email']
requests 2
```
The "ok" text is the simulator's canned completion, not a reply to the ticket.
**Learn more:** [LangGraph](https://docs.donkey-kit.dev/frameworks/langgraph.md) · [LangGraph agent](https://docs.donkey-kit.dev/examples/general/langgraph-agent.md)
**Source:**
[`demos/human-made/langgraph/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/langgraph)
---
Source: https://docs.donkey-kit.dev/examples/openai-agents.md
# OpenAI Agents SDK
The agent is given `OpenAIResponsesModel(model=…,
**donkey.openai_agents.connection_kwargs())`. `connection_kwargs()` is one
key, `openai_client`, a governed `AsyncOpenAI`, so the agent stays on the
live-verified `/responses` route and the SDK owns the transport.
`donkey.openai_agents.model()` builds a chat-completions model instead. The
`Runner` re-raises the openai error unchanged, so `classify()` types it.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `agent-and-tool.py` | An `Agent` with a `@function_tool` run by `Runner.run` | Proxy credentials (live only) |
| 02 | `typed-refusals-simulated.py` | Four simulated refusals through `Runner.run` | Nothing — any placeholder values |
| 03 | `start-gateway.py` | An agent over the local simulator | `[local]` |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,local,openai-agents]"
set -a; source .env.local; set +a
```
Every script calls `set_tracing_disabled(True)` so the Agents SDK does not
## 01 — Agent and tool
```bash
python "demos/human-made/openai-agents/01 - agent-and-tool.py"
```
A `weather` agent with one `@function_tool` (`get_weather`), run by
`Runner.run` inside `donkey.run(...)`. The runner makes two model calls: one
that decides to call the tool, one that writes the answer.
```python
model = OpenAIResponsesModel(model="gpt-4o", **donkey.openai_agents.connection_kwargs())
agent = Agent(
name="weather",
instructions="Answer in one short sentence. Use the tool for weather.",
model=model,
tools=[get_weather],
)
async with donkey.run(id="agents-sdk-weather"):
result = await Runner.run(agent, "What is the weather in Paris?")
print(result.final_output)
print("model calls ", result.context_wrapper.usage.requests)
print("total_tokens", result.context_wrapper.usage.total_tokens)
print("last_call ", donkey.last_call.status.value)
```
**Live only** — the simulator never calls tools. **You should see:** the final
answer, `model calls 2`, total tokens from the Agents SDK's own usage, and
`last_call observed`. That record is the second call only: the runner sends the
first turn from a task of its own, where the caller can't see it (see
[`donkey.last_call`](https://docs.donkey-kit.dev/frameworks/openai.md#donkeylast_call)). The SDK usage line
counts both calls.
## 02 — Typed refusals, simulated
```bash
python "demos/human-made/openai-agents/02 - typed-refusals-simulated.py"
```
`donkey.simulate(...)` replays each captured refusal while `Runner.run`
drives the agent. The runner lets the `openai.APIStatusError` through, and
`classify()` turns its response into the typed refusal.
```python
for refusal in REFUSALS:
async with donkey.run(id=f"agents-simulated-{refusal.__name__}"):
with donkey.simulate(refusal):
try:
await Runner.run(agent, "hello")
except openai.APIStatusError as err:
error = classify(err.response)
print(type(error).__name__, error.policy, error.correlation_id)
```
```text
Error getting response
Error getting response
Error getting response
Error getting response
PIIDetected pii-detection agents-simulated-PIIDetected
PromptInjectionBlocked prompt-injection-protection agents-simulated-PromptInjectionBlocked
ContentSafetyBlocked content-safety agents-simulated-ContentSafetyBlocked
TokenBudgetExceeded token-rate-limit agents-simulated-TokenBudgetExceeded
```
The `Error getting response` lines are the Agents SDK logging each failed
model call to stderr before it re-raises; they are not from DDK.
## 03 — An agent over the local simulator
```bash
python "demos/human-made/openai-agents/03 - start-gateway.py"
```
No gateway and no credentials. `start_gateway()` with `pii_block:every=2`
over two support tickets: the first gets the simulator's canned completion,
the second is the captured PII 403.
```python
gw = start_gateway()
gw.set_scenarios("pii_block:every=2")
cfg = DonkeyConfig(llm_proxy_url=gw.url, llm_proxy_client_id=..., llm_proxy_client_secret=...)
async with Donkey(cfg) as donkey:
agent = Agent(
name="triage",
instructions="Reply in one sentence.",
model=OpenAIResponsesModel(model="gpt-4o", **donkey.openai_agents.connection_kwargs()),
)
for ticket in TICKETS:
try:
result = await Runner.run(agent, ticket)
print("ok ", result.final_output[:60])
except openai.APIStatusError as err:
error = classify(err.response)
print("refused", type(error).__name__, error.entities)
```
```text
Error getting response
ok A sleepy unicorn named Luma painted soft silver stars across
refused PIIDetected ['Email']
requests 2
```
**Learn more:** [OpenAI Agents SDK](https://docs.donkey-kit.dev/frameworks/openai.md)
**Source:**
[`demos/human-made/openai-agents/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/openai-agents)
---
Source: https://docs.donkey-kit.dev/examples/agent-framework.md
# Microsoft Agent Framework
`donkey.agent_framework.chat_client("…")` builds an `OpenAIChatClient`
(verified against 1.19.0, where the keyword is `model=`) that calls
`/responses`. It sends through the SDK's shared HTTP client, so the run id
reaches the proxy and `last_call` is set in the task that made the call.
Refusals come back typed — Agent Framework wraps the
openai error in `ChatClientException`, and the response rides on `__cause__`.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-gw.py` | An `Agent` on the governed chat client | Proxy credentials |
| 02 | `typed-refusals-live.py` | `PIIDetected`, `UpstreamRequestError`, `AuthError` | Proxy + PII policy |
| 03 | `start-gateway.py` | Two tickets over the local simulator | `[local]` |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,local,agent_framework]"
set -a; source .env.local; set +a
```
## 01 — A governed agent
```bash
python "demos/human-made/agent-framework/01 - basic-gw.py"
```
```python
donkey = Donkey.from_env()
agent = Agent(
client=donkey.agent_framework.chat_client("gpt-4o"),
name="greeter",
instructions="Answer in one short sentence.",
)
result = asyncio.run(agent.run("Say hello in exactly three words."))
print(result.text)
print("total tokens", result.usage_details["total_token_count"])
print("last_call ", donkey.last_call.status.value, donkey.last_call.surface)
```
**You should see:** the reply, `total tokens` from Agent Framework's
`usage_details`, and `last_call unavailable …`. `asyncio.run(...)` runs the
call in its own context, so the script's read is a cold one, and a cold read
on this adapter reports `unavailable` rather than `unobserved` ([#740](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/740)). Read
usage from the framework here, or read `last_call` inside the coroutine that
made the call.
## 02 — Typed refusals, live
```bash
python "demos/human-made/agent-framework/02 - typed-refusals-live.py"
```
Three cases, each with its own `Donkey`: a contact record for `PIIDetected`, a
model that does not exist for `UpstreamRequestError`, and wrong credentials
for `AuthError`. The one Agent Framework-specific line is where the response
is read from:
```python
try:
asyncio.run(agent.run(prompt))
print(name, "NO REFUSAL")
except ChatClientException as err:
error = classify(err.__cause__.response)
print(name, "->", type(error).__name__, getattr(error, "entities", None))
```
**Needs:** `llm-pii-detection-policy` with `Email` and action `Reject` for the
first case. **You should see:** ` -> ` per case, or
` NO REFUSAL`.
## 03 — Two tickets over the local simulator
```bash
python "demos/human-made/agent-framework/03 - start-gateway.py"
```
No gateway and no credentials. `start_gateway()` with `pii_block:every=2`:
the first ticket gets the simulator's canned completion, the second is the
captured PII 403, unwrapped from `ChatClientException` and classified.
```python
gw = start_gateway()
gw.set_scenarios("pii_block:every=2")
donkey = Donkey(DonkeyConfig(llm_proxy_url=gw.url, llm_proxy_client_id=..., llm_proxy_client_secret=...))
for ticket in TICKETS:
agent = Agent(client=donkey.agent_framework.chat_client("gpt-4o"), instructions="Reply in one sentence.")
try:
print("ok ", asyncio.run(agent.run(ticket)).text[:60])
except ChatClientException as err:
error = classify(err.__cause__.response)
print("refused", type(error).__name__, error.entities)
```
```text
ok A sleepy unicorn named Luma painted soft silver stars across
refused PIIDetected ['Email']
requests 2
```
`TypeError … model_id` means an older Agent Framework; the scripts target
the 1.19.0 `model=` keyword.
**Learn more:** [MS Agent Framework](https://docs.donkey-kit.dev/frameworks/agent-framework.md)
**Source:**
[`demos/human-made/agent-framework/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/agent-framework)
---
Source: https://docs.donkey-kit.dev/examples/strands.md
# Strands
Strands Agents with `OpenAIModel(client=donkey.openai(), model_id=…)`. Because
the governed client is passed in, the SDK owns the transport: run id,
`last_call` and typed refusals all work. The calls go to the proxy's
`/chat/completions` route, so use a proxy whose upstream serves that route. The
scripts use `gpt-4o`. If you point them at a model the proxy routes to Gemini,
add `stream=False` to `OpenAIModel(...)`: Strands streams by default, and a
Gemini-routed stream has no chunk deltas, so every turn fails (see
[Streaming chat completions to a Gemini upstream](https://docs.donkey-kit.dev/reference/unsupported-boundary.md)).
**Why not `donkey.strands.model()`?** Strands opens and closes an OpenAI
client per request from `client_args`. That closes the shared transport
after the first call, and the second fails with a `ConfigError` saying the
HTTP client is closed. A
pre-built `client=` is reused and left open.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-gw.py` | A governed agent, Strands usage and `last_call` | Proxy credentials |
| 02 | `agent-and-tool.py` | The run id reaching a tool across two model calls | Proxy credentials (live only) |
| 03 | `typed-refusals-simulated.py` | Three simulated refusals, typed | Nothing — any placeholder values |
| 04 | `typed-refusals-live.py` | `PIIDetected`, `UpstreamRequestError`, `AuthError` | Proxy + PII policy |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm]" "strands-agents[openai]"
set -a; source .env.local; set +a
```
## 01 — A governed agent
```bash
python "demos/human-made/strands/01 - basic-gw.py"
```
```python
async with Donkey.from_env() as donkey:
model = OpenAIModel(client=donkey.openai(), model_id="gpt-4o")
agent = Agent(model=model, callback_handler=None, system_prompt="Answer in one short sentence.")
result = await agent.invoke_async("Say hello in exactly three words.")
print(str(result).strip())
print("usage ", result.metrics.accumulated_usage)
print("status ", donkey.last_call.status.value)
print("served_model", donkey.last_call.served_model)
```
**You should see:** a one-sentence answer, Strands' accumulated usage, then
`status observed` and the served model — unlike LangGraph, Strands calls the
model on the caller's task, so `last_call` is populated.
## 02 — Agent and tool
```bash
python "demos/human-made/strands/02 - agent-and-tool.py"
```
One Strands `@tool` (`lookup_sku`) inside `donkey.run(...)`. The tool loop is
two model calls, and `client=donkey.openai()` keeps the transport open across
both. The tool prints the run id it sees.
```python
@tool
def lookup_sku(sku: str) -> str:
"""Return stock for a product SKU."""
print("tool sees run id", current_correlation_id())
return "42"
async with donkey.run(id="strands-ticket-4417", team="support", project="triage"):
result = await agent.invoke_async("How many AF-1001 are in stock?")
```
**Live only.** **You should see:** `tool sees run id strands-ticket-4417`, the
answer, `model calls 2`, and `last_call observed `.
## 03 — Typed refusals, simulated
```bash
python "demos/human-made/strands/03 - typed-refusals-simulated.py"
```
`donkey.simulate(...)` replays each captured refusal while a Strands agent
runs; the `openai.APIStatusError` comes through unchanged and `classify()`
types it.
```python
model = OpenAIModel(client=donkey.openai(), model_id="gpt-4o")
for refusal in REFUSALS:
async with donkey.run(id=f"strands-simulated-{refusal.__name__}"):
with donkey.simulate(refusal):
try:
await Agent(model=model, callback_handler=None).invoke_async("hello")
except openai.APIStatusError as err:
error = classify(err.response)
print(type(error).__name__, error.policy, error.correlation_id)
```
```text
PIIDetected pii-detection strands-simulated-PIIDetected
PromptInjectionBlocked prompt-injection-protection strands-simulated-PromptInjectionBlocked
ContentSafetyBlocked content-safety strands-simulated-ContentSafetyBlocked
```
`TokenBudgetExceeded` is deliberately absent. Strands retries a 429 itself
(`ModelThrottledException`), so the one simulated 429 is absorbed and the
retry succeeds. Against a live proxy the same retry means a real budget 429
takes a while to surface.
## 04 — Typed refusals, live
```bash
python "demos/human-made/strands/04 - typed-refusals-live.py"
```
Three cases, each with its own `Donkey`: a contact record for `PIIDetected`, a
model that does not exist for `UpstreamRequestError`, and wrong credentials
for `AuthError`.
**Needs:** `llm-pii-detection-policy` with `Email` and action `Reject` for the
first case. **You should see:** ` -> ` per case, or
` NO REFUSAL`.
**Learn more:** [Strands](https://docs.donkey-kit.dev/frameworks/strands.md)
**Source:**
[`demos/human-made/strands/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/strands)
---
Source: https://docs.donkey-kit.dev/examples/crewai.md
# CrewAI
CrewAI 1.x with `donkey.crewai.llm("…")`. `openai/` models go through CrewAI's
native OpenAI provider to the proxy's `/chat/completions` route. CrewAI owns the transport: the
credentials go on the wire, but there is **no run id and no `last_call`**.
Refusals are still typed, because the native provider keeps the openai error's
response.
Both scripts need a live gateway; there is no offline CrewAI script.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-gw.py` | `llm.call(...)`, then a one-agent crew | Proxy credentials |
| 02 | `typed-refusals-live.py` | `PIIDetected`, `UpstreamRequestError`, `AuthError` | Proxy + PII policy |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,crewai]"
set -a; source .env.local; set +a
```
Both scripts set `CREWAI_TRACING_ENABLED=false` unless you set it yourself.
Otherwise CrewAI stops on an interactive "view your traces?" prompt and the
script appears to hang with no output.
## 01 — A direct call and a crew
```bash
python "demos/human-made/crewai/01 - basic-gw.py"
```
```python
donkey = Donkey.from_env()
llm = donkey.crewai.llm("gpt-4o")
print(llm.call("Say hello in exactly three words."))
agent = Agent(role="Greeter", goal="Greet people", backstory="Friendly and brief.", llm=llm)
out = agent.kickoff("Say goodbye in exactly three words.")
print(out.raw)
print("usage ", out.usage_metrics)
print("last_call", donkey.last_call.status.value, donkey.last_call.surface)
```
**You should see:** the greeting, the crew's answer, CrewAI's
`usage_metrics`, and `last_call unavailable …` — CrewAI owns the transport, so
read usage from `usage_metrics`.
## 02 — Typed refusals, live
```bash
python "demos/human-made/crewai/02 - typed-refusals-live.py"
```
Three cases, each with its own `Donkey`: a contact record for `PIIDetected`, a
model that does not exist for `UpstreamRequestError`, and wrong credentials
for `AuthError`.
```python
for name, case_cfg, model, prompt in CASES:
donkey = Donkey(case_cfg)
try:
donkey.crewai.llm(model).call(prompt)
print(name, "NO REFUSAL")
except openai.APIStatusError as err:
error = classify(err.response)
print(name, "->", type(error).__name__, getattr(error, "entities", None))
donkey.close()
```
**Needs:** `llm-pii-detection-policy` with `Email` and action `Reject` for the
first case. **You should see:** ` -> ` per case, or
` NO REFUSAL`. A `404` means the proxy's upstream has no
`/chat/completions` route.
**Learn more:** [CrewAI](https://docs.donkey-kit.dev/frameworks/crewai.md)
**Source:**
[`demos/human-made/crewai/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/crewai)
---
Source: https://docs.donkey-kit.dev/examples/llamaindex.md
# LlamaIndex
LlamaIndex with `donkey.llamaindex.llm("…")`, an `OpenAILike`. The adapter sets
`is_chat_model=True` — the `OpenAILike` default of `False` hits
`/completions` — so calls go to the proxy's `/chat/completions` route. Sync
and async calls send through the
SDK's HTTP clients, so the run id reaches the proxy and `last_call` is
populated. `OpenAILike` re-raises the openai error unchanged, so `classify()`
types refusals.
Both scripts need a live gateway; there is no offline LlamaIndex script.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-gw.py` | `complete(...)` then `chat(...)` | Proxy credentials |
| 02 | `typed-refusals-live.py` | `PIIDetected`, `UpstreamRequestError`, `AuthError` | Proxy + PII policy |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,llamaindex]"
set -a; source .env.local; set +a
```
## 01 — `complete()` and `chat()`
```bash
python "demos/human-made/llamaindex/01 - basic-gw.py"
```
```python
donkey = Donkey.from_env()
llm = donkey.llamaindex.llm("gpt-4o")
print(llm.complete("Say hello in exactly three words.").text)
reply = llm.chat([ChatMessage(role="user", content="Say goodbye in exactly three words.")])
print(reply.message.content)
print("total tokens", reply.raw.usage.total_tokens)
print("last_call ", donkey.last_call.status.value, donkey.last_call.surface)
```
**You should see:** two greetings, `total tokens`, and `last_call observed
None`. A `404 … /completions` means `is_chat_model` was overridden back to
`False`.
## 02 — Typed refusals, live
```bash
python "demos/human-made/llamaindex/02 - typed-refusals-live.py"
```
Three cases, each with its own `Donkey`: a contact record for `PIIDetected`, a
model that does not exist for `UpstreamRequestError`, and wrong credentials
for `AuthError`.
```python
try:
donkey.llamaindex.llm(model).complete(prompt)
print(name, "NO REFUSAL")
except openai.APIStatusError as err:
error = classify(err.response)
print(name, "->", type(error).__name__, getattr(error, "entities", None))
```
**Needs:** `llm-pii-detection-policy` with `Email` and action `Reject` for the
first case. **You should see:** ` -> ` per case, or
` NO REFUSAL`.
**Learn more:** [LlamaIndex](https://docs.donkey-kit.dev/frameworks/llamaindex.md)
**Source:**
[`demos/human-made/llamaindex/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/llamaindex)
---
Source: https://docs.donkey-kit.dev/examples/adk.md
# Google ADK
Google ADK with `donkey.adk.model("…")`, a `LiteLlm` model. LiteLLM calls
the proxy's `/chat/completions` route. It sends through the SDK's shared HTTP client, so the credentials and the run id
go on the wire, but LiteLLM raises its own errors, so there are **no typed
refusals**. Script 02 exists to show that gap.
Both scripts need a live gateway; there is no offline ADK script.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `basic-gw.py` | An ADK `Agent` through an `InMemoryRunner` | Proxy credentials |
| 02 | `refusal-live.py` | A PII refusal surfacing as LiteLLM's `APIError` | Proxy + PII policy |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,adk]"
set -a; source .env.local; set +a
```
## 01 — An ADK agent
```bash
python "demos/human-made/adk/01 - basic-gw.py"
```
```python
donkey = Donkey.from_env()
agent = Agent(name="greeter", model=donkey.adk.model("gpt-4o"), instruction="Answer in one short sentence.")
events = asyncio.run(InMemoryRunner(agent=agent).run_debug("Say hello in exactly three words.", quiet=True))
print(events[-1].content.parts[0].text)
print("total tokens", events[-1].usage_metadata.total_token_count)
print("last_call ", donkey.last_call.status.value, donkey.last_call.surface)
```
**You should see:** a one-sentence answer, `total tokens` from the last
event's `usage_metadata`, and `last_call unavailable …`. The record is set in
the `Runner`'s own task, not where the script reads it, and a cold read on
this adapter reports `unavailable` rather than `unobserved` ([#740](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/740)). Read it
in an `after_model_callback` instead, as shown under
[Native Gemini](https://docs.donkey-kit.dev/frameworks/adk.md#native-gemini). LiteLLM also logs a
provider-list banner, which is harmless.
## 02 — A refusal that is not typed
```bash
python "demos/human-made/adk/02 - refusal-live.py"
```
A PII prompt against a proxy with the PII policy applied. LiteLLM keeps the
status code and the message but drops the response headers, so `classify()`
has nothing to read. The script therefore catches LiteLLM's own error:
```python
try:
asyncio.run(InMemoryRunner(agent=agent).run_debug(PII_PROMPT, quiet=True))
print("NO REFUSAL")
except litellm.exceptions.APIError as err:
print(type(err).__name__, err.status_code)
print(str(err).splitlines()[0])
```
**Needs:** `llm-pii-detection-policy` with `Email` and action `Reject`. **You
should see:** `APIError 403` and the first line of the proxy's message — **not**
`PIIDetected`. Without the policy it prints `NO REFUSAL`.
If you need typed refusals with ADK, use
`donkey.adk.gemini("gemini-2.5-flash")` on a `Format=Gemini` proxy and
`classify()` its error (see [Native Gemini](https://docs.donkey-kit.dev/frameworks/adk.md#native-gemini)).
Otherwise prefer a framework path where the SDK owns the transport, such as
[OpenAI](https://docs.donkey-kit.dev/examples/openai.md) or [LangGraph](https://docs.donkey-kit.dev/examples/langgraph.md). A `404`
here means the proxy's upstream has no `/chat/completions` route.
**Learn more:** [Google ADK](https://docs.donkey-kit.dev/frameworks/adk.md)
**Source:**
[`demos/human-made/adk/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/adk)
---
Source: https://docs.donkey-kit.dev/examples/anthropic.md
# Anthropic
The native `anthropic` client from `donkey.anthropic.client()`, on the
governed transport. It sends `/v1/messages`, so it needs a proxy provisioned
**`Format=Anthropic`** (for example `ddk-anthropic-inbound`). The default DDK
proxies are `Format=OpenAI` and return 404 on `/v1/messages`. Because the
client shares the governed transport, `last_call` and `simulate()` work just
as they do for OpenAI.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `native-messages.py` | A native Messages call and `last_call` | Proxy credentials, `Format=Anthropic` |
| 02 | `typed-refusals-simulated.py` | Four simulated refusals, typed | Nothing — any placeholder values |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm,anthropic]"
set -a; source .env.local; set +a
export DONKEY_LLM_PROXY_URL=https:///ddk-anthropic-inbound/ # 01 only; overrides the file
```
**`anthropic` 0.x and 1.x both work.** The `anthropic` 1.x releases
(September 2026) moved to `httpx2` and reject an `httpx` client, so the SDK
hands them an `httpx2` client that sends through the governed transport
([Anthropic SDK](https://docs.donkey-kit.dev/frameworks/anthropic.md#manual-equivalent)). Published
`donkey-kit` builds up to 0.1.1.dev1 predate this and fail on 1.x with
`TypeError: Invalid http_client argument`; with one of those, add
`"anthropic<1"` to the install line. These examples were run against the proxy
with `anthropic` 0.125.0.
The client id and secret are the same pair as the other proxies. The model is
`claude-haiku-4-5-20251001` (the `MODEL` constant).
## 01 — Native Messages
```bash
python "demos/human-made/anthropic/01 - native-messages.py"
```
```python
async with Donkey.from_env() as donkey:
client = donkey.anthropic.client() # native AsyncAnthropic, POST /v1/messages
raw = await client.messages.with_raw_response.create(
model=MODEL,
max_tokens=32,
messages=[{"role": "user", "content": "Say hello in exactly three words."}],
)
print(raw.parse().content[0].text)
last = donkey.last_call
print("served_provider", last.served_provider)
print("served_model ", last.served_model)
print("input_tokens ", last.input_tokens)
print("output_tokens ", last.output_tokens)
print("request-id ", raw.headers.get("request-id"))
print("request_id ", last.request_id)
```
**You should see:** the reply; `served_provider`, `served_model`, and input and
output tokens from `last_call`; then Anthropic's `request-id` header next to
`last_call.request_id`. `request-id` is not one of the headers the SDK reads
for `request_id` yet, so the two can differ, or the field can be `None`.
## 02 — Typed refusals, simulated
```bash
python "demos/human-made/anthropic/02 - typed-refusals-simulated.py"
```
`donkey.simulate(...)` works for the Anthropic client too, because it shares
the governed transport. Each refusal surfaces as the matching Anthropic error
class, and `classify()` maps its response to the SDK type.
```python
client = donkey.anthropic.client()
for refusal in REFUSALS:
async with donkey.run(id=f"anthropic-simulated-{refusal.__name__}"):
with donkey.simulate(refusal):
try:
await client.messages.create(model="claude-haiku-4-5-20251001", max_tokens=32,
messages=[{"role": "user", "content": "hello"}])
except anthropic.APIStatusError as err:
error = classify(err.response)
print(type(err).__name__, "->", type(error).__name__, error.policy, error.correlation_id)
```
```text
PermissionDeniedError -> PIIDetected pii-detection anthropic-simulated-PIIDetected
BadRequestError -> PromptInjectionBlocked prompt-injection-protection anthropic-simulated-PromptInjectionBlocked
PermissionDeniedError -> ContentSafetyBlocked content-safety anthropic-simulated-ContentSafetyBlocked
RateLimitError -> TokenBudgetExceeded token-rate-limit anthropic-simulated-TokenBudgetExceeded
```
The left column is what Anthropic's client alone would tell you — a 403 is a
`PermissionDeniedError` whether it was PII or content safety. The right column
is what the gateway actually decided.
**Learn more:** [Anthropic SDK](https://docs.donkey-kit.dev/frameworks/anthropic.md)
**Source:**
[`demos/human-made/anthropic/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/anthropic)
---
Source: https://docs.donkey-kit.dev/examples/gemini.md
# Gemini
These scripts show the wire: plain `httpx` against a proxy provisioned
**`Format=Gemini`** (for example `ddk-gemini-inbound`) with the same
`client_id` / `client_secret` pair. For an agent, ADK's native `Gemini` model is
bound to the same proxy by `donkey.adk.gemini("gemini-2.5-flash")` — see
[Native Gemini](https://docs.donkey-kit.dev/frameworks/adk.md#native-gemini). The route is
`/models/:generateContent`. `DonkeyConfig` still resolves
and validates the credentials, and `classify()` still types the errors.
Both scripts need a live `Format=Gemini` proxy; there is no offline Gemini
script.
| # | Script | Shows | Needs |
| --- | --- | --- | --- |
| 01 | `native-generate-content.py` | A Gemini-shaped `contents` body, usage and routing | Proxy credentials, `Format=Gemini` |
| 02 | `openai-shape-rejected.py` | An OpenAI-shaped request typed as `UpstreamRequestError` | Proxy credentials, `Format=Gemini` |
## Install
Follow the [examples setup](https://docs.donkey-kit.dev/examples.md#setup) first, then:
```bash
python -m pip install -e "../donkey-development-kit/python[llm]" # httpx + classify()
set -a; source .env.local; set +a
export DONKEY_LLM_PROXY_URL=https:///ddk-gemini-inbound/ # overrides the file
```
The model is `gemini-2.5-flash` (the `MODEL` constant).
## 01 — Native `generateContent`
```bash
python "demos/human-made/gemini/01 - native-generate-content.py"
```
```python
cfg = DonkeyConfig.from_env().validated(need="llm")
headers = {"client_id": cfg.llm_proxy_client_id, "client_secret": cfg.llm_proxy_client_secret}
ok = httpx.post(
f"{cfg.llm_proxy_url}/models/{MODEL}:generateContent",
headers=headers,
json={"contents": [{"role": "user", "parts": [{"text": "Say hello in exactly three words."}]}]},
timeout=60,
)
ok.raise_for_status()
body = ok.json()
print(body["candidates"][0]["content"]["parts"][0]["text"])
print("total tokens ", body["usageMetadata"]["totalTokenCount"])
print("routing ", ok.headers.get("x-llm-proxy-model-based-routing-success"))
```
**You should see:** the reply, `total tokens` from `usageMetadata`, and the
proxy's model-based `routing` header. A `KeyError: 'candidates'` means the
proxy refused; print `ok.text` to see why. A `404` means the URL points at a
`Format=OpenAI` proxy.
## 02 — The wrong shape, typed
```bash
python "demos/human-made/gemini/02 - openai-shape-rejected.py"
```
An OpenAI-shaped `/chat/completions` request sent to the Gemini proxy comes
back as Gemini's error envelope. `classify()` types it as a bad request, not a
policy refusal.
```python
bad = httpx.post(
f"{cfg.llm_proxy_url}/chat/completions",
headers=headers,
json={"model": "gemini-2.5-flash", "messages": [{"role": "user", "content": "hello"}]},
timeout=60,
)
error = classify(bad)
print(type(error).__name__, bad.status_code, getattr(error, "code", None), getattr(error, "error_type", None))
print(error)
```
**You should see:** `UpstreamRequestError`, the HTTP status, the upstream code
and error type, and the message. Typing Gemini's list envelope needs
`donkey-kit` 0.1.0.dev9 or later.
**Learn more:** [Typed refusals](https://docs.donkey-kit.dev/errors.md) · [Model access](https://docs.donkey-kit.dev/frameworks.md)
**Source:**
[`demos/human-made/gemini/`](https://github.com/Donkey-Development-Kit/donkey-development-kit-demos/tree/main/demos/human-made/gemini)
---
Source: https://docs.donkey-kit.dev/roadmap.md
# Roadmap
DDK is delivered in five phases. Each phase is a GitHub milestone with a clear
goal; the progress bars and issue lists below are loaded from those milestones
when you open the page. Phase 1 shipped as `donkey-kit` 0.1.0, and its
stabilisation release, 0.1.1, is the current release on
[PyPI](https://pypi.org/project/donkey-kit/). Phase 2 is next.
Complete every issue in the milestone is closed ·{' '}
In progress work has landed and more is open ·{' '}
Planned designed, not started. Want to help
move a phase forward? See [Contribute](https://docs.donkey-kit.dev/community/contribute.md).
## Phase 1 — Build the MVP
**Version 0.1.0 · Goal:** a developer who tries DDK for fifteen minutes
finds three things they cannot get from a `base_url` and two headers — and
one of them saves them from a production incident.
The shared governed transport and the capabilities that hang off it: typed
refusals, budget & pacing, the local simulator, `simulate()` and the
conformance suite, OpenTelemetry GenAI spans, correlation IDs and cost tags.
Plus a deep LangGraph adapter, `connection_kwargs()` for seven more
frameworks, the decorators and CLI, the documentation site, and the PyPI
release. The phase closed with steering the gateway's semantic cache,
surfacing cache hits and semantic-routing matches (topic and score) on
`last_call`, and fixing model-substitution detection when the gateway strips
a provider prefix.
## Phase 1.1 — Stabilize the MVP
**Version 0.1.1 · Goal:** make the 0.1.0 surface safe and dependable to
build on, without adding Phase 2 scope.
Credentials scoped to the plane and the config source they belong to,
endpoint trust (no credentials sent to unchecked or non-loopback `http://`
endpoints), and secrets kept out of printed output and refusal messages.
Adapter and transport fixes across the frameworks, a typed `AgentKilled`
refusal for the Agent Kill Switch, an `all` extra that installs, a native
`donkey.adk.gemini()` model for Gemini-format proxies, and a pass over the
docs. Upgrading from 0.1.0? See the
[migration guide](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/MIGRATION.md).
## Phase 2 — Differentiate, go beyond
**Version 0.2.0 · Goal:** capabilities no generic LLM client offers,
because they depend on the platform behind the gateway.
Governed tool access (MCP discovery and binding), A2A `serve` / `expose` /
`dev`, on-behalf-of identity, human-in-the-loop, scan & publish to the
registry with a GitHub Action, declarative refusal handlers and a
classification registry for custom policies, typed federated guardrail
verdicts, kill-switch awareness with the kill reason, and a second deep
framework adapter chosen by demand.
Around them: an inbound correlation ID carried across MCP and A2A hops, a
stated concurrency contract for budget and telemetry under parallel fan-out,
a per-request business group for shared multi-tenant clients, bring-your-own
provider keys, and a token and registry cache. Also in this phase: governed
access to [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev),
a System One decision model, with the same typed refusals, spans, budget and
simulator support as LLM calls, plus worked patterns for confidence-gated
review, tool ranking and budget-aware routing. Now that the package is published,
the phase also picks up archiving daily PyPI adoption metrics.
## Phase 3 — Platform capabilities
**Version 0.3.0 · Goal:** agents that know the rules before they call, not
only after they are refused.
The policy handshake (read the in-force policy set, advisory only),
policy updates pushed to the code at run boundaries, a governed
structured-output path, a model catalog with honest model resolution,
evaluation hooks on the run span, monetary spend and wallet warnings, custom
cost dimensions, and run-level cost rollups. Much of this phase depends on
new gateway endpoints tracked under Upstream gaps.
## Phase 4 — Enterprise readiness
**Version 0.4.0 · Goal:** everything a security, compliance and platform
team asks for before rolling DDK out broadly.
Independent security review and supply-chain hardening, a latency and
overhead budget enforced in CI, a full pass over every error message, the
public API contract and deprecation policy, compliance evidence mapping (EU
AI Act Art. 12, ISO 42001, OWASP LLM Top 10), log shipping, data residency,
workload identity, air-gapped operation, opt-in gateway failover, a clear
boundary against the platform's infrastructure-as-code, and a support model
with a release cadence.
## Phase 5 — Complete rollout
**Version 1.0.0 · Goal:** a stable, multi-language SDK with guarantees.
A TypeScript port of the core capabilities with conformance scenarios shared
across Python and TypeScript, the remaining framework adapters brought to the
full bar by demand, framework-docs partnerships and launch channels, and the
1.0 release with stability guarantees.
## Cross-cutting tracks
These milestones run alongside every phase.
Every endpoint, header, class name and constructor argument DDK relies on is
checked against the real platform and the installed framework packages
before features are built on it. The semantic-caching and semantic-routing
header contracts are now live-captured; next up is scoping mTLS as a
consumer auth option for the LLM proxy.
Requests to the Omni Gateway team for platform capabilities DDK needs — such
as a budget-query endpoint, policy discovery, a dry-run mode and richer
guardrail verdicts. They land whenever the gateway ships them.
## What DDK will not build
At each of these boundaries the job is to make the platform's own capability
reachable and typed, not to reproduce it:
- Client-side policy enforcement — the gateway is the enforcement point.
- Client-side semantic caching — DDK steers the gateway's semantic cache and
reports hits, but never caches responses itself.
- A provisioning control plane competing with API Manager or Terraform.
- Re-implementations of Agent Scanners, Kill Switch or Trusted Agent Identity.
- An approval UI or queue.
- An evaluation framework.
- The gateway inside your agent process.
- A home-grown A2A protocol implementation — the official `a2a-sdk` is used.
---
Source: https://docs.donkey-kit.dev/community/team.md
# The DDK team
DDK is built in the open by its two creators and a growing group of
contributors, with plenty of help from AI coding agents along the way. The
goal we share: agents that are as aware of the gateway as the gateway is of
them.
## Creators
## Contributors
DDK is an open-source community project, not an official Salesforce or
MuleSoft product. Team members contribute in a personal capacity. Want to
join them? Start with [Contribute](https://docs.donkey-kit.dev/community/contribute.md).
---
Source: https://docs.donkey-kit.dev/community/contribute.md
# Contribute
DDK is open source under the Apache-2.0 licence, and contributions of every
size are welcome. This page is the short version; the full runbook is
[`CONTRIBUTING.md`](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/CONTRIBUTING.md)
in the repository.
## Ways to help
Every change starts as an issue. Search first, then file one with what you
expected and what happened.
Small, well-scoped issues that are a good way into the codebase.
This site lives in `website/` (Nextra). Every page has an "Edit this page"
link.
Runnable demos live in the companion `donkey-development-kit-demos`
repository.
## From issue to merge
### File or find the issue
No change lands without a GitHub issue — the issue is the plan. Each issue
carries exactly one milestone, which is the release it targets (see the
[Roadmap](https://docs.donkey-kit.dev/roadmap.md)).
### Cut a branch from `develop`
Branch names follow `/-`, for example
`fix/42-proxy-url-trailing-slash` or `docs/13-verified-apis-update`. Always
branch from `develop`; never from or into `main`.
```bash
git checkout develop && git pull --ff-only
git checkout -b docs/13-verified-apis-update
```
### Run the pre-PR gate
Run the same checks CI runs, from `python/`:
```bash
pip install -e ".[dev,llm,cli]"
pytest -q # tests
mypy # mypy --strict
ruff check . # lint
lint-imports # the framework-free core contract
```
If you touched an adapter, also run `python scripts/verify_frameworks.py`.
### Open a pull request into `develop`
The PR body includes `Closes #`, a `## Summary`, a `## Test plan` and
a `## Post-deploy steps` section (write `None.` when nothing applies). PRs are
squash-merged, so `develop` reads as one commit per issue.
## Contributing from a fork
Not a member of the `Donkey-Development-Kit` organisation? The flow is the
same, from a fork:
```bash
git clone https://github.com//donkey-development-kit.git
cd donkey-development-kit
git remote add upstream https://github.com/Donkey-Development-Kit/donkey-development-kit.git
git fetch upstream
git checkout -b docs/13-verified-apis-update upstream/develop
```
Open the PR from your fork into `Donkey-Development-Kit:develop` and tick
**Allow edits by maintainers**. A maintainer sets the milestone and labels,
runs the secret-gated checks that GitHub does not run on fork PRs, re-captures
any live fixture on the team sandbox before a `verified-apis.md` row flips,
and merges.
## Keep the docs in sync
When code changes what DDK does, the docs change in the same pull request —
or a `documentation` follow-up issue is filed and linked. After editing pages
here, regenerate the AI-readable docs and commit the result:
```bash
cd website
npm run generate:llms
```
**Verification discipline.** DDK never documents or codes against an
endpoint, header or class name that has not been confirmed against the real
platform. If you can't confirm one, say so in the issue rather than guessing
— see the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md).
---
Source: https://docs.donkey-kit.dev/reference/configuration.md
# Configuration
`Donkey.from_env()` builds its configuration from environment variables, then
the config files, then defaults. `Donkey(DonkeyConfig(...))` uses only what you
pass. See [Building a config](#building-a-config) for both paths,
[Precedence](#precedence) for the exact order,
[What the SDK sends where](#what-the-sdk-sends-where) for which credential goes
to which host, and [Which credentials a URL receives](#which-credentials-a-url-receives)
for the rule that ties a URL from a config file to credentials from the same files.
## Building a config
There are two ways to build a `Donkey`, and they resolve values differently:
| You write | What it reads |
|---|---|
| `Donkey.from_env(...)` or `Donkey()` | Environment variables, then the [config files](#config-files-secrets-and-trust), then defaults. `Donkey.from_env()` takes only the cost tags (`team`, `project`, `env`, `enduser_id`) and `on_model_substitution` as kwargs, and applies them on top. |
| `Donkey(DonkeyConfig(...))` | Only the values you pass. Every other field is its default. No environment variable and no config file is read. |
Building `DonkeyConfig(...)` yourself does not fill the gaps from the
environment. With `DONKEY_TIMEOUT_S=5` set,
`Donkey(DonkeyConfig(llm_proxy_url=…))` still uses the default `timeout_s` of
`60.0`. To change a few fields and keep everything else from the environment
and the files, resolve first and then override:
```python
from donkey_kit import Donkey, DonkeyConfig
donkey = Donkey(
DonkeyConfig.from_env().with_overrides(
telemetry_capture_content=True,
registry_cache_ttl_s=60,
)
)
```
`DonkeyConfig.from_env()` takes no arguments. `with_overrides(...)` accepts any
`DonkeyConfig` field name, the names in the "`DonkeyConfig` field" columns
below, and returns a new config.
## Governed model access
The three required values for the LLM proxy:
| Env var | Meaning |
|---|---|
| `DONKEY_LLM_PROXY_URL` | Proxy base URL: `https:////` — **no `/v1`**. Must be `https://`; plain `http://` is accepted for loopback hosts such as the [local simulator](https://docs.donkey-kit.dev/simulator.md), and for other hosts only with [`DONKEY_ALLOW_HTTP`](#endpoints-must-use-https). |
| `DONKEY_LLM_PROXY_CLIENT_ID` | Consumer client ID (the per-agent identity). |
| `DONKEY_LLM_PROXY_CLIENT_SECRET` | Consumer client secret. |
Auth is a `client_id` / `client_secret` **header pair** (consumer auth), **not**
a bearer token, and separate from any Anypoint control-plane credential. The
OpenAI SDK still requires a non-empty `api_key` slot, which the proxy ignores.
The stock gateway also accepts a **single** colon-joined header —
`authorization: Bearer :` or
`apikey: :` — which its
`dataweave-headers-transformation` policy splits back into the pair. DDK doesn't use that form: it always
sends the two-header pair, because `client_id` is the per-agent
[attribution](https://docs.donkey-kit.dev/telemetry.md#cost-attribution-tags) unit. The colon-joined value is not an
alternative once a `client_id` header is present — the policy ignores it.
Missing required fields are reported **all at once** with their env-var names,
so you can fix configuration in a single pass.
| Optional env var | `DonkeyConfig` field | Meaning |
|---|---|---|
| `DONKEY_LLM_PROXY_KEY` | `llm_proxy_key` | Value for the client's API-key slot, sent to the LLM proxy: `Authorization: Bearer ` from OpenAI-compatible clients, `x-api-key` from the Anthropic client, `x-goog-api-key` from ADK's `gemini()`. Unset, the SDK sends the placeholder `client-id-enforced`. A client-id proxy ignores this value, so leave it unset unless your proxy authenticates on an API key. |
## JWT / model-wallet auth mode
A **model-wallet** proxy identifies the caller from an IdP-issued **JWT** plus a
durable wallet-selector client ID, with Client ID Enforcement disabled and **no
`client_secret`**. Select it with `DONKEY_LLM_PROXY_AUTH=jwt`:
| Env var | `DonkeyConfig` field | Meaning |
|---|---|---|
| `DONKEY_LLM_PROXY_AUTH` | `llm_proxy_auth` | Data-plane auth mode: `client-id` (default), `jwt`, or [`bearer`](#bearer-token-auth-mode). |
| `DONKEY_LLM_PROXY_WALLET_CLIENT_ID` | `llm_proxy_wallet_client_id` | The wallet's system-generated client ID, sent as the `X-Client-Id` header. Required in `jwt` mode. |
In `jwt` mode the required fields are `llm_proxy_url` **and**
`llm_proxy_wallet_client_id` — **not** `client_id` / `client_secret`. The
rotating JWT is **not** a config value: supply it through an
[`AuthProvider`](#auth-providers) passed as `Donkey(llm_auth=…)`, so the SDK can
re-fetch it as it rotates and refresh it once on a `401`:
```python
from donkey_kit import Donkey, DonkeyConfig
from donkey_kit.core.auth import StaticToken # or your own rotating AuthProvider
donkey = Donkey(
DonkeyConfig(
llm_proxy_url="https:////",
llm_proxy_auth="jwt",
llm_proxy_wallet_client_id="",
),
llm_auth=StaticToken(""), # rides as Authorization: Bearer
)
```
`jwt` mode is **async-only** — the rotating credential is fetched from an
async `AuthProvider`, so the blocking client (`sync=True`) is refused with an
actionable error. Use client-id auth for a synchronous caller.
The JWT is attached only by the SDK's shared async HTTP client, so it reaches
the proxy only from async calls through adapters with **transport injection**
built from a `Donkey` that has `llm_auth`: the raw client, LangGraph, Strands,
OpenAI Agents SDK, the Anthropic SDK, ADK's `model()` and `gemini()`,
LlamaIndex and MS Agent Framework.
Sync calls through an adapter, such as LangGraph's `invoke()` or LlamaIndex's
`complete()`, go through the blocking client, so they raise the same
`ConfigError` as `sync=True` before sending anything.
These can't send the JWT, so they raise `ConfigError` before sending anything,
rather than sending the `client-id-enforced` placeholder as the bearer:
- CrewAI, in `jwt` mode at all: its native OpenAI provider builds its own
HTTP clients, which the SDK can't add the JWT to. Use client-id auth with
CrewAI;
- every adapter on a `Donkey` without `llm_auth`, and the module-level
factories (for example
`from donkey_kit.integrations.langgraph import chat_model`), which have no
`llm_auth`. Use `Donkey(llm_auth=…)` in `jwt` mode.
See [Injection depth](https://docs.donkey-kit.dev/frameworks.md#injection-depth-differs-by-framework).
## Bearer-token auth mode
A proxy that authenticates model calls on a plain bearer token, with no wallet
selector and no `client_id` / `client_secret` pair, uses
`DONKEY_LLM_PROXY_AUTH=bearer`. The only required field is `llm_proxy_url`. The
token comes from an [`AuthProvider`](#auth-providers) passed as
`Donkey(llm_auth=…)`:
```python
from donkey_kit import Donkey, DonkeyConfig
from donkey_kit.core.auth import StaticToken # or your own rotating AuthProvider
donkey = Donkey(
DonkeyConfig(
llm_proxy_url="https:////",
llm_proxy_auth="bearer",
),
llm_auth=StaticToken(""), # rides as Authorization: Bearer
)
```
The SDK's shared client adds `Authorization: Bearer ` to every model
call, replacing whatever the framework client put there. It sends no
`X-Client-Id` and no `client_id` / `client_secret`. This is the way to send a
bearer token from the Anthropic client and ADK's `gemini()`, which have no
OpenAI-style API-key slot. A `401` refreshes the token once, and the
[endpoint rule](#which-credentials-a-url-receives) covers it.
`bearer` mode has the same reach as `jwt` mode: async calls through the raw
client, LangGraph, Strands, OpenAI Agents SDK, the Anthropic SDK, ADK's
`model()` and `gemini()`, LlamaIndex and MS Agent Framework, on a `Donkey`
built with `llm_auth`. The forms that can't carry the token raise
`ConfigError` instead of sending an unauthenticated request: CrewAI's
`connection_kwargs()` and `llm()`, the blocking client (`sync=True`), sync
calls through an adapter such as LangGraph's `invoke()` or LlamaIndex's
`complete()`, and the module-level factories, which have no `llm_auth`.
### Auth providers
An `AuthProvider` (in `donkey_kit.core.auth`) is any object with two async
methods: `token()` returns the current credential and `invalidate()` drops a
cached one.
Each provider belongs to one credential plane:
- `Donkey(llm_auth=…)` supplies the LLM-proxy credential. It is used only in
`jwt` and `bearer` mode.
- `Donkey(auth=…)` supplies the Anypoint control-plane credential. By default
it is an `AnypointConnectedApp` built from `ANYPOINT_CLIENT_ID` /
`ANYPOINT_CLIENT_SECRET`. It is used only for control-plane calls, never for
model calls. Its token counts as set in code for the
[endpoint rule](#which-credentials-a-url-receives).
What happens on a `401`:
| Request | On `401` |
|---|---|
| Model call, client-id mode | Raises `AuthError` straight away. There is no token to refresh. |
| Model call, `jwt` or `bearer` mode | Calls `invalidate()` on the `llm_auth` provider, re-sends once with the new token, then raises `AuthError`. |
| Control-plane call | Calls `invalidate()` on the `auth` provider, re-sends once, then raises `AuthError`. |
| Blocking client (`sync=True`) | Raises `AuthError` straight away. |
| Provider | Use it for |
|---|---|
| `StaticToken(token)` | A token injected out-of-band, for example from CI. Never refreshes. |
| `AnypointConnectedApp(client_id=…, client_secret=…, control_plane_url=…, http_client=…)` | OAuth2 client credentials against the Anypoint token endpoint. Caches the token in memory and refreshes it 60 seconds before expiry. The token endpoint must be `https://` (loopback excepted, or any host with `DONKEY_ALLOW_HTTP`). `AnypointConnectedApp.from_config(config, http_client=…)` builds one from a `DonkeyConfig` and also applies the [endpoint check](#which-credentials-a-url-receives) before the first token request. |
| `ChainedAuth(*providers)` | Tries providers in order; the first that yields a token wins. |
For a rotating JWT from your IdP, implement the two methods yourself and pass
the object as `llm_auth`.
## Optional attribution
| Env var | `DonkeyConfig` field | Sent as |
|---|---|---|
| `DONKEY_APP_NAME` | `application_name` | The `X-Anypoint-Client-Application` request header. |
| `DONKEY_BUSINESS_GROUP` | `business_group` | The `X-Anypoint-Business-Group` request header. |
These two values go only on request headers. No span attribute carries them.
The gateway attributes traffic to the `client_id` credential (or, in `jwt`
mode, the JWT's `client_id` claim), so you don't need either value for
attribution.
Both header names are unconfirmed guesses. The gateway hasn't been seen
reading them (see §3 of the
[verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md)).
The first time the SDK sends one,
it emits an `UnverifiedValueWarning`. Unlike the correlation and cost-tag
headers, these names have no config key, so you can't override them. To
stop the warning, leave both values unset.
## Correlation headers
Per-call and per-run correlation IDs ride on request headers. The IDs also
appear on spans and exceptions regardless of the header names. See
[Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md#correlation-ids).
- `X-Correlation-Id` carries the run ID. The gateway reads this header and
echoes it on the response `x-correlation-id`, so a client log line joins to
the gateway's record.
- `X-Donkey-Request-Id` carries the per-call ID. The name is the SDK's own
convention, and the gateway doesn't read it.
Both are confirmed names, so neither emits a warning. Override one only if a
proxy of your own in front of the gateway expects a different name:
| Env var | `DonkeyConfig` field | Default | Meaning |
|---|---|---|---|
| `DONKEY_CORRELATION_HEADER` | `correlation_header` | `X-Correlation-Id` | Request header that carries the per-run correlation ID. |
| `DONKEY_CALL_ID_HEADER` | `call_id_header` | `X-Donkey-Request-Id` | Request header that carries the per-call ID. |
### Header names you can't use
`correlation_header`, `call_id_header` and the four `cost_*_header` keys
([Cost-attribution tags](#cost-attribution-tags)) are checked whenever a config
is built, whether from the environment, a file or code. Names compare
case-insensitively. A name must start with `X-` (all the defaults do), and a
key can't name:
| Header names | Why |
|---|---|
| `Host`, `Forwarded`, any `X-Forwarded-*`, `X-Real-IP`, `Content-Length`, `Transfer-Encoding`, `Connection`, `Upgrade`, `TE`, `Trailer`, `Expect` | They decide where the request goes or how its body is framed. |
| `X-HTTP-Method-Override`, `X-HTTP-Method`, `X-Method-Override`, `X-Original-URL`, `X-Original-URI`, `X-Rewrite-URL` | Some servers and gateways read them to change the request's method or path. |
| `Authorization`, `Proxy-Authorization`, `Cookie`, `x-api-key`, `api-key`, `apikey`, `api_key`, `x-goog-api-key`, `client_id`, `client_secret`, `X-Client-Id` | They carry or select credentials. |
| `X-Anypoint-Client-Application`, `X-Anypoint-Business-Group`, the `x-cache-*` headers, any `x-stainless-*`, `Content-Type`, `Accept`, `Accept-Encoding`, `User-Agent` | The SDK, its HTTP client or a framework's SDK already sets them. |
| A name another of these keys already uses, including its default | Two values would share one header. |
| Anything that isn't a valid HTTP header name (for example, one with a space) | It can't be sent. |
| Any other name that doesn't start with `X-` (for example `Team`, `Origin`, `Range`) | Standard and framework headers (`Content-Encoding`, `Via`, `anthropic-version`, …) change how the request is handled. |
```text
ConfigError: cost_team_header names the header 'Host', which can't be used: it controls where the request goes or how it is framed. cost_team_header is set in the environment (DONKEY_COST_TEAM_HEADER). Choose a different header name.
```
## Cost-attribution tags
A fixed set of dimensions set once and emitted on every call as `donkey.cost.*`
span attributes (and, only if you opt in, as request headers). Override them per run with
`donkey.run(team=…, project=…, env=…, enduser_id=…)`. The key set is fixed — an
unknown dimension is a configuration error, not a silently dropped tag. See
[Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md).
| Env var | `Donkey.from_env` kwarg | Meaning |
|---|---|---|
| `DONKEY_COST_TEAM` | `team` | Owning team. |
| `DONKEY_COST_PROJECT` | `project` | Project / workload. |
| `DONKEY_COST_ENV` | `env` | Deployment environment (e.g. `prod`). |
| `DONKEY_COST_ENDUSER_ID` | `enduser_id` | End-user ID (the `enduser.id` tag). |
In a config file these live under a `[donkey.cost]` table (the end-user
dimension keeps its dotted key):
```toml
[donkey.cost]
team = "support"
project = "triage-v2"
env = "prod"
"enduser.id" = "user-42"
```
A `[donkey.cost]` table in `.donkey-kit.local.toml` merges into the one in
`.donkey-kit.toml` key by key: a local `"enduser.id"` adds to the project
file's `team` and `project` instead of replacing them. `DONKEY_COST_*`
variables then override single dimensions.
The LLM Gateway doesn't read cost tags from request headers, so by default DDK
sends **no** cost-tag header. The `donkey.cost.*` span attributes carry every
tag either way. To also send the tags as request headers, for example to a
proxy of your own that reads them, opt in:
| Env var | `DonkeyConfig` field | Default | Meaning |
|---|---|---|---|
| `DONKEY_SEND_COST_HEADERS` | `send_cost_headers` | `false` | Also send the cost tags as request headers. |
With it enabled, the headers use the SDK's own names (`X-Anypoint-Cost-Team`,
…). Because no gateway reads them, they're a convention rather than a
placeholder for some real name, and they emit no warning. Rename them to match
whatever receiver you send them to: `DONKEY_COST_TEAM_HEADER`, `DONKEY_COST_PROJECT_HEADER`,
`DONKEY_COST_ENV_HEADER`, `DONKEY_COST_ENDUSER_HEADER` (or the matching
`cost_*_header` config keys), within the
[names you can't use](#header-names-you-cant-use). `send_cost_headers` is not a `Donkey.from_env`
kwarg; in code, use
`DonkeyConfig.from_env().with_overrides(send_cost_headers=True)`.
With `send_cost_headers` on, the headers, including the end-user ID, go only
on model requests to the LLM proxy, never to the Anypoint platform. Where the
tags go is described in
[Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md#cost-attribution-tags).
## Telemetry
| Env var | `DonkeyConfig` field | Meaning |
|---|---|---|
| `DONKEY_TELEMETRY` | `telemetry` | Emit OTel spans at all (default `true`). |
| `DONKEY_TELEMETRY_CAPTURE_CONTENT` | `telemetry_capture_content` | Put prompt/completion text on spans (default **`false`**). |
`telemetry_capture_content` defaults to `false` on purpose: spans are emitted
inside your process, **upstream of the gateway's PII masking**, so capturing
content re-exports the very text the platform masks. Enable it only for a
trusted collector. See [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md).
## Behaviour
| Env var | `DonkeyConfig` field | Default | Meaning |
|---|---|---|---|
| `DONKEY_TIMEOUT_S` | `timeout_s` | `60.0` | HTTP timeout for governed calls, in seconds. |
| `DONKEY_MAX_RETRIES` | `max_retries` | `3` | Retries for transient upstream failures (`502` / `503` / `504`) with backoff. Policy refusals are never retried, and a gateway fallback is never retried twice. |
| `DONKEY_ON_MODEL_SUBSTITUTION` | `on_model_substitution` | `off` | `off` surfaces a model substitution on `donkey.last_call`; `raise` turns it into `ModelSubstituted`. See [Telemetry & cost](https://docs.donkey-kit.dev/telemetry.md#two-behaviours-worth-knowing). |
| `DONKEY_REGISTRY_CACHE_TTL_S` | `registry_cache_ttl_s` | `300` | How long registry lookups (used by [tool access](https://docs.donkey-kit.dev/tool-access.md)) are cached in memory, in seconds. |
| `DONKEY_NO_CACHE` | — | unset | Set to `1`, `true` or `yes` to bypass that in-memory registry cache. |
| `DONKEY_TRUST_PROJECT_CONFIG` | — | unset | Set to `1` (or `true`, `yes`, `on`) to let a URL from the working directory's config files receive credentials from elsewhere. Read only from the environment. See [Which credentials a URL receives](#which-credentials-a-url-receives). |
| `DONKEY_ALLOW_HTTP` | — | unset | Set to `1` (or `true`, `yes`, `on`) to allow plain `http://` endpoints on non-loopback hosts. Read only from the environment. See [Endpoints must use `https://`](#endpoints-must-use-https). |
### Invalid values
Every `DonkeyConfig` is checked when it is built, including by
`with_overrides()`, and an invalid value raises `ConfigError` before any
request is sent. One error lists every bad field and where each was set:
| Field | Accepted |
|---|---|
| `timeout_s` | A number greater than `0` |
| `max_retries`, `registry_cache_ttl_s` | A whole number, `0` or more |
| `telemetry`, `telemetry_capture_content`, `send_cost_headers` | `True` or `False` in code; `1`, `true`, `yes`, `on`, `0`, `false`, `no` or `off` (any case) in the environment or a config file |
| `region` | `us`, `eu`, `ca` or `jp` |
| `llm_proxy_auth` | `client-id` or `jwt` (any case in the environment or a config file) |
| `on_model_substitution` | `off` or `raise` (any case in the environment or a config file) |
A misspelt switch such as `DONKEY_TELEMETRY=flase` is an error, not `false`:
```text
ConfigError: Configuration is invalid:
- timeout_s is 'abc', set in the environment (DONKEY_TIMEOUT_S); expected a number
- telemetry is 'flase', set in the environment (DONKEY_TELEMETRY); expected one of 1, true, yes, on, 0, false, no, off
- max_retries is -1, set in the environment (DONKEY_MAX_RETRIES); expected a whole number, 0 or more
Fix each value where it is set: in code, an environment variable, or .donkey-kit.toml.
```
The environment-only switches above (`DONKEY_NO_CACHE`,
`DONKEY_TRUST_PROJECT_CONFIG`, `DONKEY_ALLOW_HTTP`) are not checked: any value
other than the ones listed leaves them off.
## Anypoint control plane
A separate credential from the LLM proxy, for features that call the Anypoint
platform: registry and tool discovery, publication and provisioning. The SDK
exchanges `ANYPOINT_CLIENT_ID` / `ANYPOINT_CLIENT_SECRET` for a connected-app
token at the Anypoint token endpoint and sends only that token to the platform.
You don't need these values for governed model access. Model calls never use
them.
The features that call the Anypoint platform are
Roadmap in this release and stop before sending
anything. [Which features use the control plane](#which-features-use-the-control-plane)
lists them.
| Env var | Meaning |
|---|---|
| `ANYPOINT_CLIENT_ID` | Connected-app client ID. |
| `ANYPOINT_CLIENT_SECRET` | Connected-app client secret. |
| `ANYPOINT_ORG_ID` | Anypoint organization ID. |
| `ANYPOINT_ENV` | Anypoint environment (default `Sandbox`). |
| `ANYPOINT_REGION` | Control-plane region: `us` (default), `eu`, `ca`, or `jp`. |
| `ANYPOINT_BASE_URL` | Explicit control-plane base URL; overrides the region. Must be `https://` (loopback hosts excepted, or any host with `DONKEY_ALLOW_HTTP`). |
## What the SDK sends where
`Donkey` keeps one HTTP client per credential plane, so a credential never
rides the other plane's requests:
- **Data plane**: the LLM proxy at `llm_proxy_url`. Used by `donkey.llm` and
every framework adapter.
- **Control plane**: the Anypoint platform at `base_url` (default: the host for
`region`). Used for the connected-app token request and for registry and tool
discovery.
### Credentials
| Destination | Client-id mode (default) | `jwt` mode | `bearer` mode | Never sent here |
|---|---|---|---|---|
| **LLM proxy** (`llm_proxy_url`) | `client_id: ` and `client_secret: ` headers. The API-key slot1 carries `llm_proxy_key`, or the placeholder `client-id-enforced`. | `X-Client-Id: ` and `Authorization: Bearer `2. No `client_secret`. | `Authorization: Bearer `2. The API-key slot1 as in client-id mode. No `X-Client-Id`, no `client_secret`. | `ANYPOINT_CLIENT_ID`, `ANYPOINT_CLIENT_SECRET`, the connected-app token, anything from `Donkey(auth=…)` |
| **Anypoint token endpoint** (`/accounts/api/v2/oauth2/token`) | Form body: `grant_type=client_credentials`, `client_id`, `client_secret` (the `ANYPOINT_*` values). | Same. | Same. | Every `llm_proxy_*` value, the JWT or bearer token, `X-Client-Id` |
| **Registry and tools** (`base_url`) | `Authorization: Bearer ` (or the token from `Donkey(auth=…)`). | Same. No `X-Client-Id`. | Same. | Every `llm_proxy_*` value, the JWT or bearer token, `X-Client-Id` |
1 The framework client decides the header: `Authorization: Bearer`
for OpenAI-compatible clients, `x-api-key` for the Anthropic client,
`x-goog-api-key` for ADK's `gemini()`. In `jwt` and `bearer` mode the SDK's
shared client replaces `Authorization` with the token from `llm_auth`.
2 Only from async calls through adapters with transport
injection; see the [`jwt` mode note](#jwt--model-wallet-auth-mode) and the
[`bearer` mode note](#bearer-token-auth-mode).
Model calls never request, send or refresh the connected-app token.
### Credentials go only to checked endpoints
Each of the SDK's HTTP clients attaches credentials only to the endpoints it
was checked for, compared by scheme, host and port:
- the configured endpoint: `llm_proxy_url` on the data plane, `base_url` on the
control plane;
- a URL you pass to a factory in code: `base_url` to `donkey.llm.client()` or
an adapter factory, `api_base` (LlamaIndex, ADK's `model()`, CrewAI),
`openai_api_base` (LangGraph), or `client_args["base_url"]` (Strands). It
must pass the same [`https://` rule](#endpoints-must-use-https) as the
config, and then receives the same credentials.
A request to any other origin is still sent, but without credentials: the SDK
adds none and removes the credential headers the framework set
(`Authorization`, `x-api-key`, `client_id`, `client_secret`, `X-Client-Id` and
the other names in [What printed output hides](#what-printed-output-hides)).
The correlation, attribution and other headers in the next table still go.
This covers:
- a URL changed outside the factories, for example
`ChatOpenAI(**{**kwargs, "base_url": other})`;
- every redirect hop. The SDK's clients don't follow redirects; if your code
asks for it on a request, a hop to another origin carries no credentials.
CrewAI builds its own HTTP client, so its `connection_kwargs()` includes an
`interceptor` that removes the credential headers the same way.
### Which features use the control plane
| Feature | Control-plane requests | In this release |
|---|---|---|
| Model calls: `donkey.llm`, `donkey.openai()`, every adapter | None | Available |
| `donkey doctor` | None. It checks the control-plane settings without sending, and makes one model call to the LLM proxy. | Available |
| `donkey init`, `donkey mock`, `donkey test`, `donkey.simulate()` | None | Available |
| Registry: `donkey.registry.search()`, `resolve_mcp()`, `resolve_agent()`, `explain()`, `warm()` | Token request, then Exchange reads | Roadmap: raises `NotImplementedError` before sending |
| Tool discovery: `donkey.tools.discover()`, `donkey.tools.lock()` | Token request, then Exchange reads | Roadmap: raises `NotImplementedError` before sending |
| Publication and governance: `Publication.preview()` / `verify()`, `Governance.resolve()` / `apply()` / `simulate()` | Token request, then platform calls | Roadmap: raises `NotImplementedError` before sending |
| Hidden provisioning commands (`donkey plan`, `apply`, `drift`, `lint`, `generate`, `status`, `publish`, `verify`) | Token request, then platform calls | Roadmap: exits with "blocked on verification" before sending |
| Your own code calling `AnypointConnectedApp.token()` | Token request | Available |
So in this release the SDK sends the `ANYPOINT_*` credentials only if your own
code asks for a connected-app token.
### Other headers
The correlation and per-call IDs go on every request the SDK's HTTP clients
send, on both planes. The other headers go only on model requests to the LLM
proxy: the connected-app token request and the registry and tool requests
never carry them.
| Header (default name) | Value | When | Sent to |
|---|---|---|---|
| `X-Correlation-Id` (`correlation_header`) | The run ID | Always | Both planes |
| `X-Donkey-Request-Id` (`call_id_header`) | A per-call ID | Always | Both planes |
| `X-Anypoint-Client-Application` | `application_name` | When set | The LLM proxy only |
| `X-Anypoint-Business-Group` | `business_group` | When set | The LLM proxy only |
| `X-Anypoint-Cost-Team`, `-Project`, `-Env`, `-Enduser-Id` | Your cost tags, including the end-user ID | Only with `send_cost_headers`; see [Cost-attribution tags](https://docs.donkey-kit.dev/telemetry.md#cost-attribution-tags) | The LLM proxy only, or a receiver of yours in front of it |
| `x-cache-*` | Your cache controls, including `principal_id` | Inside a `donkey.cache(...)` block | The LLM proxy's semantic cache only |
CrewAI, which gets only a header snapshot, sends the auth headers, the
attribution headers and, with `send_cost_headers`, the configured cost tags. It
doesn't send per-run values or cache controls.
## Config files, secrets and trust
Instead of env vars you can put values in a `[donkey]` table. Keys are the
`DonkeyConfig` field names.
| File | Where | Meant for |
|---|---|---|
| `.donkey-kit.toml` | The current working directory (not searched upward) | Non-secret project config you commit: URLs, client IDs, attribution. |
| `.donkey-kit.local.toml` | The current working directory | Secrets and personal overrides. Keep it gitignored. Same keys. Read with or without `.donkey-kit.toml`. |
| `$XDG_CONFIG_HOME/.donkey-kit.toml`, or `~/.config/.donkey-kit.toml` when `XDG_CONFIG_HOME` is unset or empty | Your user config directory | Personal defaults for every project. Read only when the working directory has neither file above. When `XDG_CONFIG_HOME` is set, `~/.config` isn't read. |
`.donkey-kit.local.toml` merges into `.donkey-kit.toml` key by key, and nested
tables such as `[donkey.cost]` merge the same way. A scalar or an array in the
local file replaces the project file's value; it isn't appended to. Each key
remembers which file it came from.
Both working-directory files must be regular files, or links whose target is
inside the working directory. A link that points elsewhere is refused, because
its keys would be treated as the working directory's:
```text
ConfigError: /home/me/my-agent/.donkey-kit.toml links to /home/me/shared/donkey.toml, outside the working directory. Replace the link with a regular file in the working directory, or set those values in the environment.
```
The user file isn't merged with the working-directory files. If either
`.donkey-kit.toml` or `.donkey-kit.local.toml` exists, the user file is not
read at all, even when only `.donkey-kit.local.toml` exists
([#727](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/727)).
### Precedence
Highest first, per key:
1. Values set in code: anything you change on a resolved config, with
`with_overrides(...)`, `dataclasses.replace(...)` or any other copy, and the
cost-tag and `on_model_substitution` kwargs of `Donkey.from_env(...)`.
2. Environment variables.
3. `./.donkey-kit.local.toml`
4. `./.donkey-kit.toml`
5. The user file (`$XDG_CONFIG_HOME/.donkey-kit.toml`, or
`~/.config/.donkey-kit.toml`), only when 3 and 4 don't exist.
6. Defaults.
A value counts as set in code once it differs from the value that was loaded,
whichever way you changed it; a copy that keeps the value keeps its source.
That includes `DonkeyConfig(**dataclasses.asdict(cfg))` in the same process. A
config rebuilt in another process, for example from a serialized `asdict()`,
keeps where its URLs came from, but its credentials count as set in code. So a
worker in another process (`multiprocessing` spawn, Celery, Ray) that receives
a config with a URL from `.donkey-kit.toml` and credentials from
`.donkey-kit.local.toml` refuses to send them; call `DonkeyConfig.from_env()` in
the worker instead. `DonkeyConfig(...)` built directly reads no environment variable and no file:
every value is the one you pass or the field default. `donkey.run(...)` cost
tags apply on top of all of this for their block. `DONKEY_TRUST_PROJECT_CONFIG`
and `DONKEY_ALLOW_HTTP` are read from the environment only.
```toml
# .donkey-kit.toml (committed)
[donkey]
llm_proxy_url = "https:////"
llm_proxy_client_id = "my-client-id"
```
```toml
# .donkey-kit.local.toml (gitignored)
[donkey]
llm_proxy_client_secret = ""
```
`donkey init` generates `.donkey-kit.toml` from your current environment and
never writes a secret into it. If `.donkey-kit.toml` contains `client_secret`,
`llm_proxy_client_secret` or `llm_proxy_key`, the SDK emits a `ConfigWarning`
(`donkey_kit.core.errors.ConfigWarning`) pointing you to
`.donkey-kit.local.toml`. Under `python -W error` or pytest's
`filterwarnings = error`, that warning is raised as an exception.
### Which credentials a URL receives
Two checks run before the SDK sends a credential. Both raise `ConfigError`, and
nothing is sent.
**When they run:** the LLM-proxy check runs when you build a client:
`donkey.llm.client()`, `donkey.openai()`, any `donkey.` factory, or
`connection_kwargs()`. Creating `Donkey(...)` doesn't run it. The
control-plane check runs at the first control-plane call and again right before
each connected-app token request. `donkey doctor` runs both without sending
anything.
**Scope:** the checks cover `llm_proxy_url`, `base_url` and the connected-app
token endpoint. A URL you pass to a factory in code gets the `https://` check
when the factory is called; see
[Credentials go only to checked endpoints](#credentials-go-only-to-checked-endpoints). They don't cover OTLP exporter endpoints (configured by
OpenTelemetry) or governance `[targets.*].base_url` profiles. No request goes
to a target profile today, because `resolve()`, `apply()` and `simulate()` are
not available yet
([#832](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/832)).
#### Endpoints must use `https://`
This applies wherever the URL comes from: environment, file or code.
| URL | Result |
|---|---|
| `https://…` | Accepted |
| `http://` on a loopback host: `localhost`, `127.0.0.0/8`, `::1` | Accepted |
| `http://` on any other host, including `0.0.0.0`, LAN addresses and container hostnames such as `http://simulator:8080` | `ConfigError` |
| `http://` on any other host, with `DONKEY_ALLOW_HTTP=1` in the environment | Accepted, with a `ConfigWarning` |
| Any other scheme, or a URL with no host | `ConfigError`, whatever `DONKEY_ALLOW_HTTP` says |
```text
ConfigError: llm_proxy_url must be an https:// URL (got http scheme, llm.example.test). Plain http:// is accepted for loopback hosts (localhost, 127.0.0.0/8, ::1), such as the local gateway simulator. Change it to the https:// address of the service, or set DONKEY_ALLOW_HTTP=1 in the environment to allow plain http:// to other hosts.
```
`DONKEY_ALLOW_HTTP` (`1`, `true`, `yes` or `on`) is for networks you control,
such as a simulator in another container or an in-cluster proxy without TLS.
Credentials sent over plain `http://` travel unencrypted, so don't use it across
a network you don't trust. It is read from the environment only; a `[donkey]`
key with that name is ignored. It changes only the `https://` rule: a URL from
the project files still receives only credentials from those files (next
section). When it lets a URL through, the SDK emits a `ConfigWarning` naming the
key (`base_url`, `llm_proxy_url` or `token endpoint`) and the host. Python shows
each distinct warning once per location by default:
```text
ConfigWarning: llm_proxy_url uses plain http:// to llm.example.test because DONKEY_ALLOW_HTTP=1 is set. Credentials and data sent to it are not encrypted in transit.
```
The warning quotes the value as you set it, for example
`DONKEY_ALLOW_HTTP=true is set`. `donkey doctor` adds a `plain http` line while
it is on. See
[CLI](https://docs.donkey-kit.dev/cli.md#donkey-doctor).
#### A URL from the project files gets credentials from the project files only
An endpoint URL read from `./.donkey-kit.toml` or `./.donkey-kit.local.toml`
only receives credentials read from those two files. This applies to loopback
URLs too: they are exempt from the `https://` rule only.
These are the credentials checked for each endpoint:
| Endpoint | Credentials checked |
|---|---|
| `base_url` | `client_id`, `client_secret`, or the token from a `Donkey(auth=…)` provider, which always counts as coming from outside the files |
| `llm_proxy_url`, client-id mode | `llm_proxy_client_id`, `llm_proxy_client_secret`, `llm_proxy_key` (the ones that are set) |
| `llm_proxy_url`, `jwt` mode | `llm_proxy_wallet_client_id`, and the JWT from `llm_auth`, which always counts as coming from outside the files |
| `llm_proxy_url`, `bearer` mode | `llm_proxy_key` if set, and the token from `llm_auth`, which always counts as coming from outside the files |
| The URL comes from | A credential it would receive comes from | Result |
|---|---|---|
| Environment, user file, code, or the region default | Anywhere | Sent |
| `.donkey-kit.toml` / `.donkey-kit.local.toml` | `.donkey-kit.toml` / `.donkey-kit.local.toml` | Sent |
| `.donkey-kit.toml` / `.donkey-kit.local.toml` | An environment variable | `ConfigError` |
| `.donkey-kit.toml` / `.donkey-kit.local.toml` | The user file | `ConfigError` |
| `.donkey-kit.toml` / `.donkey-kit.local.toml` | Code (a value you set or changed on the config) | `ConfigError` |
| `.donkey-kit.toml` / `.donkey-kit.local.toml`, `jwt` mode | The JWT from `llm_auth` (always) | `ConfigError`, every time |
| `.donkey-kit.toml` / `.donkey-kit.local.toml`, `bearer` mode | The token from `llm_auth` (always) | `ConfigError`, every time |
| `.donkey-kit.toml` / `.donkey-kit.local.toml` (`base_url`) | The token from a `Donkey(auth=…)` provider (always) | `ConfigError`, every time |
| `.donkey-kit.toml` / `.donkey-kit.local.toml`, `base_url` on `anypoint.mulesoft.com`, `eu1.anypoint.mulesoft.com`, `ca1.anypoint.mulesoft.com` or `jp1.anypoint.mulesoft.com` | Anywhere | Sent |
| Any of the above, with `DONKEY_TRUST_PROJECT_CONFIG` set in the environment | Anywhere | Sent |
| Any of the above, with only `DONKEY_ALLOW_HTTP` set | As in the rows above | Unchanged: `DONKEY_ALLOW_HTTP` affects the `https://` rule only |
**What counts as code:** any value passed to `DonkeyConfig(...)`; any value on
a resolved config that differs from what was loaded, whether you changed it with
`with_overrides(...)`, `dataclasses.replace(...)` or otherwise; the token from a
`Donkey(auth=…)` provider; and in `jwt` and `bearer` mode the token returned by
`llm_auth`. A
config built entirely with `DonkeyConfig(...)` has its URLs in code too, so the
rule never applies to it. Likewise, a URL you change in code is no longer
treated as coming from the file.
The error names every credential that came from elsewhere, where each came
from, the host, the key and the file:
```text
ConfigError: Not sending llm_proxy_client_secret (from env) to llm-proxy.example.com: llm_proxy_url is set in /home/me/my-agent/.donkey-kit.toml, and credentials from outside the working directory's config files are only sent to hosts those files name when you opt in. To continue, do one of:
- set the URL in the environment instead (DONKEY_LLM_PROXY_URL=https://...)
- keep the credentials in /home/me/my-agent/.donkey-kit.local.toml, next to the project file
- trust this directory's config files by setting DONKEY_TRUST_PROJECT_CONFIG=1
```
The source labels are `env`, `user file` and `code`.
#### Resolving it
1. **Set the URL in the environment** (`DONKEY_LLM_PROXY_URL` /
`ANYPOINT_BASE_URL`). The environment wins over the file, so the URL no
longer comes from the file.
2. **Keep the credentials in `.donkey-kit.local.toml`** in the same directory.
Not available for the `llm_auth` token in `jwt` or `bearer` mode or for a
`Donkey(auth=…)` token, because none of them comes from a file.
3. **Opt in:** set `DONKEY_TRUST_PROJECT_CONFIG` to `1` (or `true`, `yes`,
`on`) in the environment if you trust this directory's config files. The
setting decides whether those files are trusted, so it is read from the
environment only; a `[donkey]` key with that name is ignored.
In `jwt` or `bearer` mode with the URL in a project file, and for a `Donkey(auth=…)` token
with `base_url` in a project file, only 1 and 3 apply; the error lists only
those two.
`donkey doctor` prints each endpoint's host and where it came from (`env`,
`project file`, `local overlay`, `user file` or `default`), and reports this
error on its `config` line without sending a request. See
[CLI](https://docs.donkey-kit.dev/cli.md#donkey-doctor).
## Programmatic
```python
import dataclasses
from donkey_kit import Donkey, DonkeyConfig
# Entirely in code: no environment variable or config file is read.
donkey = Donkey(DonkeyConfig(
llm_proxy_url="https:////",
llm_proxy_client_id="my-client-id",
llm_proxy_client_secret="",
))
# Resolved from env and files, with code values on top. Either form records
# the changed value as set in code:
cfg = DonkeyConfig.from_env().with_overrides(send_cost_headers=True)
cfg = dataclasses.replace(DonkeyConfig.from_env(), send_cost_headers=True)
# Or from the environment, with lifecycle:
async with Donkey.from_env() as donkey:
...
```
## What printed output hides
The SDK keeps credential values out of `repr()` and `str()`, so printing,
logging, a test-failure diff or a traceback with locals doesn't show them. It
only changes how objects are printed: every value is still there, and the SDK
sends the real values.
| Object | Hidden when printed | Still shown |
|---|---|---|
| `DonkeyConfig` | `client_secret`, `llm_proxy_client_secret` and `llm_proxy_key` are left out | Every other field, including client IDs, `llm_proxy_wallet_client_id`, URLs and `cost=CostTags(…, enduser_id=…)` |
| Every adapter's `connection_kwargs()`, ADK's `gemini_connection_kwargs()`, `donkey_kit.core.transport.proxy_auth_headers()` | The value under any key named `api_key`, `apikey`, `api-key`, `x-api-key`, `x-goog-api-key`, `client_secret`, `authorization`, `proxy-authorization` or `cookie` (any case), at any depth, shows as `'***'` | `base_url`, `client_id`, `X-Client-Id`, attribution headers, and `X-Anypoint-Cost-*` headers including the end-user ID |
| `PIIDetected` (`str`, `repr`, `.args`) | The flagged values | Entity types, count and character offsets; see [Errors](https://docs.donkey-kit.dev/errors.md#refusal-messages-dont-repeat-blocked-content) |
The masked mappings are ordinary `dict`s. The table shows how common operations behave:
| Operation | Result |
|---|---|
| `**kwargs`, `kwargs["api_key"]`, `==`, passing it to a framework | Real values |
| `.copy()`, `copy.copy`, `copy.deepcopy`, `kwargs \| other` | Still masked when printed |
| `dict(kwargs)`, `{**kwargs}`, `kwargs.items()` | Top level printed in full, including `api_key`. Nested header mappings stay masked. |
| `json.dumps(kwargs)` | Writes the **real** values |
Some framework objects built from the kwargs print credentials in their own
`repr()` / `str()`:
| Object | Prints |
|---|---|
| LangGraph `ChatOpenAI` (`donkey.langgraph.chat_model()`) | `client_secret` |
| LlamaIndex `OpenAILike` (`donkey.llamaindex.llm()`) | `client_secret` and `api_key` |
| CrewAI `OpenAICompletion` (`donkey.crewai.llm()`) | `api_key` |
Don't print or log these objects, and don't leave them in locals that a
traceback renders.
Not covered by the masking:
- Attribute access and conversions: `config.client_secret`,
`dataclasses.asdict(config)`, `vars(...)`. (Where each value came from is
kept as a keyed digest, so `asdict()` holds each secret once, in its own
field.)
- Objects built from the kwargs, such as a framework client or the
`openai_client` returned by the OpenAI Agents adapter, and whatever those
objects log.
- `attribution_headers()` and `cost_headers()`, which return plain `dict`s.
- `.response` on any `DonkeyError`: `exc.response.request.headers` holds the
real `client_secret` and `Authorization` headers.
- `PIIDetected.gateway_message` and `.response`, which contain the flagged
values.
- `UpstreamRequestError` messages, which include the upstream provider's own
error text and can quote parts of your request.
- OpenTelemetry spans: cost tags, including `donkey.cost.enduser.id`, and
message content when `telemetry_capture_content` is on.
---
Source: https://docs.donkey-kit.dev/reference/last-call.md
# `last_call` field reference
`donkey.last_call` is what the gateway said about the **most recent governed
model call** in the current context. It is an immutable snapshot: every governed
response replaces the record wholesale rather than mutating it, so a reader
always sees one internally-consistent call. It is contextvar-scoped, so a
fan-out of concurrent calls each reads its own record. When a framework sends a
single call from a task of its own, as LangChain's `ainvoke` does, the
`donkey.langgraph(...)` model still brings that call's record back to the caller
(see [LangGraph](https://docs.donkey-kit.dev/frameworks/langgraph.md#which-provider-served-this)).
```python
donkey = Donkey.from_env()
await donkey.openai().responses.create(model="gpt-5.1", input="…")
r = donkey.last_call
r.status # LastCallStatus.OBSERVED
r.served_provider # "openai"
r.total_tokens # 1730
```
Every response-derived field defaults to `None`. `None` always means **not
observed** — never `0`, and never a fabricated value. A count of `0` is a real
observation (an empty completion) and is distinct from `None` (no usage was
reported at all). See [Verification discipline](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/ARCHITECTURE.md#verification-discipline) for why
the SDK never guesses a value it did not see on the wire.
## Observability status
Whether — and from where — the call was observed. These are always meaningful,
even on a cold read.
| Field | Type | Meaning |
|---|---|---|
| `status` | `LastCallStatus` | `OBSERVED` (a governed response populated this record), `UNOBSERVED` (no governed model call has returned in this context yet), or `UNAVAILABLE` (this surface structurally cannot be observed). |
| `observed` | `bool` | `True` iff a governed response actually populated the record (`status is OBSERVED`). |
| `available` | `bool` | `False` only when the current surface structurally cannot be observed; a plain cold read is still `available` — it just has not observed anything yet. |
| `observed_at` | `datetime \| None` | When the record was observed (UTC), for freshness. `None` unless `OBSERVED`. |
| `surface` | `str \| None` | For `UNAVAILABLE`, the adapter surface(s) that cannot observe (e.g. `"adk"`); else `None`. |
`donkey.last_call` is populated only when the governed response passes through
the SDK's shared httpx client. CrewAI, whose framework owns the transport,
routes outside it. A cold read on a `Donkey` that resolved only CrewAI,
LlamaIndex, MS Agent Framework or ADK's `model()` reports `UNAVAILABLE` with
the surface named. See
[When `last_call` is unavailable](https://docs.donkey-kit.dev/telemetry.md#when-last_call-is-unavailable).
## Gateway identity
Who served the call — for correlating a request with the platform's own logs and
for quoting to a provider's support team.
| Field | Type | Meaning |
|---|---|---|
| `request_id` | `str \| None` | The upstream provider's own request id, passed through by the gateway (`x-request-id` / `x-amzn-requestid` / `apim-request-id` / Anthropic's `request-id`). `None` when the gateway forwarded none. |
| `api_instance_id` | `str \| None` | The API Manager instance id that served the call. |
| `environment_id` | `str \| None` | The Anypoint environment id that served the call. |
## Routing & fallback
What the gateway *did* with the request — which provider and model served it, how
it routed, and whether that was a failover. Read live off the shared transport,
so the raw `donkey.llm.client()` path gets them with no framework required. See
[Routing & resilience](https://docs.donkey-kit.dev/telemetry.md#routing--resilience) for the operational story,
or [Which provider served this?](https://docs.donkey-kit.dev/frameworks/langgraph.md#which-provider-served-this)
for why a LangGraph message's own `model_provider` field disagrees with these.
| Field | Type | Meaning |
|---|---|---|
| `requested_model` | `str \| None` | The model the caller **asked for** (from the request body). The reference point for `substituted`. |
| `served_provider` | `str \| None` | The provider the gateway actually routed to (`x-llm-proxy-llm-provider`). |
| `served_model` | `str \| None` | The model the gateway actually served (`x-llm-proxy-llm-model`). May differ from `requested_model` after a fallback. |
| `routing_type` | `str \| None` | The routing strategy the gateway applied (`x-llm-proxy-routing-type`), e.g. `"ModelBased"` or `"Semantic"`. |
| `fallback` | `bool \| None` | Whether the gateway performed a routing **fallback**. `None` when the header is absent (non-proxy / simulated response) — distinct from a definitive `False` ("no fallback occurred"). |
| `substituted` | `bool` | `True` iff the gateway served a *different* model than requested — a silent substitution your cost model and evals are otherwise blind to. A `provider/` prefix on the requested model that names `served_provider` is ignored (`openai/gpt-5-mini` served as `gpt-5-mini` by `openai` is not a substitution). Requires both models known; a missing either side is not a substitution claim. |
| `matched_topic` | `str \| None` | On a **semantic-routing** proxy, the topic the prompt matched (`x-llm-proxy-semantic-routing-success`). `None` on a model-based proxy (the header is semantic-only) or when the message did not parse. |
| `routing_score` | `float \| None` | On a **semantic-routing** proxy, the similarity score of the matched topic (a bare `0.xx` float). `None` on a model-based proxy or when the score did not parse. |
`matched_topic` and `routing_score` are populated **only** on a semantic-routing
proxy (`routing_type == "Semantic"`), which reports *why* it picked a provider.
A model-based proxy emits no semantic header, so both stay `None`.
## Token usage
The per-call token counts from the response body's `usage` object. On a streamed
response these land once the terminal SSE event is scanned, not at record time.
Each is `None` (never `0`) when unobserved or absent.
| Field | Type | Meaning |
|---|---|---|
| `input_tokens` | `int \| None` | Prompt/input tokens billed for this call, as the provider reports them. OpenAI's count includes `cached_tokens`. Anthropic's excludes both `cached_tokens` and `cache_write_tokens`. |
| `output_tokens` | `int \| None` | Completion/output tokens produced (includes `reasoning_tokens`). |
| `total_tokens` | `int \| None` | Total tokens the gateway attributed to this call. |
| `cached_tokens` | `int \| None` | Input tokens served from the prompt cache (billed at the cached rate). |
| `cache_write_tokens` | `int \| None` | Input tokens written into the prompt cache this call. |
| `reasoning_tokens` | `int \| None` | Output tokens spent on model reasoning the developer never sees. |
## Semantic cache
When the proxy is fronted by the Anypoint **semantic-caching** policy, the
gateway reports what it did with each request. Steer it per block with
[`donkey.cache(...)`](https://docs.donkey-kit.dev/budget.md#semantic-cache-steering); read the outcome here.
| Field | Type | Meaning |
|---|---|---|
| `cache_status` | `str \| None` | What the caching policy did (`x-semantic-cache-status`): `"hit"` / `"miss"` / `"bypass"` / `"no-store"`. `None` on a proxy with no caching policy (the header is absent) or a simulated response. |
| `cache_score` | `float \| None` | On a cache **hit**, the similarity score of the matched entry (`x-semantic-cache-score`). `None` on miss/bypass/no-store (the header is hit-only) or when the score did not parse. |
| `cache_hit` | `bool` | `True` iff `cache_status == "hit"` — a verbatim replay with no provider round-trip. A hit never advances the [budget](https://docs.donkey-kit.dev/budget.md) (a replay is no fresh spend). |
A cache **hit** replays a stored completion byte-for-byte, including its
original `usage` block — so the token counts above describe the *cached*
call, not fresh spend. `cache_hit` is the signal that they should not be
counted again.
## On the span
The routing, usage, and cache fields also land on the OpenTelemetry GenAI span
for each governed call, under the pinned `gen_ai.*` keys and the stable
`donkey.*` namespace: `gen_ai.response.model`, `donkey.routing.type`,
`donkey.routing.fallback`, — on a semantic route —
`donkey.routing.matched_topic` / `donkey.routing.score`, and — on a cached
proxy — `donkey.cache.status` / `donkey.cache.score`, alongside the usage
counts. A field that is `None` is omitted from the span entirely. See
[Telemetry](https://docs.donkey-kit.dev/telemetry.md) for the full attribute list.
---
Source: https://docs.donkey-kit.dev/reference/unsupported-boundary.md
# Unsupported boundary
Which platform APIs does the SDK call, and are they supported for third-party
use? This page answers that question for security and procurement reviews.
The maintained list lives in the repository at
[`docs/unsupported-boundary.md`](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/unsupported-boundary.md).
It is separate from the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md): the
ledger records how a fact was established, while this page records whether
MuleSoft publishes the contract for third-party use.
Every platform API the SDK calls is classified:
| Classification | Meaning |
|---|---|
| **Documented and public** | Safe to depend on. |
| **Documented, no SLA for third-party use** | May break; we'll fix. |
| **Undocumented** | Should be empty. Anything here needs a written justification and an owner. |
## Current boundary
The SDK can reach two platform destinations: the Model Proxy, for every model
call, and the Anypoint connected-app token endpoint, only when a control-plane
feature needs a token. Model calls never request that token. The features that
would request it are listed in
[Which features use the control plane](https://docs.donkey-kit.dev/reference/configuration.md#which-features-use-the-control-plane);
all of them are Roadmap in this release, so the token endpoint is contacted
only if your own code calls `AnypointConnectedApp.token()`. The rows below
classify each contract the SDK consumes. Any feature that would need an unconfirmed endpoint stops
before making a network request.
| Destination / contract | Classification | SDK use |
|---|---|---|
| Anypoint connected-app token endpoint | **Documented and public** | Retrieves an OAuth bearer token with client credentials. |
| Model Proxy OpenAI-format `/responses` endpoint | **Documented and public** | Sends buffered or streaming model requests with the documented `client_id` / `client_secret` headers and reads OpenAI-format usage. The raw client can also call documented OpenAI-native routes such as `/chat/completions`; `/responses` is the route tested against a deployed proxy. A streamed `/chat/completions` request that the proxy routes to a Gemini upstream does not stream OpenAI chunks; see the warning below. |
| Model Proxy policy refusals (observed) | **Documented and public** | Classifies Client ID Enforcement, token-rate-limit, PII, Injection Protection, Regex Prompt Guard, Azure Content Safety, and Amazon Bedrock Guardrails responses captured from a deployed proxy. |
| Upstream provider error pass-through | **Documented, no SLA for third-party use** | Classifies the nested non-`429` `4xx` provider envelope as `UpstreamRequestError`; generic `5xx` responses become `UpstreamModelError` by status only. The envelope schema belongs to the upstream provider, and MuleSoft's public Model Proxy page states no pass-through compatibility contract. |
| `x-llm-proxy-ratelimit` success-budget sentence | **Documented, no SLA for third-party use** | Updates `donkey.budget`; an absent or changed value is ignored. |
| Gateway identity and routing extension headers | **Documented, no SLA for third-party use** | Populates `donkey.last_call`; missing or unrecognised values become `None`. |
**Streaming chat completions to a Gemini upstream.** On an OpenAI-format
proxy that routes to Gemini, `chat.completions.create(..., stream=True)`
returns a `text/event-stream` whose events are whole `chat.completion`
objects: the text is in `choices[0].message`, not `choices[0].delta`, every
event has `finish_reason: "stop"`, and no `data: [DONE]` is sent. The openai
client does not validate the events, so `chunk.choices[0].delta` is `None`
and `client.chat.completions.stream(...)` fails. This was observed live
(VERIFIED (LIVE), non-conformant) and is a gateway gap, not an SDK one.
Until it is fixed, use `stream=False` for chat completions on Gemini routes.
`donkey.last_call` still records the call's token totals. OpenAI-routed
proxies stream normal chunks, and a `Format=Gemini` proxy's native
`streamGenerateContent` route streams correctly.
Exchange search and resolution, API Manager governed-state reads, MCP
discovery and binding, and provisioning/publication are
Roadmap — not hidden dependencies. Today they
raise `NotImplementedError` before making any network request. The SDK also
never calls a Model Proxy `/models` endpoint, because the proxy has no model
catalog endpoint.
The full ledger links each contract to its official documentation, SDK
consumer, evidence, and maintenance owner. Its **Undocumented surfaces**
section is empty.
## Local Mode is not supported
Donkey Development Kit does not support Omni/Flex Gateway **Local Mode**.
The [local simulator](https://docs.donkey-kit.dev/simulator.md) (`donkey mock` / `simulate()`) is the
supported local dev loop; real-gateway behavior is exercised against a
hosted (Connected Mode) proxy instead. A stock Flex Gateway 1.14.0 Local
Mode image rejects the LLM Proxy and MCP Bridge policies as missing
extensions and refuses configuration until registered to a control plane —
enough to make this scope call, bounded to that image version. See the
[verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md#6-governance--local-mode-the-verification-milestone)
and [issue #661](https://github.com/Donkey-Development-Kit/donkey-development-kit/issues/661).
## Support statement
Donkey Development Kit is an independent, community-maintained project with
best-effort maintainer support and no SLA. It is not affiliated with, endorsed
by, or supported by Salesforce or MuleSoft. "Agent Fabric", "Anypoint", and
"Omni Gateway" are Salesforce trademarks.
## Why the boundary stays small
The SDK doesn't invent endpoints: every call it makes is against a classified,
known API, or it doesn't happen at all. See the
[verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/main/docs/verified-apis.md).