Strands Agents
Strands Agents gets a governed OpenAIModel pointed at the Agent Fabric LLM
proxy. The connection details travel in Strands’ client_args, which Strands
passes straight through to its underlying OpenAI client.
What you get
- A native
strands.models.openai.OpenAIModel. - Full header and transport injection through
client_args— per-run correlation IDs anddonkey.last_callwork, as with LangGraph. - Supported at
connection_kwargs().
Install
pip install "donkey-kit[strands]"Quickstart
Python
from donkey_kit.integrations.strands import model
llm = model("gpt-4o")llm is a real strands.models.openai.OpenAIModel instance — pass it to your
Agent as you would any other Strands model.
Three ways to construct
1. Off a shared Donkey instance:
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.strands.model("gpt-4o")2. Module-level factory (shortest):
from donkey_kit.integrations.strands import model
llm = model("gpt-4o")3. Governed kwargs, native constructor:
from donkey_kit import Donkey
from strands.models.openai import OpenAIModel
async with Donkey.from_env() as donkey:
llm = OpenAIModel(model_id="gpt-4o", **donkey.strands.connection_kwargs())Manual equivalent
from strands.models.openai import OpenAIModel
llm = OpenAIModel(
model_id="gpt-4o",
client_args={
"base_url": ..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
"api_key": ...,
"default_headers": ..., # client_id / client_secret header pair
"http_client": ..., # a non-owning view of the SDK's shared client
"max_retries": 0, # the SDK retries in its own transport layer
},
stream=False, # see Notes
)Everything the SDK injects lives inside the single client_args dict that
Strands forwards to its internal OpenAI client.
Notes
-
Build the agent with
retry_strategy=None. A StrandsAgentretries a throttled model call by default (up to 6 attempts), and Strands treats every429as throttling. On the proxy a429is a budget refusal, so turn the agent’s retry off and let the SDK’s transport handle the transient5xx:from strands import Agent agent = Agent(model=donkey.strands.model("gpt-4o"), retry_strategy=None)The model itself has
max_retries=0, so the OpenAI client under it doesn’t retry either. -
Strands forwards
client_argsverbatim to the underlying OpenAI client, so both header injection (default_headers) and transport injection (http_client) are available. -
The client stays open. Strands opens and closes an OpenAI client for every request (
async with AsyncOpenAI(**client_args)). Thehttp_clientit gets is a view whose close is a no-op, so the SDK’s shared client survives every call. Onlydonkey.aclose()ends the connection pool. -
Streaming is off by default. The governed model sets
stream=False. A proxy routing to a Gemini upstream answers a streamed request with one wholechat.completionand no chunk deltas, and Strands fails on it. Passdonkey.strands.model("gpt-4o", stream=True)on routes that stream.
See the error taxonomy for how proxy rejections surface as typed exceptions, and the verification ledger for the current status of every constructor signature this adapter depends on.