Skip to Content
FrameworksStrands

Strands Agents

Strands Agents gets a governed OpenAIModel pointed at the Agent Fabric LLM proxy. The connection details travel in Strands’ client_args, which Strands passes straight through to its underlying OpenAI client.

What you get

  • A native strands.models.openai.OpenAIModel.
  • Full header and transport injection through client_args — per-run correlation IDs and donkey.last_call work, as with LangGraph.
  • Supported at connection_kwargs().

Install

pip install "donkey-kit[strands]"

Quickstart

from donkey_kit.integrations.strands import model llm = model("gpt-4o")

llm is a real strands.models.openai.OpenAIModel instance — pass it to your Agent as you would any other Strands model.

Three ways to construct

1. Off a shared Donkey instance:

from donkey_kit import Donkey async with Donkey.from_env() as donkey: llm = donkey.strands.model("gpt-4o")

2. Module-level factory (shortest):

from donkey_kit.integrations.strands import model llm = model("gpt-4o")

3. Governed kwargs, native constructor:

from donkey_kit import Donkey from strands.models.openai import OpenAIModel async with Donkey.from_env() as donkey: llm = OpenAIModel(model_id="gpt-4o", **donkey.strands.connection_kwargs())

Manual equivalent

from strands.models.openai import OpenAIModel llm = OpenAIModel( model_id="gpt-4o", client_args={ "base_url": ..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix "api_key": ..., "default_headers": ..., # client_id / client_secret header pair "http_client": ..., # a non-owning view of the SDK's shared client "max_retries": 0, # the SDK retries in its own transport layer }, stream=False, # see Notes )

Everything the SDK injects lives inside the single client_args dict that Strands forwards to its internal OpenAI client.

Notes

  • Build the agent with retry_strategy=None. A Strands Agent retries a throttled model call by default (up to 6 attempts), and Strands treats every 429 as throttling. On the proxy a 429 is a budget refusal, so turn the agent’s retry off and let the SDK’s transport handle the transient 5xx:

    from strands import Agent agent = Agent(model=donkey.strands.model("gpt-4o"), retry_strategy=None)

    The model itself has max_retries=0, so the OpenAI client under it doesn’t retry either.

  • Strands forwards client_args verbatim to the underlying OpenAI client, so both header injection (default_headers) and transport injection (http_client) are available.

  • The client stays open. Strands opens and closes an OpenAI client for every request (async with AsyncOpenAI(**client_args)). The http_client it gets is a view whose close is a no-op, so the SDK’s shared client survives every call. Only donkey.aclose() ends the connection pool.

  • Streaming is off by default. The governed model sets stream=False. A proxy routing to a Gemini upstream answers a streamed request with one whole chat.completion and no chunk deltas, and Strands fails on it. Pass donkey.strands.model("gpt-4o", stream=True) on routes that stream.

See the error taxonomy for how proxy rejections surface as typed exceptions, and the verification ledger  for the current status of every constructor signature this adapter depends on.

Last updated on