Google ADK
Google’s Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy in one of two ways, depending on the proxy’s ingress Format:
model()— ADK’sLiteLlmmodel wrapper, for aFormat=OpenAIproxy (the default). The adapter translates the governed connection into LiteLLM’s own model-string and kwarg conventions for you.gemini()— ADK’s nativeGeminimodel, for aFormat=Geminiproxy. The SDK’s shared HTTP client is injected, so every governed feature works: per-run correlation, spans, usage anddonkey.last_call. See Native Gemini.
What you get
- A native
google.adk.models.lite_llm.LiteLlmorgoogle.adk.models.Gemini, with the proxy auth and attribution headers set. - For
model(): theopenai/model prefix and LiteLLM kwarg names handled automatically. Supported atconnection_kwargs(). LiteLLM gets a pre-built OpenAIclientthat sends through the SDK’s shared HTTP client, so per-run correlation, retries, spans,donkey.last_calland thejwt-mode JWT apply (see Notes). - For
gemini(): the shared client injected throughHttpOptions.httpx_async_client, round-trip verified live against aFormat=Geminiproxy.
Install
pip install "donkey-kit[adk]"Quickstart
Python
from donkey_kit.integrations.adk import model
llm = model("gpt-4o")llm is a real google.adk.models.lite_llm.LiteLlm instance. The model string
is prefixed with openai/ before it reaches LiteLLM (openai/gpt-4o), which
is the prefix LiteLLM’s OpenAI-compatible route expects — you don’t add it
yourself.
Three ways to construct
1. Off a shared Donkey instance:
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
llm = donkey.adk.model("gpt-4o")2. Module-level factory (shortest):
from donkey_kit.integrations.adk import model
llm = model("gpt-4o")3. Governed kwargs, native constructor:
from donkey_kit import Donkey
from google.adk.models.lite_llm import LiteLlm
async with Donkey.from_env() as donkey:
llm = LiteLlm(model="openai/gpt-4o", **donkey.adk.connection_kwargs())Manual equivalent
from google.adk.models.lite_llm import LiteLlm
llm = LiteLlm(
model="openai/gpt-4o",
api_base=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
extra_headers=..., # client_id / client_secret header pair
client=..., # an AsyncOpenAI on the SDK's shared client
max_retries=0, # the SDK retries in its own transport layer
)LiteLLM uses api_base and extra_headers, not base_url /
default_headers — connection_kwargs() already translates for you.
client is present when the openai package is installed; LiteLLM’s OpenAI
route uses it in place of the client it would build. max_retries=0 has to
be passed to LiteLLM itself: LiteLLM sets the client’s retry count on every
call, and its default is 2. An api_base /
base_url passed to model() must pass the
https:// rule, and the
adapter builds client on that URL.
Native Gemini
A proxy provisioned with Format = Gemini exposes the native Gemini API at
POST /<base-path>/models/<model>:generateContent (and
:streamGenerateContent). ADK’s own Gemini model speaks that API, so
gemini() returns a native google.adk.models.Gemini bound to the proxy, with
the SDK’s shared HTTP client injected.
The three forms mirror model(). Pass the bare Gemini model id — there is no
prefix, because the URL path carries the model:
from donkey_kit import Donkey
from google.adk.models import Gemini
async with Donkey.from_env() as donkey:
# 1. Off a shared Donkey instance
llm = donkey.adk.gemini("gemini-2.5-flash")
# 3. Governed kwargs, native constructor
llm = Gemini(model="gemini-2.5-flash", **donkey.adk.gemini_connection_kwargs())# 2. Module-level factory
from donkey_kit.integrations.adk import gemini
llm = gemini("gemini-2.5-flash")Pointing at the Gemini proxy. DONKEY_LLM_PROXY_URL usually names a
Format=OpenAI proxy. If your Gemini proxy is a different one, pass its URL
as base_url; the client_id/client_secret pair must be contracted on that
proxy, and the URL must pass the
https:// rule:
llm = donkey.adk.gemini("gemini-2.5-flash", base_url="https://…/ddk-gemini-inbound/")Any other keyword is passed to ADK’s Gemini and overrides the governed
default — for example retry_options, or your own client_kwargs.
Manual equivalent. This is what gemini_connection_kwargs() returns:
from google.adk.models import Gemini
llm = Gemini(
model="gemini-2.5-flash",
base_url=..., # the Format=Gemini proxy URL
client_kwargs={
"api_key": ..., # a placeholder; google-genai requires one
"http_options": {
"base_url": ..., # same URL
"api_version": "", # the proxy path has no /v1beta segment
"headers": ..., # client_id / client_secret header pair
"timeout": ..., # milliseconds
"httpx_async_client": ..., # the SDK's shared client
},
},
)client_kwargs replaces ADK’s default HTTP options wholesale, so all five
http_options keys are passed together. google-genai requires an API key and
always sends it as x-goog-api-key; the gateway authenticates on the
client_id/client_secret pair and ignores it.
donkey.last_call. Usage (input_tokens, output_tokens,
total_tokens, cached_tokens, reasoning_tokens) is read from Gemini’s
usageMetadata, and requested_model from the URL path. Gemini’s
total_tokens includes the thinking tokens it also reports as
reasoning_tokens. A Format=Gemini proxy is a passthrough, so it sends no
routing headers: served_provider, served_model and routing_type stay
None, and substituted is False. request_id and api_instance_id are
populated.
last_call is contextvar-scoped, and ADK’s Runner makes the model call in a
task of its own, so the caller of runner.run_async(...) reads UNOBSERVED.
Read it in an after_model_callback, which runs in the same task as the call:
from google.adk.agents import LlmAgent
def record_usage(callback_context, llm_response):
r = donkey.last_call
print(r.requested_model, r.input_tokens, r.output_tokens)
return None # keep the model's response
agent = LlmAgent(
name="assistant",
model=donkey.adk.gemini("gemini-2.5-flash"),
after_model_callback=record_usage,
)When streaming, the callback runs once per partial response; usage lands on the last one.
Errors. google-genai raises its own google.genai.errors.APIError
(ClientError for a 4xx) and the SDK does not wrap it. The error’s .response
is the proxy’s HTTP response, so classify() gives you the typed Donkey error:
from donkey_kit.core.errors import classify
from google.genai import errors as genai_errors
try:
async for event in runner.run_async(...):
...
except genai_errors.APIError as exc:
err = classify(exc.response) # e.g. AuthError on 401, UpstreamRequestError on 404Reusing a model across asyncio.run(...) calls works. The shared HTTP
client keeps one connection pool per event loop, so a sync app that wraps each
run in asyncio.run(...) can build the Donkey (or call the module-level
gemini()) once and reuse it.
Notes
- Both factories send through the SDK’s shared HTTP client, so the
correlation ID bound by
donkey.run()reaches every request. donkey.last_callis set in the context that made the call. ADK’sRunnercalls the model in a task of its own, so read it in anafter_model_callback(see Native Gemini). Untilgemini()has been called on aDonkey, a cold read on aDonkeythat resolved only ADK reportsUNAVAILABLErather thanUNOBSERVED, becausemodel()is still listed as not observing calls. Aligning that, and the matching conformance exemptions, is tracked in #740 .- Refusals on the
model()path aren’t typed. LiteLLM raises its own error without the response headers, soclassify()has nothing to read; see the ADK examples. gemini()needsgoogle-adk>=2.4, the first release whoseGeminiacceptsclient_kwargs; theadkextra declares that floor. On an older ADK,Geminidrops the governed client without an error and talks to Google directly, sogemini()raisesNotImplementedErrorinstead of returning a model that bypasses the gateway.google-adkrequireslitellm>=1.84as a floor, not a ceiling — pin your own upper bound if you need one.
See the error taxonomy for how proxy rejections surface through
ADK’s LiteLlm model, and Native Gemini for gemini().