LlamaIndex
LlamaIndex gets a governed OpenAILike LLM pointed at the Agent Fabric LLM
proxy, with the chat-model flag a chat-only gateway requires already set.
What you get
- A native
llama_index.llms.openai_like.OpenAILike. is_chat_model=Trueandis_function_calling_model=Trueset for you, andmax_retries=0soOpenAILikedoesn’t retry on top of the SDK’s transport.- Supported at
connection_kwargs(). Sync and async calls send through the SDK’s HTTP clients (http_client,async_http_client), so per-run correlation, retries, spans anddonkey.last_callapply (see Notes).
Install
pip install "donkey-kit[llamaindex]"Quickstart
Python
from donkey_kit.integrations.llamaindex import llm
model = llm("gpt-4o")model is a real llama_index.llms.openai_like.OpenAILike instance, ready to
hand to any LlamaIndex query engine, chat engine, or agent.
Three ways to construct
1. Off a shared Donkey instance:
from donkey_kit import Donkey
async with Donkey.from_env() as donkey:
model = donkey.llamaindex.llm("gpt-4o")2. Module-level factory (shortest):
from donkey_kit.integrations.llamaindex import llm
model = llm("gpt-4o")3. Governed kwargs, native constructor:
from donkey_kit import Donkey
from llama_index.llms.openai_like import OpenAILike
async with Donkey.from_env() as donkey:
model = OpenAILike(model="gpt-4o", **donkey.llamaindex.connection_kwargs())Manual equivalent
from llama_index.llms.openai_like import OpenAILike
model = OpenAILike(
model="gpt-4o",
api_base=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix
api_key=...,
default_headers=..., # client_id / client_secret header pair
is_chat_model=True, # required — see below
is_function_calling_model=True,
http_client=..., # the SDK's blocking client
async_http_client=..., # the SDK's async client
max_retries=0, # the SDK's transport retries
)LlamaIndex uses api_base rather than base_url; connection_kwargs()
already translates for you. An api_base passed to llm() must pass the
https:// rule.
Notes
-
Always set
is_chat_model=True.OpenAILikedefaults tois_chat_model=False, which routes requests to the completions endpoint instead of chat — and that fails against a chat-only proxy like the Omni Gateway LLM proxy.connection_kwargs()always sets it (andis_function_calling_model=True); set it yourself if you constructOpenAILikeoutside the adapter. -
llm()sets model defaults from the bare name. LlamaIndex looks models up by their exact OpenAI name, so a provider-prefixed name likeopenai/gpt-5-miniwould miss its reasoning-model handling and its context-window table.llm()strips one<provider>/prefix before both lookups, and still sends the prefixed name to the proxy:context_windowcomes from LlamaIndex’s table (400,000 forgpt-5-mini). For a name the table doesn’t list, such as a Gemini or Bedrock model, it stays atOpenAILike’s default of 3,900 tokens. That default also caps agent memory and RAG prompt packing, so passcontext_window=for those models.- A prefixed reasoning model (gpt-5, o-series) gets
temperature=1.0, sendsmax_tokensasmax_completion_tokens, and sendsreasoning_effort. LlamaIndex already does the same for the bare name.
Any
temperature=,context_window=oradditional_kwargs=you pass takes precedence.connection_kwargs()carries no model, so if you buildOpenAILikeyourself, set these yourself. -
donkey.last_callis set in the context that made the call. A cold read (no call yet in this context) on aDonkeythat resolved only adapters like this one still reportsstatus == LastCallStatus.UNAVAILABLErather thanUNOBSERVED, because LlamaIndex is still listed as not observing calls. Aligning that, and the matching conformance exemptions, is tracked in #740 . -
jwtmode is async-only.acomplete()/achat()carry the JWT; synccomplete()/chat()go through the blocking client, which can’t fetch one, so they raiseConfigErrorbefore sending anything. -
Printing the model shows credentials.
OpenAILike’s ownrepr()/str()includeclient_secretandapi_key. Don’t print or log it; see What printed output hides.
See the error taxonomy for how proxy rejections surface as typed exceptions, and the verification ledger for the current status of every constructor signature this adapter depends on.