Skip to Content
FrameworksLlamaIndex

LlamaIndex

LlamaIndex gets a governed OpenAILike LLM pointed at the Agent Fabric LLM proxy, with the chat-model flag a chat-only gateway requires already set.

What you get

  • A native llama_index.llms.openai_like.OpenAILike.
  • is_chat_model=True and is_function_calling_model=True set for you, and max_retries=0 so OpenAILike doesn’t retry on top of the SDK’s transport.
  • Supported at connection_kwargs(). Sync and async calls send through the SDK’s HTTP clients (http_client, async_http_client), so per-run correlation, retries, spans and donkey.last_call apply (see Notes).

Install

pip install "donkey-kit[llamaindex]"

Quickstart

from donkey_kit.integrations.llamaindex import llm model = llm("gpt-4o")

model is a real llama_index.llms.openai_like.OpenAILike instance, ready to hand to any LlamaIndex query engine, chat engine, or agent.

Three ways to construct

1. Off a shared Donkey instance:

from donkey_kit import Donkey async with Donkey.from_env() as donkey: model = donkey.llamaindex.llm("gpt-4o")

2. Module-level factory (shortest):

from donkey_kit.integrations.llamaindex import llm model = llm("gpt-4o")

3. Governed kwargs, native constructor:

from donkey_kit import Donkey from llama_index.llms.openai_like import OpenAILike async with Donkey.from_env() as donkey: model = OpenAILike(model="gpt-4o", **donkey.llamaindex.connection_kwargs())

Manual equivalent

from llama_index.llms.openai_like import OpenAILike model = OpenAILike( model="gpt-4o", api_base=..., # from DONKEY_LLM_PROXY_URL, no /v1 suffix api_key=..., default_headers=..., # client_id / client_secret header pair is_chat_model=True, # required — see below is_function_calling_model=True, http_client=..., # the SDK's blocking client async_http_client=..., # the SDK's async client max_retries=0, # the SDK's transport retries )

LlamaIndex uses api_base rather than base_url; connection_kwargs() already translates for you. An api_base passed to llm() must pass the https:// rule.

Notes

  • Always set is_chat_model=True. OpenAILike defaults to is_chat_model=False, which routes requests to the completions endpoint instead of chat — and that fails against a chat-only proxy like the Omni Gateway LLM proxy. connection_kwargs() always sets it (and is_function_calling_model=True); set it yourself if you construct OpenAILike outside the adapter.

  • llm() sets model defaults from the bare name. LlamaIndex looks models up by their exact OpenAI name, so a provider-prefixed name like openai/gpt-5-mini would miss its reasoning-model handling and its context-window table. llm() strips one <provider>/ prefix before both lookups, and still sends the prefixed name to the proxy:

    • context_window comes from LlamaIndex’s table (400,000 for gpt-5-mini). For a name the table doesn’t list, such as a Gemini or Bedrock model, it stays at OpenAILike’s default of 3,900 tokens. That default also caps agent memory and RAG prompt packing, so pass context_window= for those models.
    • A prefixed reasoning model (gpt-5, o-series) gets temperature=1.0, sends max_tokens as max_completion_tokens, and sends reasoning_effort. LlamaIndex already does the same for the bare name.

    Any temperature=, context_window= or additional_kwargs= you pass takes precedence. connection_kwargs() carries no model, so if you build OpenAILike yourself, set these yourself.

  • donkey.last_call is set in the context that made the call. A cold read (no call yet in this context) on a Donkey that resolved only adapters like this one still reports status == LastCallStatus.UNAVAILABLE rather than UNOBSERVED, because LlamaIndex is still listed as not observing calls. Aligning that, and the matching conformance exemptions, is tracked in #740 .

  • jwt mode is async-only. acomplete() / achat() carry the JWT; sync complete() / chat() go through the blocking client, which can’t fetch one, so they raise ConfigError before sending anything.

  • Printing the model shows credentials. OpenAILike’s own repr() / str() include client_secret and api_key. Don’t print or log it; see What printed output hides.

See the error taxonomy for how proxy rejections surface as typed exceptions, and the verification ledger  for the current status of every constructor signature this adapter depends on.

Last updated on