@kindgi/adapter-model-openai-compat
npm install @kindgi/adapter-model-openai-compat · source
OpenAI-compatible ModelProvider for @kindgi/capabilities. Wraps the official openai SDK client behind the framework's provider-neutral interface and works against any endpoint that speaks the Chat Completions wire format — OpenAI, Ollama, vLLM, llama-server, LM Studio, OpenRouter, Groq, Together, Fireworks, DeepInfra, or a LiteLLM proxy. One adapter, many backends: baseURL selects the endpoint.
Purpose
Section titled “Purpose”Translate between the framework's ModelCallInput / ModelCallResult and a non-streaming chat.completions.create call:
- Messages map role for role. Assistant
toolCallsbecometool_callswith JSON-encoded arguments;toolmessages carrytool_call_id. toolsbecomefunctiontools;structuredOutputbecomesresponse_format: { type: 'json_schema', json_schema: { name, schema, strict: true } };temperatureandmaxOutputTokens(asmax_tokens) are sent only when set.- Chat Completions endpoints reject
.in function names, so tool names are encoded.→__on send and decoded on receive (acme.orders.lookup↔acme__orders__lookup). Tool ids must not contain a literal__. Tool-call arguments that are not valid JSON decode to{}. finish_reasonmapstool_calls/function_call→tool-use,content_filter→content-filter,length→length, anything else →stop.costUsdis computed from the matchingModelInfo.cost: prompt and completion tokens, each per 1K tokens. The wire format carries no pricing, so rates come from the caller.
Exports
Section titled “Exports”createOpenAICompatModelProvider(options)— returns aModelProviderwhosemetadataisoptions.metadataunchanged. The SDK client is created on the firstinvoke()and reused while the key stays the same.invoke()throws wheninput.modelis not one ofmetadata.models[].name, and passes the turn'sabortSignalto the request, so a cancelled turn cancels the call.OpenAICompatProviderOptions:baseURL: string— endpoint root, e.g.BASE_URLS.OLLAMA_LOCAL.apiKey: string | (() => string | Promise<string>)— a static key, or a resolver called on everyinvoke(): a rotated key takes effect on the next call. Local runners that ignore keys still need a non-empty string such as'unused'.metadata: ProviderMetadata— surfaced to the router.models[]must list every model invoked through this connection, with its per-1K-token cost.clientOptions?— otheropenaiclient options (timeout,maxRetries,defaultHeaders,fetch, …);apiKeyandbaseURLare excluded.extraBody?— fields merged into every request body: settings an endpoint takes that the OpenAI format has no field for. A Qwen thinking model served by vLLM, SGLang or llama-server needs{ chat_template_kwargs: { enable_thinking: false } }, or its answer starts with its thinking and a typed (JSON) answer fails. The fields the adapter sets (EXTRA_BODY_RESERVED:model,messages,tools,response_format,temperature,max_tokens,stream) are refused.
openAICompatAdapterFactory(OPENAI_COMPAT_ADAPTER_ID) — theAdapterFactorya runtime registers. A provider registration names its endpoint inadapter_config.baseURL(an http(s) URL;openAICompatBaseUrlreads and checks it), any extra request fields asextraBody.*keys (adapter_configis flat, so one key per field and dots nest:"extraBody.chat_template_kwargs.enable_thinking": false;openAICompatExtraBodyexpands and checks them), and, for an endpoint that needs a key, its secret insecret_ref; without one the adapter sends'unused', as local runners expect.BASE_URLS— well-known endpoints:OPENAI,OLLAMA_LOCAL,VLLM_LOCAL,LLAMA_SERVER_LOCAL,LM_STUDIO_LOCAL,OPENROUTER,GROQ,TOGETHER,FIREWORKS,DEEPINFRA,LITELLM_LOCAL. Any other URL works too.ModelProvider— type re-export from@kindgi/capabilities.
Example
Section titled “Example”import { createProviderRegistry } from '@kindgi/capabilities';import { BASE_URLS, createOpenAICompatModelProvider } from '@kindgi/adapter-model-openai-compat';
const apiKey = process.env.OPENAI_API_KEY;if (apiKey === undefined) throw new Error('OPENAI_API_KEY is not set');
const openai = createOpenAICompatModelProvider({ baseURL: BASE_URLS.OPENAI, apiKey, metadata: { id: 'openai', region: 'us', models: [ { name: 'gpt-4o-mini', contextWindow: 128_000, features: ['tool-use', 'structured-output'], // Rates from the provider's published pricing. cost: { promptUsdPer1kTokens: 0.00015, completionUsdPer1kTokens: 0.0006 }, }, ], }, clientOptions: { timeout: 30_000 },});
const { registry } = createProviderRegistry();registry.register(tenantId, openai);
const result = await openai.invoke({ model: 'gpt-4o-mini', messages: [{ role: 'user', content: 'What is the status of order 1042?' }], tools: [ { name: 'acme.orders.lookup', description: 'Fetch an order by id.', inputSchema: { type: 'object', properties: { orderId: { type: 'string' } }, required: ['orderId'] }, }, ],});for (const call of result.message.toolCalls ?? []) { console.log(call.name, call.arguments); // acme.orders.lookup { orderId: '1042' }}console.log(result.finishReason, result.costUsd);A local runner uses the same factory: baseURL: BASE_URLS.OLLAMA_LOCAL, apiKey: 'unused', and zero cost rates in metadata.
Non-goals
Section titled “Non-goals”- Streaming. Requests are sent with
stream: false;invoke()resolves with the complete response. - Pricing discovery. Cost rates are caller-supplied per model.
- Cached-token accounting.
usagereports prompt and completion tokens only; all prompt tokens are billed atpromptUsdPer1kTokens.
Related
Section titled “Related”@kindgi/capabilities—ModelProvider,ProviderMetadata, and the provider registry.@kindgi/adapter-model-anthropic— the Anthropic Messages API adapter, with a lazily resolved API key.@kindgi/adapter-model-in-process— local models inside the Node.js process, no HTTP endpoint.