TurboPrivate docsCreate API account

Context & caching

Your application owns the conversation. Send the context needed for each answer; the model service can reuse computation for an identical beginning of that context within your account.

One request contract

Direct API, SDK and CLI use the same model service. A Room is an authenticated, encrypted connection, not a stored conversation. Every turn must include all instructions, messages and tool results the model needs. Sending only the latest message does not restore earlier context from the cache.

There is no separate cache-create, cache-update or cache-delete API, no client-selected cache identifier and no promised retention period. Prefix reuse happens during inference when compatible cached computation is available. Updating your application's local state does not send a request or refresh a server cache.

With the Direct API, place messages and supported model parameters inside the ordinary encrypted Room request. Through the SDK or CLI adapter, use the compatible request format below. Do not send plaintext prompts to the public broker's /v1 paths.

Build a reusable prefix

  1. Stable instructions first. Keep system instructions and tool definitions identical and in a deterministic order while their meaning is unchanged.
  2. Preserve history. Append messages and exact tool results in causal order. Do not rewrite previous messages, change tool-call identifiers or reorder messages just for caching.
  3. Fresh state near the end. Put the latest time, file revisions and current state after reusable content, before the new question when appropriate. Clearly identify which observation supersedes older data. Do not promote untrusted retrieved text into system instructions.
  4. Keep the context useful. Correct or remove stale information and compact history when needed. That can reduce cache reuse, but preserving an irrelevant long prompt can cost more and delay output.

Reuse depends on matching the beginning of the model's token sequence. A change near the start can invalidate reuse after that point. Different models, templates, tool definitions or execution settings can also change the sequence. The cache accelerates repeated input processing; it does not return a saved answer or guarantee faster output generation.

JavaScript: two turns and measured cache usage

Install the packages as shown in the Quickstart. This example keeps context in your application and reads actual counters rather than assuming that a repeated request was cached.

import OpenAI from 'openai'
import { connect } from '@turboprivate/sdk'

const room = await connect({
  apiKey: process.env.TURBOPRIVATE_API_KEY,
  model: 'gemma-4-26b-a4b',
})
try {
  const api = new OpenAI({ baseURL: room.baseURL, apiKey: room.apiKey })
  const messages = [
    { role: 'system', content: 'Explain concepts clearly and accurately.' },
    { role: 'user', content: 'What is a database index?' },
  ]
  const first = await api.chat.completions.create({ model: room.model, messages })
  messages.push(first.choices[0].message)
  messages.push({ role: 'user', content: 'When can an index slow down writes?' })
  const second = await api.chat.completions.create({ model: room.model, messages })
  const totalInput = second.usage?.prompt_tokens ?? 0
  const cachedInput = second.usage?.prompt_tokens_details?.cached_tokens ?? null
  console.log({ totalInput, cachedInput,
    uncachedInput: cachedInput === null ? null : totalInput - cachedInput })
} finally {
  await room.close()
}

Zero means the engine reported no cached input tokens. A null counter means the engine did not report cache reads; preserve that distinction in your measurements. Keep tool-call messages, their identifiers and matching results intact in applications that use tools. Each Room accepts one active turn at a time; concurrent work needs separate Rooms.

Usage fields

Optional cache and reasoning counters use null when unreported. Total input and output remain measured integers. Reasoning tokens, when reported, are already included in output tokens.

InterfaceTotal inputCached input
Direct Room receipt / Chat Completionsusage.prompt_tokensusage.prompt_tokens_details.cached_tokens, already included in total input
Responses adapterusage.input_tokensusage.input_tokens_details.cached_tokens, already included in total input
Messages adapterSum of the input counters when cache reads are known; otherwise input_tokens contains total measured inputusage.cache_read_input_tokens; a number excludes those tokens from input_tokens, while null means the split is unknown. Cache creation is reported as zero.

For streaming Chat Completions, consume the stream to completion and read the usage-bearing chunk. Responses supplies usage in response.completed; Messages supplies final usage in message_delta. A direct Room client must also verify the signed receipt and terminal frame. Interrupted work may have estimated usage; do not use that estimate as evidence of a cache hit.

Cached input is not an extra set of tokens. Cost is ((total input − cached input) × input rate + cached input × cached rate + output × output rate) / 1,000,000. Read current rates from https://api.turboprivate.ai/api/config or Pricing & billing. Some models have equal cached and uncached rates. Reservations assume uncached input, so an expected cache hit does not make an otherwise insufficient balance eligible.

Availability and privacy

When cache reads are unknown, all measured input uses the ordinary input rate. The signed receipt preserves null; billing does not turn an unreported counter into a measured cache miss. Exclude unknown counters from cache-hit comparisons rather than substituting zero.

Cache reuse is isolated by billing account, not by Room or API key. Separate keys on the same billing account are not separate cache privacy boundaries. Your application remains responsible for authorization between its own end users. Reuse never authorizes a client to retrieve another client's prompt or history.

Routing favors consistent placement only when eligible servers have equivalent load. A busy server, a restart, eviction or a deployment change can produce a cache miss. Keep working correctly with zero cached tokens; do not retry a completed paid request solely to obtain a cache hit.

Closing a Room removes its live transport state. It does not prove immediate erasure of account-isolated computation retained by a model engine. There is no per-conversation purge guarantee. See Security & privacy for the host trust boundary.

Schedule flexible background work

Read https://api.turboprivate.ai/api/status and select models[modelId]. The catalog's contextLength is the model context limit; reserve space for output and account for your client's transport limits. The load object reports score (1: idle, 10: admission full), observedAtMs, expiresAtMs and backgroundAvailable. Intermediate scores describe pressure against the server's adaptive admission window; they are not hardware utilization percentages. For a model with several servers this is the best eligible server's signal, not a fleet average.

Only consider optional work when backgroundAvailable is true and the observation is current. A missing or expired observation means unknown capacity, not an idle server. Poll conservatively (for example every 15 seconds), then send the ordinary encrypted inference request with scheduling: "background" alongside messages and the model parameters. Chat Completions, Responses and Messages adapters preserve this field. It is service metadata, not a model instruction.

The server checks capacity again when admitting each turn. A recent availability signal does not reserve a slot. A busy response (503 model_busy) means your application should defer optional work and retry later with bounded backoff; do not switch it to foreground just to bypass this check. Background requests use ordinary authentication, account limits, measured usage and billing. They are not free or guaranteed to finish before new interactive work arrives.

Time-critical work uses the default foreground scheduling mode, which still obeys normal admission. The service does not maintain your task queue or decide which conversation facts to preserve. Your application owns those decisions and must remain correct when capacity is unavailable.

Measure the benefit

Compare equivalent workloads with the same model, controls and output limit. Record total input, cached input, output tokens, time to first token, completion time and cost. Use summed cached tokens divided by summed input tokens for a workload's cache ratio. A cache ratio alone cannot establish lower latency, lower cost or server capacity.