Models & controls
Capabilities, context, output policy, speed references, controls and token prices come from the same catalog and live metrics used by the CLI, SDKs and Room adapters.
Catalog and live state
| id | model | availability | reads | context | speed | input /M | cached /M | output /M |
|---|---|---|---|---|---|---|---|---|
gemma-4-26b-a4b | Gemma 4 26B A4B | offline | text, images | 128k | not measured | $0.50 | $0.50 | $2.00 |
qwen3.8-27b | Qwen3.8 27B | offline | text, images | 256k | not measured | $1.50 | $1.50 | $6.00 |
Live means an eligible engine is currently registered. Offline means the model cannot be selected. The context shown for a live model is the window reported by the running fleet; otherwise it is the catalog value. A live speed value is the measured average per active request in the latest metrics window. A benchmark value is a historical single-stream reference.
Model reference
Gemma 4 26B A4B gemma-4-26b-a4b
Quick private replies, fast enough for voice
- Best for
- conversation,voice,quick tool use
- reads
- text, images
- Produces
- text, tool calls, reasoning
- Context
- 128k advertised; up to 256k supported by the catalog record
- Output
- 8,192 tokens by default; up to the context remaining after input
- Speed
- No public benchmark; no live rate while the model is offline
- Price
- $0.50 input, $0.50 cached input, $2.00 output per million tokens
Controls
effort — Thinking
Gemma 4 either answers directly or thinks first; its template has one thinking level.
| value | meaning |
|---|---|
none — default | Answers directly, without thinking. |
high | Thinks before answering. |
Wire field: reasoning_effort.
Qwen3.8 27B qwen3.8-27b
Deeper private reasoning on the same attested machine
- Best for
- reasoning,coding,long documents
- reads
- text, images
- Produces
- text, tool calls, reasoning
- Context
- 256k advertised
- Output
- 8,192 tokens by default; up to the context remaining after input
- Speed
- No public benchmark; no live rate while the model is offline
- Price
- $1.50 input, $1.50 cached input, $6.00 output per million tokens
Controls
effort — Thinking effort
How much Qwen3.8 reasons before it answers. Its own default is xhigh; ours is low until measured, as for every unmeasured model.
| value | meaning |
|---|---|
none | Answers directly, without thinking. |
low — default | Brief thinking. |
medium | Moderate thinking. |
xhigh | The model's own full thinking; the slowest and dearest. |
Wire field: reasoning_effort.
Reading speed correctly
Generated tokens per second measures decoding after generation begins. It does not include queue time, prompt processing or time to first token. It is therefore only one part of the latency a person experiences.
- Live average/request is current aggregate generated-token throughput divided by active decoding requests. It reflects current sharing, but remains an interval average rather than a promise to the next request.
- Sustained fleet capacity is learned separately from repeated, stable-concurrency samples. It remains “learning” until queueing, memory pressure, latency degradation or a throughput plateau proves the knee.
- Benchmark is a historical single-stream measurement for the documented serving configuration. It is a comparison reference, not a live result or minimum guarantee.
- Per-request speed changes with prompt length, requested output, reasoning effort, batching, concurrent load and the active serving configuration.
TurboPrivate charges measured input, cached-input and output tokens—not elapsed time or tokens per second. See Pricing & billing for the unit prices and settlement rules.
Control precedence
- A valid request-level wire field applies to that request.
- Otherwise the standing control chosen when the room opened applies.
- Otherwise the model's catalog default applies.
The API returns normalized effective settings when the room opens. Unknown control names and explicit illegal values fail with 400; they are not silently mapped to a convenient default. Generic client aliases listed above are normalized to a legal value.
Context and output
Context is shared by input and output. “Up to remaining context” means the output ceiling decreases as the input grows. A model can also have a smaller default output allowance so an ordinary request does not reserve the entire remaining window. Clients should not infer one room's concurrency or total fleet capacity from the context number.
Reasoning output
Reasoning controls change how much scratch work a model may perform; they do not guarantee a longer or better final answer. The local compatible wires keep raw scratch text opaque and carry it across tool rounds without presenting it as an explanation. OpenAI Chat Completions can expose the model's dedicated reasoning_content field to a developer who deliberately uses that wire.
max can consume the output allowance in reasoning. Use it only when the task justifies the extra latency and tokens. If a direct answer matters more, choose none or low.Images and unsupported input
Use the standard image content block for the chosen wire format. If the selected model does not read images, the adapter replaces the image with a note and the answer indicates that it was not read; it does not pretend vision occurred.
Machine-readable catalog
curl https://api.turboprivate.ai/api/config curl https://api.turboprivate.ai/api/status
/api/config is the stable source for model ids, controls, modalities and public limits. /api/status supplies current availability. Do not copy the table above into application code.