JETINFER QWEN3.8-27B · PRAGUE, EU

AI inference infrastructure

More tokens per dollar.

OpenAI-compatible Qwen3.8-27B for agent and coding workloads. Repeated prompt prefix bills at one tenth of the input rate.

See the measurements How you reach it no account, no card · billed through OpenRouter
Effective input cost by prefix cache hit rate · share billed at full rate
0% cachedevery token new
100%
50% cachedtypical agent session
55% −45%
90% cachedlong-running agent
19% −81%

A cached prefix bills at one tenth of the input rate: a 6,500-token prompt with 90% cached bills as 1,235 tokens. The ratio holds at any rate; the rates themselves are set on OpenRouter. Output is never cached. Our own production hit rate is not yet measured. How this was measured.

Qwen3.8-27B int4 W4A16 · 150,000 context · 65,536 max output
1/10rate on repeated prefix
OpenRoutersets the current rate
tools · JSON schemaverified live

What you get

Works with your code

An OpenAI-compatible endpoint. Change the base URL and the key, nothing else. Streaming, tool calling and JSON schema output all work.

Scales with you

Capacity follows the live worker pool, not a hard-coded number. Add a node and the limit grows with it; lose one and traffic moves across without a dropped request.

Nothing is kept

Prompts and answers are never written to disk, logs or analytics, and never used for training. Servers are in the EU.

Numbers you can check

Every throughput and latency figure is published with the prompt length, concurrency and endpoint that produced it. See them.

What is deployed today

One node in Prague: 12 concurrent requests, 150,000 tokens of context, 65,536 max output. Capacity is added as demand arrives.

Exceed a limit and the request fails with an explicit error. Prompts are never silently truncated to fit. Full limits.

How you reach it

We sell through OpenRouter, not directly. You keep your existing OpenRouter account, they bill you, and there is no contract, no signup and no card here. When our provider listing goes live, picking JetInfer on the Qwen3.8-27B page is the entire integration.

Technical questions, measurement methodology and retention terms: hello@jetinfer.com.