Qwen3.8-27B API · OpenAI-compatible

Qwen3.8-27B for coding agents.
Flat-rate from $15/month.

Every request runs Qwen3.8-27B with FP8 weights and FP8 KV caching. Get predictable monthly billing, 128K context, and no routine prompt or response retention while keeping the chat-completions stack your tools already use.

  • 43Btokens served
  • 1.3Mrequests handled
Start free →15 requests/day · no card required Compare plans

Works with Pi · OpenClaw · Hermes · OpenCode · Cline · OpenAI SDK

Paid plans are for normal interactive use with rolling workload limits, bounded queues, and shared-capacity availability.

Proof from live use

Real workloads. Real users.

Review current model health and generation speed on the live metrics page, or ask users directly in Discord.

b
badideaguy @badideaguy

Yolo-Auto gives me fast, consistent access to quality inference at a fraction of the cost of any other provider I have seen.

R
Riknar @riknarr

Good speed, high uptime, and very responsive support.

Quotes lightly edited for spelling and punctuation. Usernames are shown with permission.

One endpoint

Change three settings, not your stack.

Set the base URL, add a Yolo-Auto API key, and choose qwen3.8-27b. Keep the familiar request shape.

POST /v1/chat/completions
Simple pricing

Test for free, then choose predictable monthly access.

No token credits, no per-token overages, and no surprise usage invoice.

Free

$0

Prove Yolo-Auto works in your stack before you pay.

  • Fair use applies
  • Designed for 1 coding agent
  • Qwen3.8-27B
  • 128K context for real prompts and docs
  • No card required, free forever
What this can handle
15 requests per day to prove the API works in your stack.

Free forever. No card required and no trial period.

Builder

Best value
$15/mo

For coding agents and daily development.

  • No per-token billing
  • Designed for roughly 3-4 coding agents
  • 128K context for repositories and long conversations
  • Best-effort overflow after Standard limits
  • Fair use applies
  • Cancel anytime
What this can handle
Example: about 1,739 Standard-lane requests per 24 hours with 32K context and 90% input-cache reuse.

Cancel anytime. No lock-in, no per-token overage.

Pro

$25/mo

For heavier interactive workloads that need more agent capacity and headroom.

  • Everything in Builder
  • Designed for roughly 5-6 coding agents
  • 256K context for larger repositories
  • More Standard workload headroom
  • Fair use applies
  • Cancel anytime
What this can handle
Example: about 4,348 Standard-lane requests per 24 hours with 32K context and 90% input-cache reuse.

Cancel anytime. No lock-in, no per-token overage.

Compare exact limits and estimated costs →
Fits your existing workflow

Focused model capacity for everyday agent work.

Works with clients that support a custom base URL and OpenAI-compatible chat completions. Run coding sessions, retries, tool calls, documents, and long-context work without changing request formats.

Coding-agent clients

Pi, OpenClaw, Hermes, OpenCode, Aider, Cline, Roo Code, Continue, and other configurable clients.

SDKs and frameworks

OpenAI SDK clients, LangChain, LlamaIndex, curl scripts, and custom chat-completions integrations.

Predictable cost

Use focused Qwen capacity for daily work and reserve frontier-model spend for tasks that genuinely need it.

Privacy by default

Your prompts stay out of our history and training data.

We process request content to generate a response. We do not collect prompt or response bodies into a conversation history, and we do not use them to train models.

No stored prompt history

Prompt and response bodies are not routinely retained as a browsable conversation archive.

No training on your requests

Your requests and model responses are not fed into model-training pipelines.

Metadata, not content

We keep operational data such as model, token counts, status, and latency to run the service.

Narrow safety, security, abuse, and legal exceptions apply. Read the Privacy Policy.

FAQ

The short version before setup.

Does this work with my existing tools?

Works with clients that support a custom base URL and OpenAI-compatible chat completions. See the documentation for tested configurations and supported routes.

What does flat-rate access mean?

Paid plans are not billed per token. They remain subject to plan limits, fair-use controls, bounded queues, and shared-capacity availability.

How much work fits in a plan?

Capacity depends on uncached input, cached input, and generated output. The pricing page includes familiar workload examples and the documentation provides the exact pressure formula.

Do you store my prompts or conversations?

No routine prompt or response retention. Request content is not used for training. Narrow safety, security, abuse, and legal exceptions apply. Operational metadata excludes routine prompt and response text.

Test the API before you pay.

Start with 15 requests per day. Upgrade when Yolo-Auto fits your stack and workload.

Start free →