Yolo-Auto gives me fast, consistent access to quality inference at a fraction of the cost of any other provider I have seen.
Qwen3.8-27B for coding agents.
Flat-rate from $15/month.
Every request runs Qwen3.8-27B with FP8 weights and FP8 KV caching. Get predictable monthly billing, 128K context, and no routine prompt or response retention while keeping the chat-completions stack your tools already use.
- 43Btokens served
- 1.3Mrequests handled
Works with Pi · OpenClaw · Hermes · OpenCode · Cline · OpenAI SDK
Paid plans are for normal interactive use with rolling workload limits, bounded queues, and shared-capacity availability.
Real workloads. Real users.
Review current model health and generation speed on the live metrics page, or ask users directly in Discord.
Good speed, high uptime, and very responsive support.
Quotes lightly edited for spelling and punctuation. Usernames are shown with permission.
Change three settings, not your stack.
Set the base URL, add a Yolo-Auto API key, and choose qwen3.8-27b. Keep the familiar request shape.
Test for free, then choose predictable monthly access.
No token credits, no per-token overages, and no surprise usage invoice.
Free
Prove Yolo-Auto works in your stack before you pay.
- Fair use applies
- Designed for 1 coding agent
- Qwen3.8-27B
- 128K context for real prompts and docs
- No card required, free forever
15 requests per day to prove the API works in your stack.
Free forever. No card required and no trial period.
Builder
Best valueFor coding agents and daily development.
- No per-token billing
- Designed for roughly 3-4 coding agents
- 128K context for repositories and long conversations
- Best-effort overflow after Standard limits
- Fair use applies
- Cancel anytime
Example: about 1,739 Standard-lane requests per 24 hours with 32K context and 90% input-cache reuse.
Cancel anytime. No lock-in, no per-token overage.
Pro
For heavier interactive workloads that need more agent capacity and headroom.
- Everything in Builder
- Designed for roughly 5-6 coding agents
- 256K context for larger repositories
- More Standard workload headroom
- Fair use applies
- Cancel anytime
Example: about 4,348 Standard-lane requests per 24 hours with 32K context and 90% input-cache reuse.
Cancel anytime. No lock-in, no per-token overage.
Focused model capacity for everyday agent work.
Works with clients that support a custom base URL and OpenAI-compatible chat completions. Run coding sessions, retries, tool calls, documents, and long-context work without changing request formats.
Pi, OpenClaw, Hermes, OpenCode, Aider, Cline, Roo Code, Continue, and other configurable clients.
OpenAI SDK clients, LangChain, LlamaIndex, curl scripts, and custom chat-completions integrations.
Use focused Qwen capacity for daily work and reserve frontier-model spend for tasks that genuinely need it.
Your prompts stay out of our history and training data.
We process request content to generate a response. We do not collect prompt or response bodies into a conversation history, and we do not use them to train models.
Prompt and response bodies are not routinely retained as a browsable conversation archive.
Your requests and model responses are not fed into model-training pipelines.
We keep operational data such as model, token counts, status, and latency to run the service.
Narrow safety, security, abuse, and legal exceptions apply. Read the Privacy Policy.
The short version before setup.
Does this work with my existing tools?
Works with clients that support a custom base URL and OpenAI-compatible chat completions. See the documentation for tested configurations and supported routes.
What does flat-rate access mean?
Paid plans are not billed per token. They remain subject to plan limits, fair-use controls, bounded queues, and shared-capacity availability.
How much work fits in a plan?
Capacity depends on uncached input, cached input, and generated output. The pricing page includes familiar workload examples and the documentation provides the exact pressure formula.
Do you store my prompts or conversations?
No routine prompt or response retention. Request content is not used for training. Narrow safety, security, abuse, and legal exceptions apply. Operational metadata excludes routine prompt and response text.
Test the API before you pay.
Start with 15 requests per day. Upgrade when Yolo-Auto fits your stack and workload.
Start free →