Anthropic

Claude Opus 4.8 Fast

Claude Opus 4.8 Fast is Anthropic's high-speed variant of its flagship model. It reads images. It takes up to 1M tokens of context and writes up to 128K per response.

Vision inputStreaming

Output price

$31

USD / 1M tokens

38% below list price

Add $5, get $5 more on your first top-up.

Same Claude Opus 4.8 Fast. 38% less.

Price
USD / 1M tokens, next to Anthropic's list price.
TokensList pricePromptsForLess
Input$10$6.20
Output$50$31
Your monthly bill
Enter your monthly usage to see the difference.
At list price$200
With PromptsForLess$124

That's $76 saved every month.

USD, before tax and prompt caching.

Specs.

Model ID
claude-opus-4.8-fast
Made by
Anthropic
Context window
1M tokens
Max output
128K tokens
Endpoint
/v1/chat/completions
Billing
Per token, from one balance
Built for
  • Text and code. Drafting, summarising, extraction, chat and code, with the output streamed back as it is generated.
  • Vision. Send screenshots, photos, charts and documents alongside your prompt.

Edge gateways

Requests enter at the gateway nearest your app and go straight to the model, so routing adds very little to each call.

One key, every model

Switch models by changing a string. One prepaid balance covers all of them, billed per token.

Nothing stored

We keep token counts for billing, never your prompts or responses.

Call Claude Opus 4.8 Fast in a minute.

Point any OpenAI-compatible client at our base URL, keep your key in an environment variable, and use the model ID claude-opus-4.8-fast.

Copy
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.promptsforless.com/v1",
    api_key=os.environ["PFL_API_KEY"],
)

response = client.chat.completions.create(
    model="claude-opus-4.8-fast",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

More from Anthropic.

Questions.

How much does Claude Opus 4.8 Fast cost?

$31 / 1M tokens output and $6.20 / 1M tokens input, 38% below Anthropic's list price. You pay per token from a prepaid balance, with no subscription or minimum. We match your first $5 top-up, so you start with $10 of credit.

Is it the real Claude Opus 4.8 Fast?

Yes. Your requests run on Anthropic's Claude Opus 4.8 Fast, not a distilled copy or a substitute. The discount comes from how we buy capacity, not from changing the model.

Which tools can use Claude Opus 4.8 Fast?

Anything that speaks the OpenAI API: set the base URL to our endpoint and use the model ID claude-opus-4.8-fast. Claude Code and the Anthropic SDKs connect through our Anthropic-compatible endpoint. We have setup guides for Cursor, Codex, OpenCode, the OpenAI and Vercel AI SDKs, LangChain and more.

How fast is it?

Requests go through our edge gateways, so the network hop adds very little on top of Claude Opus 4.8 Fast's own generation time. Responses stream back token by token.