Alibaba

Qwen3.8 Flash

Qwen3.8 Flash is Alibaba's fast, low-cost model for high-volume work. It reads images, understands video and reasons step by step before it answers. It takes up to 1M tokens of context and writes up to 131.07K per response.

Vision inputVideo inputReasoningStreaming

Output price

$0.2867

USD / 1M tokens

39% below list price

Add $5, get $5 more on your first top-up.

Same Qwen3.8 Flash. 39% less.

Price
USD / 1M tokens, next to Alibaba's list price.
TokensList pricePromptsForLess
Input$0.15$0.0915
Output$0.47$0.2867
Your monthly bill
Enter your monthly usage to see the difference.
At list price$2.44
With PromptsForLess$1.4884

That's $0.9516 saved every month.

USD, before tax and prompt caching.

Specs.

Model ID
qwen3.8-flash
Made by
Alibaba
Context window
1M tokens
Max output
131.07K tokens
Endpoint
/v1/chat/completions
Billing
Per token, from one balance
Built for
  • Text and code. Drafting, summarising, extraction, chat and code, with the output streamed back as it is generated.
  • Reasoning. Works through multi-step problems before answering, for planning, analysis and agent workflows.
  • Vision. Send screenshots, photos, charts and documents alongside your prompt.
  • Video. Pass video in and ask questions about what happens in it.

Edge gateways

Requests enter at the gateway nearest your app and go straight to the model, so routing adds very little to each call.

One key, every model

Switch models by changing a string. One prepaid balance covers all of them, billed per token.

Nothing stored

We keep token counts for billing, never your prompts or responses.

Call Qwen3.8 Flash in a minute.

Point any OpenAI-compatible client at our base URL, keep your key in an environment variable, and use the model ID qwen3.8-flash.

Copy
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.promptsforless.com/v1",
    api_key=os.environ["PFL_API_KEY"],
)

response = client.chat.completions.create(
    model="qwen3.8-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

More from Alibaba.

Questions.

How much does Qwen3.8 Flash cost?

$0.2867 / 1M tokens output and $0.0915 / 1M tokens input, 39% below Alibaba's list price. You pay per token from a prepaid balance, with no subscription or minimum. We match your first $5 top-up, so you start with $10 of credit.

Is it the real Qwen3.8 Flash?

Yes. Your requests run on Alibaba's Qwen3.8 Flash, not a distilled copy or a substitute. The discount comes from how we buy capacity, not from changing the model.

Which tools can use Qwen3.8 Flash?

Anything that speaks the OpenAI API: set the base URL to our endpoint and use the model ID qwen3.8-flash. We have setup guides for Cursor, Codex, OpenCode, the OpenAI and Vercel AI SDKs, LangChain and more.

How fast is it?

Requests go through our edge gateways, so the network hop adds very little on top of Qwen3.8 Flash's own generation time. Responses stream back token by token.