← model directory

qwen

Qwen2.5 72B Instruct

qwen/qwen-2.5-72b-instruct

↓ runs free on your own hardware⚙ tool calling

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

specs & pricing
Type
text
Provider
qwen
Model ID
qwen/qwen-2.5-72b-instruct
Capabilities
tools
Context window
33K tokens
Self-hostable
Yes — runs on your own GPU
Input price
$0.40 / 1M tokens
Output price
$0.44 / 1M tokens

Cloud price is billed from prepaid credits when a request fails over to the cloud. Open-weight models run free on GPUs you own — the gateway routes to your nodes first.

What it costs to run

cloud · 1M in + 1M out

$0.84

$0.40 for a million input tokens plus $0.44 for a million output tokens, billed from credits only when a request fails over to the cloud.

your hardware

$0 / token

Served from a workstation or on-prem server you own, there is no per-token fee — the marginal cost is power.

Run it on your own hardware

Estimated GPU memory for Qwen2.5 72B Instruct (72B parameters), for the weights alone — add room for the context cache.

Q4_K_M
~41.7 GBfits a 48 GB card
Q8_0
~71.8 GBfits a 80 GB card
FP16 / BF16
~134.9 GBmulti-GPU or large unified memory

Install the agent on the machine and it dials out to the gateway over a single WebSocket — no port forwarding or inbound firewall rules. Requests for this model go to your node first and fail over to the cloud only when it can't serve.

/// initialize

Call Qwen2.5 72B Instruct through one OpenAI-compatible endpoint.

Serve it from hardware you own, with cloud failover when your nodes are busy or offline.

no credit card · 2 nodes free · openai-compatible

Frequently asked questions

What is the context window of Qwen2.5 72B Instruct?
Qwen2.5 72B Instruct accepts up to 32,768 tokens (33K) of context per request.
How much does Qwen2.5 72B Instruct cost per million tokens?
Through Wide Area Intelligence, Qwen2.5 72B Instruct costs $0.40 per 1M input tokens and $0.44 per 1M output tokens when served from the cloud, billed from prepaid credits. A workload of 1M input plus 1M output tokens costs $0.84.
Can I run Qwen2.5 72B Instruct on my own hardware?
Yes — Qwen2.5 72B Instruct is an open-weight model you can run on a workstation or on-prem server you own. At Q4_K_M it needs roughly 41.7 GB of GPU memory, plus room for the context cache — it fits a single 48 GB card. Wide Area Intelligence routes requests to your own nodes first and fails over to the cloud only when they can't serve.

More from qwen

← browse all models