← model directory

mistralai

Mistral: Mistral Small 3.2 24B

mistralai/mistral-small-3.2-24b-instruct

↓ runs free on your own hardware⚙ tool calling

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

specs & pricing
Type
text
Provider
mistralai
Model ID
mistralai/mistral-small-3.2-24b-instruct
Capabilities
vision, tools
Context window
256K tokens
Self-hostable
Yes — runs on your own GPU
Input price
$0.10 / 1M tokens
Output price
$0.28 / 1M tokens

Cloud price is billed from prepaid credits when a request fails over to the cloud. Open-weight models run free on GPUs you own — the gateway routes to your nodes first.

What it costs to run

cloud · 1M in + 1M out

$0.38

$0.10 for a million input tokens plus $0.28 for a million output tokens, billed from credits only when a request fails over to the cloud.

your hardware

$0 / token

Served from a workstation or on-prem server you own, there is no per-token fee — the marginal cost is power.

Run it on your own hardware

Estimated GPU memory for Mistral: Mistral Small 3.2 24B (24B parameters), for the weights alone — add room for the context cache.

Q4_K_M
~14.4 GBfits a 16 GB card
Q8_0
~24.4 GBfits a 32 GB card
FP16 / BF16
~45.5 GBfits a 48 GB card

Install the agent on the machine and it dials out to the gateway over a single WebSocket — no port forwarding or inbound firewall rules. Requests for this model go to your node first and fail over to the cloud only when it can't serve.

/// initialize

Call Mistral: Mistral Small 3.2 24B through one OpenAI-compatible endpoint.

Serve it from hardware you own, with cloud failover when your nodes are busy or offline.

no credit card · 2 nodes free · openai-compatible

Frequently asked questions

What is the context window of Mistral: Mistral Small 3.2 24B?
Mistral: Mistral Small 3.2 24B accepts up to 256,000 tokens (256K) of context per request.
How much does Mistral: Mistral Small 3.2 24B cost per million tokens?
Through Wide Area Intelligence, Mistral: Mistral Small 3.2 24B costs $0.10 per 1M input tokens and $0.28 per 1M output tokens when served from the cloud, billed from prepaid credits. A workload of 1M input plus 1M output tokens costs $0.38.
Can I run Mistral: Mistral Small 3.2 24B on my own hardware?
Yes — Mistral: Mistral Small 3.2 24B is an open-weight model you can run on a workstation or on-prem server you own. At Q4_K_M it needs roughly 14.4 GB of GPU memory, plus room for the context cache — it fits a single 16 GB card. Wide Area Intelligence routes requests to your own nodes first and fails over to the cloud only when they can't serve.

More from mistralai

← browse all models