← model directory

huggingface

Qwen3 14B Claude 4.5 Opus High Reasoning Distill

hf/TeichAI/Qwen3-14B-Claude-4.5-Opus-High-Reasoning-Distill-GGUF

↓ runs free on your own hardware⚙ tool calling

Open-weight model — runs free on your own GPU via the node runtime (GGUF). 15,952 downloads and 331 likes on Hugging Face (TeichAI/Qwen3-14B-Claude-4.5-Opus-High-Reasoning-Distill-GGUF).

specs & pricing
Type
text
Provider
huggingface
Model ID
hf/TeichAI/Qwen3-14B-Claude-4.5-Opus-High-Reasoning-Distill-GGUF
Capabilities
tools
Self-hostable
Yes — runs on your own GPU

Listed in the directory for discovery. textmodels aren't callable through the chat gateway yet — image and video models can run on your own nodes today.

Run it on your own hardware

Estimated GPU memory for Qwen3 14B Claude 4.5 Opus High Reasoning Distill (14B parameters), for the weights alone — add room for the context cache.

Q4_K_M
~8.7 GBfits a 12 GB card
Q8_0
~14.6 GBfits a 16 GB card
FP16 / BF16
~26.8 GBfits a 32 GB card

Install the agent on the machine and it dials out to the gateway over a single WebSocket — no port forwarding or inbound firewall rules. Requests for this model go to your node first and fail over to the cloud only when it can't serve.

/// initialize

Call Qwen3 14B Claude 4.5 Opus High Reasoning Distill through one OpenAI-compatible endpoint.

Serve it from hardware you own, with cloud failover when your nodes are busy or offline.

no credit card · 2 nodes free · openai-compatible

Frequently asked questions

Can I run Qwen3 14B Claude 4.5 Opus High Reasoning Distill on my own hardware?
Yes — Qwen3 14B Claude 4.5 Opus High Reasoning Distill is an open-weight model you can run on a workstation or on-prem server you own. At Q4_K_M it needs roughly 8.7 GB of GPU memory, plus room for the context cache — it fits a single 12 GB card. Wide Area Intelligence routes requests to your own nodes first and fails over to the cloud only when they can't serve.

More from huggingface

← browse all models