← model directory

huggingface

gemma 4 12B agentic fable5 composer2.5 v2 3.5x tau2

hf/yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF

↓ runs free on your own hardware⚙ tool calling

Open-weight model — runs free on your own GPU via the node runtime (GGUF). 678,289 downloads and 1,576 likes on Hugging Face (yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF).

specs & pricing
Type
text
Provider
huggingface
Model ID
hf/yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
Capabilities
tools
Self-hostable
Yes — runs on your own GPU

Listed in the directory for discovery. textmodels aren't callable through the chat gateway yet — image and video models can run on your own nodes today.

Run it on your own hardware

Estimated GPU memory for gemma 4 12B agentic fable5 composer2.5 v2 3.5x tau2 (12B parameters), for the weights alone — add room for the context cache.

Q4_K_M
~7.6 GBfits a 8 GB card
Q8_0
~12.6 GBfits a 16 GB card
FP16 / BF16
~23.1 GBfits a 24 GB card

Install the agent on the machine and it dials out to the gateway over a single WebSocket — no port forwarding or inbound firewall rules. Requests for this model go to your node first and fail over to the cloud only when it can't serve.

/// initialize

Call gemma 4 12B agentic fable5 composer2.5 v2 3.5x tau2 through one OpenAI-compatible endpoint.

Serve it from hardware you own, with cloud failover when your nodes are busy or offline.

no credit card · 2 nodes free · openai-compatible

Frequently asked questions

Can I run gemma 4 12B agentic fable5 composer2.5 v2 3.5x tau2 on my own hardware?
Yes — gemma 4 12B agentic fable5 composer2.5 v2 3.5x tau2 is an open-weight model you can run on a workstation or on-prem server you own. At Q4_K_M it needs roughly 7.6 GB of GPU memory, plus room for the context cache — it fits a single 8 GB card. Wide Area Intelligence routes requests to your own nodes first and fails over to the cloud only when they can't serve.

More from huggingface

← browse all models