← model directory

openai

OpenAI: GPT-3.5 Turbo 16k

openai/gpt-3.5-turbo-16k

⚙ tool calling

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up...

specs & pricing
Type
text
Provider
openai
Model ID
openai/gpt-3.5-turbo-16k
Capabilities
tools
Context window
16K tokens
Input price
$3.30 / 1M tokens
Output price
$4.40 / 1M tokens

Cloud price is billed from prepaid credits when a request fails over to the cloud. Open-weight models run free on GPUs you own — the gateway routes to your nodes first.

What it costs to run

cloud · 1M in + 1M out

$7.70

$3.30 for a million input tokens plus $4.40 for a million output tokens, billed from credits only when a request fails over to the cloud.

your hardware

cloud only

OpenAI: GPT-3.5 Turbo 16k is a hosted model, so it is always served from the cloud. Route easy traffic to an open model on your own hardware and keep this one for the requests that need it.

/// initialize

Call OpenAI: GPT-3.5 Turbo 16k through one OpenAI-compatible endpoint.

Mix it with open models on your own hardware — the gateway routes each request to the cheapest place that can serve it.

no credit card · 2 nodes free · openai-compatible

Frequently asked questions

What is the context window of OpenAI: GPT-3.5 Turbo 16k?
OpenAI: GPT-3.5 Turbo 16k accepts up to 16,385 tokens (16K) of context per request.
How much does OpenAI: GPT-3.5 Turbo 16k cost per million tokens?
Through Wide Area Intelligence, OpenAI: GPT-3.5 Turbo 16k costs $3.30 per 1M input tokens and $4.40 per 1M output tokens when served from the cloud, billed from prepaid credits. A workload of 1M input plus 1M output tokens costs $7.70.

More from openai

← browse all models