GPT-6.1 Sol: $2 input, $0.10 cached and $10 output per million tokens

GPT‑6.1 Sol makes agentic tasks cheaper. Here's how to work out by how much

Sol costs a fifth of GPT‑6 Astra, but what you pay for an agent depends on caching, context length and the number of steps. We go through the pricing and run the numbers on an example.

OpenAI has added GPT‑6.1 Sol to its API for coding, computer use and professional work. The model page pitches it as a cheaper alternative to GPT‑6 Astra with similar quality. The price gap is large: $2 and $10 per million input and output tokens, against $10 and $50 for Astra.

Pricing

Standard prices per 1M tokens for a prompt of up to 272K input tokens:

  • input tokens: $2;
  • cached input: $0.10, which is 5% of the normal price;
  • cache write: $2.50, 1.25 times the normal input price;
  • output tokens: $10.
Bar chart of GPT-6.1 Sol prices for short and long prompts
Standard prices per 1M tokens. Data: OpenAI API pricing table.

If a prompt has more than 272K input tokens, the whole request is billed at the higher rate: input and cache cost twice as much, output 1.5 times as much. Batch and Flex cost half of Standard; Fast mode costs twice as much. The full table is on OpenAI's pricing page.

What the model can do

The context window is 1.05M tokens, with up to 922K for input and up to 128K for the response. The model accepts text and images and responds with text.

Sol works through both the Responses API and Chat Completions, but it can only call tools through the Responses API. If your agent uses search, a shell or MCP servers, Chat Completions won't do.

The model can hand parts of a task to subagents. In the GPT‑6 model guide OpenAI warns that it does this less often than it sometimes should. If a task benefits from running in parallel, say so in the prompt: when to start subagents and how many.

Price the agent, not the request

The price per million tokens tells you little about the bill for an agentic task. On every step the agent resends almost the whole context, repeats tool calls and checks its own results. Each subagent also pays for its own context.

We worked through a hypothetical example: the agent takes 30 steps, each with 50K tokens of context of which only 5K are new, and writes 2K output tokens. Without caching the task costs $3.60. If the unchanging part of the context is written to the cache once and then read at the cached rate, it comes to about $1.15. The same task on GPT‑6 Astra without caching would cost $18.

Cost of a 30-step agent: $3.60 without caching and $1.15 with caching
Our estimate at Standard prices. The actual amount depends on how much of the context really hits the cache.

Before moving a large SEO audit or a server log analysis to Sol, run one real task on both models and compare the totals in the usage dashboard. Also watch whether prompts go past 272K tokens. With long logs that happens quickly, and then the whole request gets more expensive, not just the tail.

No comments yet