OpenAI introduced GPT-6.1 Sol on September 29 with a striking split in its API prices: $2 per million ordinary input tokens and $0.10 per million cached input tokens. If your agent sends the same substantial context with every request, the cheap cached portion could matter more to your bill than the headline model price. The question is whether your actual requests qualify for that price and whether the work gets finished.
What the new price covers
A request can include instructions, reference material, and the current task. OpenAI lists separate prices for ordinary input, cached input, and output: $2, $0.10, and $10 per million tokens respectively for GPT-6.1 Sol. The cached rate applies to eligible reused input, not to every token in a repeated-looking prompt. Output remains at the output rate. OpenAI says cached input is 50% cheaper than GPT-6 Sol's cached input, while standard input and output are one-fifth of GPT-6 Astra's listed prices.
This makes a recurring agent more interesting than a one-off question. A workflow that repeatedly supplies the same stable instructions and reference material may reuse a larger share of its input. A workflow that rebuilds its prompt each time, changes its references, or produces long outputs may see a different cost profile. You need the usage records to know which case you have.

Compare completed work and actual spend, not just the price of a token.
Compare finished tasks, including retries
OpenAI reports stronger performance than GPT-6 Sol on its selected coding, document, business-workflow, and computer-use evaluations. On AutomationBench, it reports an improvement of 4.8 percentage points over GPT-6 Sol at the same medium reasoning setting. Those are vendor-reported evaluation results, not a promise about your own systems. GitHub announced a gradual Copilot rollout on the same day and says Copilot usage-based billing uses provider list pricing; a Copilot bill should not be assumed to follow the API token accounting described here.
For an API workflow, record the model, reasoning setting, ordinary and cached input, output, retries, and whether the result met the task's acceptance criteria. Divide total spend by accepted tasks rather than by initial attempts. If a cheaper request needs more retries or more human repair, its apparent savings may vanish. Conversely, better completion at the same effort can justify a higher output bill. Human review time belongs beside the API bill, even if it is not part of token pricing.
If a workflow can spend without a clear stopping point, set a spending boundary before optimizing its model.
Run a small comparison before changing the default
Take a representative batch from a recurring workflow and keep its inputs, tools, success criteria, and review process fixed while comparing the current model with GPT-6.1 Sol. Check the actual cached-token fields in your usage data, then inspect failed and retried tasks. This separates a real saving from one projected from a price sheet.
GPT-6.1 Sol is available through the OpenAI API as gpt-6.1-sol. OpenAI says it is also available in ChatGPT Work and Codex for eligible plans but was not yet in Chat at launch. GitHub says its Copilot availability is rolling out gradually and administrators can control model access. Those are separate product surfaces; confirm the one you plan to use before budgeting the switch.
BaristaLabs can help compare one recurring workflow's usage records and accepted outcomes, then decide whether its model choice or its prompt structure deserves attention first.
AI workflow economics
Compare models on completed work
BaristaLabs can help measure the token usage, retries, cache hits, and human review time of one recurring workflow before changing its model.
Bring a representative task and its existing usage records.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.
