Back to the Archive

LibraryAI News10 min read

The expensive model got cheaper. Your software did not.

On July 30 OpenAI cut its cheap model by 80 percent and left the flagship sitting at $5 and $30. Over the weekend the flagship moved too, quietly, to $4 and $20, with an expiry date printed beside it and the software you actually pay for not a cent cheaper.

A model booth shows prices cut from $5 and $30 to $4 and $20, while software boxes keep their original prices.

A 40-person company that handles 200 customer conversations a week can now have the best model OpenAI sells read every one of them, tag it, check the account history, and draft the reply, for about $208 a year in model cost. The help desk that sells that same capability as a feature charges 99 cents every time its AI closes a conversation, which on that volume lands somewhere between five and ten thousand dollars a year. Over the weekend the gap got wider, because OpenAI quietly cut the price of its flagship model to $4 per million input tokens and $20 per million output, down from $5 and $30. That is 20 percent off the input and 33 percent off the output. It is also the least interesting number on the page.

The tier that was not supposed to move

Three weeks ago OpenAI ran a round of price cuts and everybody read them the same way. Luna, the cheap high-volume model, dropped 80 percent, from $1 and $6 per million tokens to $0.20 and $1.20. Terra, the middle tier, came down 20 percent. Sol, the frontier model, did not move at all. It sat at $5 and $30 while the bottom of the range got gutted.

ToolDirectory's editors wrote that up on August 13 and called the shape of it plainly: a company defending the commodity end of its range while protecting the premium end. That was the correct read of the evidence available. Chinese models had taken a large share of enterprise token traffic on the open routers, DeepSeek V4 Flash had come out of preview at $0.14 and $0.28, and the obvious defensive move was to fight on price where the fighting was happening and leave the expensive end alone.

Nine days later the expensive end moved. There was no blog post, no launch video, no thread. The model reference page simply started saying $4 and $20, somebody noticed on Saturday morning and posted the docs link to Hacker News, and that is how most of the people paying for it found out. The page carries one more sentence worth reading twice: "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." That sentence is repeated on the pricing page. It is not a price. It is a runway.

What that is actually worth to you

Here is the part that matters for anyone who has built, or is thinking about building, a tool of their own instead of filing a ticket or buying an add-on.

Take the support example above. Two hundred conversations a week, roughly 10,400 a year. A reasonable shape for one pass is about 3,000 tokens in (the ticket, the customer's history, your instructions) and 400 tokens out (a summary, a tag, a draft reply). At the new price that is 1.2 cents of input and 0.8 cents of output, so two cents a conversation, so $208 a year. At last week's price it was $281. The cut saved you seventy-three dollars.

Now hold that against the alternative. Intercom's Fin is priced at $0.99 per outcome, and their own definition of an outcome is broader than most people assume when they see the number: it counts when a customer confirms the issue is resolved, or when they simply do not ask for more help after Fin responds, or when Fin completes a workflow, and their documentation says that includes handoffs. So a conversation where the bot tried, failed, and passed the customer to a human can still be an outcome you pay for. At a 50 percent outcome rate that is about $5,100 a year. If most conversations end up counting, it is closer to $10,300.

Two hundred and eight dollars against five thousand. That is the number people mean when they say the build math changed, and notice what did not do the heavy lifting: the price cut. On the same conservative outcome rate the gap was already eighteen to one last week. This weekend it went to about twenty-five to one, and if prices had gone the other way and risen 20 percent instead, it would still be twenty to one.

Second example, for the RevOps side of the room. Say you build the account-brief tool you have been asking engineering for: pull the company's site, pull the CRM record, write a 150-word brief back to the account. Call it 12,000 tokens in and 500 out. That is 5.8 cents an account, so $290 for five thousand accounts. Last week it was $375.

Both of those savings are real and both of them are rounding errors against a single seat of almost any software you are comparing them to. Which is the thing to actually take away from this weekend. The intelligence was already the cheapest line in the build, and it got cheaper. The expensive parts of doing this yourself are your time, the place you run it, and the maintenance in month seven. None of those moved.

The honest take: read the whole price list, not the headline

Four things on that page will cost you more than the cut just saved you, and three of them are invisible from inside a no-code builder.

The same model has an eightfold spread on it. Look at the pricing page properly. Sol on Batch or Flex is $2 and $10. Sol on Standard is $4 and $20. Sol on Fast mode, which used to be called Priority, is $8 and $40. Identical model, identical tokens, four times the bill depending on which lane the request went down. Stretch it across the long-context boundary as well and the top of the range is $16 and $60 against a floor of $2 and $10, which is eight times the input rate for the same work. If the tool you built defaults to the fast tier because a template you copied set it that way, you have been paying double this whole time and this weekend's cut did not reach you at all. Go find out which lane you are in. That single check is worth more than the discount.

There is a cliff at 272,000 tokens. Cross it in a single request and the entire request reprices at double the input rate and one and a half times the output rate. Not the overage. The whole thing. This is aimed directly at the instinct every non-engineer builder has, which is to stop being clever and paste everything in, because the context window says 1,050,000 and a million sounds like permission. Using more than about a quarter of the advertised window doubles your input rate on that call. A 300,000-token request costs $2.40 in input where a 270,000-token one would have cost $1.08.

The cut is 20 percent off input. Caching is 90 percent off, and most people cannot get at it. Cached input on Sol is $0.40 per million against $4.00 uncached, on the biggest line in most workloads. Go back to the account-brief example: if 8,000 of those 12,000 input tokens are a stable instruction and schema block that never changes between accounts, caching takes the job from $290 to roughly $146. Halved. But getting that discount means structuring the prompt so the unchanging part comes first and stays byte-identical, and a great many no-code and workflow tools give you a single text box and reassemble it however they like. The largest saving available on this price list is the one gated behind control that the people who most need the saving usually do not have.

"At least through November 21" is a fuse. Whatever you cost out at $4 and $20, cost it out again at $5 and $30 before you commit to anything with a contract on the other end. This company cut a tier by 80 percent one month and left the flagship untouched, then cut the flagship nine days later with no announcement. That is a market moving on competitive pressure, week to week, and pressure works in both directions. If your business case only clears at the promotional rate, you do not have a business case, you have a coupon.

And then the thing nobody selling you software wants to discuss. By ToolDirectory's reckoning the marginal cost of a token has fallen roughly an order of magnitude in a year. The price of the finished tools built on those tokens has not moved at all. Their editors counted the catalogue on August 13 and found that 65.1 percent of 2,457 active AI tools require payment and 5 percent are genuinely free, after a summer of cuts like this one. Worse, the categories where free tiers are rarest are exactly ours: sales and RevOps at 18.9 percent free or freemium, customer support at 21.2, analytics at 17.9. Free tiers cluster where the buyer can leave on a Tuesday afternoon. They vanish where switching takes six months and a security review.

The four reasons for that are structural and none of them require anybody to be acting in bad faith. Inference is a minority of what it costs to run a software company. Software is priced against the value it delivers, not against what it costs to make. Most tools sell you a seat, not a token, so cheaper tokens show up as margin rather than as a smaller invoice. And switching costs absorb whatever pressure is left over. Expect the benefit to arrive as wider usage limits at the same price, not as a lower number on the bill, and treat any vendor whose product is a thin wrapper over an API whose output price just fell 33 percent as a fair conversation at renewal.

Who this is genuinely wrong for

If you have not built anything yet, a 20 percent cut on a bill you do not have is worth exactly nothing, and this is not the week to start a project on the strength of it. If you are paying for an AI feature inside software you already own, the API price list has no bearing on your invoice and will not next quarter either. If your volume is low, and most small operators' volume is low, the difference between the old price and the new one on a real workload is a lunch. The people this actually reaches are the ones already metered: already building, already shipping tokens every day, already watching a usage dashboard. For them, the answer is not to celebrate. It is to open the pricing page, work out which lane their requests are in, and fix that.

Five days ago this archive ran the news that the eighteen-dollar server every build-it-yourself piece here has quietly assumed is gone, because AI data centers bought the memory supply and hosting providers have raised prices three times this year. So the room is getting more expensive and the thinking is getting cheaper, at the same time, on the same project. That is a genuinely strange shape for a market and it tells you where the cost of building your own tools is heading: away from the model, back toward the infrastructure and the hours, which is exactly where it has always actually been.

The token is now the cheapest thing in your stack, and it is the only line anyone bothered to discount.

Sources

Every claim above traces back to one of these. Go read them yourself.

  1. 01
    GPT-5.6 Sol model reference

    OpenAI / developers.openai.com / retrieved Aug 24, 2026

  2. 02
    OpenAI API pricing

    OpenAI / developers.openai.com / retrieved Aug 24, 2026

  3. 03
  4. 04
    Intercom pricing

    Intercom / intercom.com / retrieved Aug 24, 2026

  5. 05
    GPT 5.6 Sol 20% price reduction

    Hacker News / news.ycombinator.com / retrieved Aug 24, 2026