It is not that it became generous β it started designing how you use AI for the first time. Understand the three things behind this price cut and you understand the whole second half of the industry.
Last month I was still picking apart DeepSeek's price-increase bill, and this month the plot turned again β OpenAI, on the other side of the ocean, suddenly cut the price of its entry model Luna by 80%.
Input price went from 1 US dollar per million tokens to 0.2, and output from 6 dollars to 1.2. Converted into yuan, that is input from about 6.8 yuan down to 1.35, and output from about 40.6 yuan down to 8.1.
The mid-tier model Terra also came down 20%. Only the flagship Sol did not move at all.
By the usual script, a Silicon Valley giant cutting prices to grab share should be another familiar round of price war. But after reading through OpenAI's technical notes, I found the thing really worth watching here is not the price figures at all.
1. For the first time it advises you: do not use the most expensive model
This is the most counter-intuitive part.
For the past years the whole AI industry told the market as hard as it could that its own model was the smartest. And now OpenAI stands up and says:
Many tasks do not actually need the strongest model.
It even gives a very concrete official suggestion: for a complex task, use the flagship Sol first for requirement analysis and solution design, then hand execution, coding and testing to Luna.
The most expensive model thinks; the cheapest model works.
Two years ago this strategy would have been commercial self-denial. OpenAI does it today because it has realised the real rival is not whose model is stronger, but whose model companies cannot do without.
2. Models started optimising models, and the cost really did come down
This detail is more interesting than the price cut itself.
OpenAI says the cut did not come from buying more GPUs, nor from greater scale. The real reason is that the model started taking part in optimising the model.
Led by human engineers, Sol autonomously rewrote and optimised the underlying production kernel: it designed hundreds of experiments itself, raised token generation efficiency, and even took part in monitoring the model training pipeline, stepping in directly when it found a problem.
The result: end-to-end running cost down 20%, token generation efficiency up more than 15%.
Efficiency used to come from engineers. After a model shipped, people would grind away at the inference framework, CUDA and caching strategy, and squeezing out a dozen-odd points in a year was already a good result. Now the model itself takes over the engineering optimisation of low-level code and compute scheduling, iterating around the clock.
The cost really fell, and only then could the price really fall. This is a different thing from burning money to buy market share β it rests on hard technical cost reduction and it is sustainable.
3. The price-cut flywheel is turning agents into the default
Luna, the model discounted hardest this time, is not only a classic small model for classification and extraction.
In OpenAI's positioning it can call tools and carry out multi-step tasks, aimed mainly at high-frequency, cost-sensitive workloads.
So cutting Luna by 80% lowers the threshold for running agents over the long term β review, verification and monitoring jobs that were too expensive to run often can now become routine steps in a workflow.
Add Sol taking part in optimising its own production system, and a new loop appears:
The model helps raise running efficiency β cost falls β price falls
β the model enters more high-frequency workflows β call volume grows β the next round of optimisation
The shape of OpenAI's price-cut flywheel is already visible. It is not selling one model; it is designing a division of labour between models β Sol thinks, Luna works, and what companies buy from now on is not a single model but a whole system.
4. Three reminders for ordinary people
After all that industry logic, what lands on you and me is really just three things:
First, you can no longer choose a model by price alone.
In the past it was use whichever is cheapest; now it is layered competition: light work on a cheap entry model, heavy work on the flagship. OpenAI cutting the entry models while holding the flagship steady follows exactly this logic. When we use AI ourselves, we should learn the same habit β the cheap one works, the expensive one thinks.
Second, the idea that AI has peak hours too will become more and more common.
OpenAI uses tiered pricing; DeepSeek prices peaks and off-peaks separately. The industry is moving from one-size-fits-all to fine-grained, and knowing which model fits which scenario is the real way to save money.
Third, use the window to turn AI into a handy tool.
The price war will settle down eventually. Rather than agonising over a few tenths of a cent, use the time to get familiar with the tools and build your own workflows β that is the real compounding.
In the end, what OpenAI cut this time is not only the price, but the posture with which it enters the second half of AI.
From competing on model capability to competing on ecosystem stickiness; from selling the most expensive model to designing a whole division of labour. When a giant starts fighting with a penetration-rate mindset, it means AI really is turning from a toy into infrastructure.
And what we ordinary people can do is start running one step earlier, before it becomes universal.
*When you use AI, do you distinguish between the model that should be cheap and the one that should be expensive? Tell us in the comments.*
#AI second brain #knowledge management #automation workflows #content creation #productivity
Do you also store a lot of documents and then fail to find what you need when it actually matters?
Do not worry, this is not your fault β your knowledge base is missing an intelligent steward.
I am SavantCat, and I help turn scattered knowledge into assets you can use. If document management, content creation or productivity is giving you trouble, leave a message in the answer library.