The cost of AI keeps falling

As recently as a few weeks ago, I was worried about the cost of AI rising, and about what that would do to our business. Data centers, accelerators, electricity and memory were all getting harder to secure, and I assumed the price of using the models would follow. That worry sent me back to fact-check myself. Here’s what I learned:
For most of the software era, the infrastructure story was simple: compute, storage and bandwidth all got cheaper, and software scaled on top of them. AI has split the inputs from the output. Demand is still climbing, and the largest technology companies are still building data centers as fast as they can secure power and chips. Accelerators remain hard to manufacture, electricity decides where a new building can go, and memory has been tight enough that Apple pulled its highest memory configuration off the shelf for months. Governments treat semiconductors, energy and AI infrastructure as matters of national security. Through all of that, the price of using a model has kept falling.
On September 22, OpenAI released GPT-6 Sol and Luna. At the standard rates for shorter requests, Sol is $2 per million input tokens and $10 per million output tokens, half the current listed price of GPT-5.6 Sol. Luna, the high-volume tier, is $0.10 and $0.50, against $0.20 and $1.20 for GPT-5.6 Luna. The July price on that tier was higher still. Before OpenAI’s July 30 cut, GPT-5.6 Luna was $1 per million input tokens and $6 per million output tokens. Input on that tier has gone from $1 on GPT-5.6 Luna in July to $0.10 on GPT-6 Luna, a 90% drop in under three months across the two models. Longer requests cost more than these rates.
The same day, Anthropic released Claude Opus 5.5 at $4 and $20 per million input and output tokens, 20% below Opus 5. Anthropic says the model costs less per token and uses fewer tokens per task, and that the two together make a typical workload about 40% cheaper at default settings. Cache reads, which Anthropic says are most of the cost of agentic and coding work, fell from $0.50 to $0.20 per million tokens.
What I learned makes me even more optimistic. Power, chips and memory can all get tighter while the cost of finishing a task still falls. That is a more workable problem for our business than the cost increase I was planning around a few weeks ago.
Cheaper tokens, more expensive inputs
A token is the output of a production process that includes a model, accelerators, memory, electricity, networking, and a growing amount of software aimed at getting more work out of the same hardware. Those inputs can get more expensive while the output gets cheaper. A factory can pay more for electricity and still sell a cheaper widget if a new process turns out twice as many widgets on the same machines.
Several of those process changes are underway together. Mixture-of-experts architectures activate only part of a very large model for each request, quantization reduces the memory and computation a model needs, and prompt caching skips context the system has already processed. Speculative decoding speeds up generation, and better GPU kernels get more out of the same chip.
The labs are also using the models on their own serving stack. OpenAI says that, inside a human-led process, GPT-5.6 Sol rewrote production kernels, and that this kernel work, together with other kernel improvements, cut the end-to-end cost of serving the model by 20%. Those savings helped fund the July price cuts.
The number worth watching has shifted with the prices. A cheaper token does little if the model needs twice as many of them to finish the job, which is why Anthropic’s figure for a typical Opus 5.5 task matters more than the 20% change in the list price.
Data centers, chips and memory
The International Energy Agency estimates that data centers used around 415 TWh of electricity in 2024, about 1.5% of global consumption. Its base case takes that to around 945 TWh by 2030, with AI the largest driver of the increase.
The United States accounts for a large share of the growth. A 2025 update from Lawrence Berkeley National Laboratory, prepared with Department of Energy support, puts U.S. data center use at 192 TWh in 2024, or 4.7% of U.S. electricity. The reference case projects 11.8% by 2030, within a scenario range of 9.5% to 15.3%. In that reference case, more than one in ten kilowatt-hours generated in the United States would go to data centers. The lab is clear that these are projections from current shipment trends, and that the model does not try to price in a break in chip supply, regulation, or available grid power.
The chips inside those buildings still depend on a concentrated supply chain. A May 2024 projection from the Semiconductor Industry Association and Boston Consulting Group has the U.S. share of global advanced-logic capacity, meaning chips on processes smaller than 10nm, rising from 0% in 2022 to 28% by 2032. That is a forecast of fabs announced and planned, and leading-edge production remains slow to add.
Memory is the shortage that showed up in ordinary hardware prices. Manufacturers shifted capacity toward the high-bandwidth memory stacked inside AI accelerators, and the rest of the market felt it. TrendForce raised its forecast for first-quarter 2026 conventional DRAM contract prices to a 90 to 95 percent increase on the previous quarter. The squeeze arrived in the price of machines you would buy to run a model yourself, while API prices moved down.
Efficiency and the work it unlocks
In the 19th century, William Stanley Jevons noticed that more efficient steam engines did not cut Britain’s coal use. Coal use rose, because efficiency made coal worth applying to more jobs.
AI is set up the same way. When a task costs a dollar, you spend it on work you already meant to do, and when it costs a cent, you attach it to work that was not worth the earlier price: agents left running, media files analyzed, documents classified, a model step added to a workflow that did not have one, individual frames read by a system. Inside a company, the figure that matters is the cost of a finished task, which can fall while the total bill rises, because the number of tasks worth doing grows faster than the price declines.
China, and compute in the wrong places
China is often credited with an advantage in this buildout, and some of that description is fair. It has a large electrical system, deep manufacturing capacity, and a state that can direct capital at infrastructure. Access to advanced AI chips is still a hard limit, sharpened by U.S. export controls, and that limit has pushed spending into domestic accelerators and into models that ask less of the hardware companies can get.
China also shows that scarcity depends on where the buildings sit. Utilization at some of the capacity built in western China, under the national plan to put computing near cheaper power, has been reported at 20 to 30 percent. Much of it was sponsored by local governments and investors, far from the labs and companies that wanted the compute. Power is tightening fastest in the United States, and advanced chips remain the tighter constraint in China, with capital and location still binding in both countries.
There is a broader economic argument running alongside the infrastructure one. China already has more productive capacity than household consumption absorbs. On September 19, Huang Yiping, a member of the People’s Bank of China’s monetary policy committee, said wider AI deployment could deepen that imbalance. Money spent on AI infrastructure cannot also be spent somewhere else, and the United States lives with that same tradeoff.
A note on Bitcoin
A friend recently commented on the Bitcoin market as a signal about whether compute is scarce. If miners are still buying hardware at that scale, maybe the hardware is plentiful. If mining gets more expensive, maybe computing is getting scarce. I could not make either reading hold.
Mining and modern AI mostly run on different machines. Bitcoin miners use ASICs built to calculate SHA-256 hashes, and those boards do not become AI servers when you point them at a language model. The Bitcoin price is a weak proxy for the price of AI compute. Miners do control something AI builders want: grid interconnections, power contracts, substations, land, cooling and fiber.
Crusoe began by capturing stranded natural gas and using the electricity to mine Bitcoin. It sold that mining business in 2025 so it could concentrate on AI infrastructure. On September 17 it raised $3.9 billion in a Series F, at a $30.9 billion post-money valuation, for an AI infrastructure business that kept the power and the sites after the mining hardware was sold.
Hardware under the desk
Some of the compute that used to require a data center now fits in an office. Apple’s M5 Ultra Mac Studio can be configured with 512GB of unified memory, with that configuration scheduled for late October. NVIDIA’s DGX Spark puts 128GB of unified memory in a much smaller box, and NVIDIA designs it to run models of up to 200 billion parameters. An organization can buy the hardware once, download an open-weight model, and shift inference cost toward hardware, electricity and operations. That is mostly a fixed cost, where an API bill is tokens times a price.
The memory shortage followed the work onto the desk. In March, Apple removed the 512GB option from the previous Mac Studio. The M5 Ultra brings it back, later than the rest of that lineup, which started reaching customers on September 22. In February, NVIDIA raised the DGX Spark Founders Edition from $3,999 to $4,699, attributed the change to memory supply, and left the hardware the same.
Local hardware got more expensive this year while the variable price of cloud inference fell by half on Sol and by 90% on high-volume input since July. For a lot of workloads, buying the box takes longer to pay back than it did a year ago, which leaves control of the data as the stronger reason to run a model locally.
The open models that dominate download counts are also smaller than the launch headlines. Hugging Face’s summer report says that, among repositories which declare a parameter count, models under a billion parameters account for 83% of all-time downloads, and models above 100 billion account for 1%. Continuous-integration jobs inflate those counts, so the 83% describes what gets fetched, including by automated pipelines, more than it describes fleets running in production, and the pattern that remains is a small model doing one task.
Sensitive data, and where the model runs
As models move further into the work of a company, the material they see gets more sensitive, well past someone asking for a summary of a public page. Models now touch intellectual property, unreleased media, financial information, customer data, source code, internal communications, and the processes that separate one company from another. A frontier API is a reasonable choice for some of that work. Other work should stay in infrastructure the company controls.
A local model, or a model in a cloud account the enterprise runs, is that other choice. Many of the strongest open-weight models now come from China. Hugging Face’s spring report found that Chinese models accounted for about 41% of downloads over the prior year. By summer, other developers had published 151,448 models derived from Alibaba’s Qwen on the Hub.
Running those weights needs a distinction. Downloading a Chinese-developed model and executing verified weights in an isolated environment is a different risk from sending proprietary data to a Chinese API. The prompts do not go back to the organization that trained the model. Provenance, software dependencies, tampered weights, model behavior and license terms are still live questions, and they apply to open models from anywhere. Which of those risks you will accept is an AI data sovereignty decision, separate from which model posted the best score this month.
More than one place to run the work
Some workloads will justify the most capable frontier model on offer. Others will use an inexpensive cloud model, a specialized model, an open model in AWS, Azure or Google Cloud, a model in private infrastructure, or a machine sitting in an office. The same workflow can use more than one of them. A media team might run a small local model over sensitive metadata, a vision model over the pictures and video, and a frontier model when the decision is hard.
Cost will matter, along with latency, quality, security, data residency, sovereignty and whether the endpoint is up when the job runs. Those answers have already moved within weeks, from the July price cuts through the September launches, so a plan that freezes this quarter’s price list for five years will be stale while the hardware is still being depreciated.
How we’re building Flo
Flo is an Agentic Media Platform. The workflow sits above the model that happens to perform a step, so a team can change the intelligence under a job without rebuilding the job. A step can call a different model on cost, quality, latency or policy, which is the practical version of swappable models, with no single model vendor as the spine of the system. A frontier model fits the step where its capability earns the spend, a cheaper model fits where it can finish the work, and a private or local model fits where security, sovereignty or the economics require it. The files can stay in storage the team already runs, and Flo connects to that storage and does the work from there, including media discovery.
Compute can stay scarce while token prices keep falling, and local machines may get cheaper again once memory supply catches up. Open models may keep closing the distance to frontier systems, while energy prices, regulation and security requirements keep moving. The companies that get the most from AI will be the ones that can change the model, and the place it runs, without rebuilding the workflow around them.
See the workflow above the model
Look at how Flo connects models to media discovery and the systems you already run, with the model choice kept separate from the work.
Book a demo