For years, the most visible question in AI hardware was how much compute it took to train the next model. Bigger clusters, faster accelerators, and longer training runs became shorthand for progress. But a trained model has no practical value until people and applications can use it.
That second task is inference. Every generated answer, transcribed recording, analyzed camera frame, and automated decision requires a model to run again. Once AI becomes part of a product, the question changes from Can we train this model? to Can we serve it reliably, at an acceptable speed and cost, whenever it is needed?
Training remains essential, expensive, and technically demanding. The point is not that inference has replaced it. The point is that building a capable model and making that model useful at scale are different engineering problems. The industry now has to solve both.
A model is trained once, then used again and again
Training adjusts a model’s parameters using large amounts of data and computation. Inference uses those learned parameters to process a new input and produce a result. Training can be concentrated into a major computing project. Inference is usually spread across the lifetime of the model, wherever and whenever requests arrive.
Consider a company that adds an AI assistant to its software. The model may already exist, whether it was developed in-house or obtained from another provider. The company still needs to decide how to handle incoming requests, how quickly responses must begin, what happens when traffic spikes, and what each successful request costs.
The same questions arise in less visible applications. A camera may run detection repeatedly throughout the day. A transcription service receives files of different lengths. An AI coding tool may make several model calls to complete one user task. In each case, a useful model is only one part of a working service.
This is why the economics of inference cannot be reduced to the cost of a single model run. A rough operational question is: How much does each completed task cost, multiplied by how often the task runs? The answer also depends on whether the output is good enough, whether requests arrive in bursts, and how much capacity sits idle between them.
Cheaper inference changes what developers can build
Inference has become dramatically cheaper for some levels of model capability, though the price of a frontier model is not the same as the price of a smaller one.
Stanford’s 2025 AI Index compared the price of querying models at a fixed performance threshold. It found that the price per million tokens for a model reaching approximately GPT-3.5-level performance on its chosen benchmark fell from $20 in November 2022 to $0.07 in October 2024, a decline of more than 280-fold. That is a historical comparison at one performance threshold, not a claim that every model or workload became 280 times cheaper.
Lower prices can make applications that once seemed uneconomical worth trying. A developer can use a model to classify incoming documents, summarize logs, or assist a search system without building an expensive custom training pipeline first.
But cheaper requests do not automatically produce a smaller total bill. An application that makes one model call per user action has a different cost profile from an agent that makes repeated calls, reads large tool outputs, and revisits a growing conversation. Better capabilities can also create demand for more complex tasks. Whether overall spending rises or falls depends on usage, model choice, and efficiency gains.
The engineering challenge is to find the smallest and least expensive system that meets the task’s actual requirements, rather than assume the largest available model is necessary for every step.
Inference performance is more than tokens per second
Two inference servers can generate the same answer at similar average speeds and still feel very different to a user.
A model first processes the incoming prompt. For a language model, this stage is commonly called prefill. The server then generates the response, typically token by token, in the decode stage. A long prompt can delay the first token; a slow decode can make the remainder of the answer arrive painfully slowly.
Serving systems therefore have several distinct questions to answer:
- Time to first token: How long does the user wait before a response starts?
- Generation speed: How quickly does the response arrive once generation begins?
- Throughput: How much total work can the system finish over time?
- Concurrency: What happens when several requests arrive together?
- Memory use: How much room remains for model weights, context, intermediate data, and other services?
- Reliability: Does the service recover after a restart, full queue, or failed request?
The trade-offs are real. Batching requests can improve total throughput while increasing the wait for an individual request. Keeping a model loaded avoids repeated startup delays but occupies memory during idle periods. Longer context may improve a task, yet it can also increase memory pressure and prompt-processing time.
The DistServe research paper examines the interference between prefill and decode in large-scale language-model serving. Its proposed design separates those stages so resources can be assigned according to their different latency requirements. A small local server does not need that architecture, but the underlying lesson applies at any scale: one headline speed number does not describe the user experience.
This is also why benchmark methodology is changing. In September 2026, MLCommons added end-to-end retrieval-augmented generation and edge agentic workloads to MLPerf Inference v6.1. These tests examine more than one isolated prompt. They reflect systems in which retrieval, tools, repeated model calls, and limited edge-device memory affect the result. Benchmark submissions show performance under defined test conditions, not a guarantee that every application will see the same result.
The bottleneck is not always the accelerator
AI hardware discussions often start with advertised compute capacity. For a deployed service, that is only part of the story.
A language model needs memory for its weights and for information retained while it processes a conversation. As context length and simultaneous requests grow, this additional memory can become a constraint. If the model cannot fit in available accelerator memory, offloading work elsewhere may hurt responsiveness. Faster arithmetic alone cannot remove those limits.
Software matters just as much. Model format, quantization, runtime support, request scheduling, and the handling of intermediate data can all change what a system can serve. NVIDIA’s technical explanation of KV-cache optimization describes how cache size affects memory capacity and bandwidth during inference. The specific techniques depend on the hardware and runtime, but the broader constraint is familiar: capacity and data movement matter alongside raw compute.
There is a practical consequence for developers. Before buying a faster accelerator, identify the problem: does the chosen model fit? Is startup slow? Is prompt processing the bottleneck? Are users waiting behind other requests? Is the application spending time on retrieval, networking, or storage rather than inference itself?
Those questions lead to better decisions than comparing TOPS or tokens per second without a defined workload.
Where inference runs is an application decision
A large data center can serve powerful models to many users and absorb traffic across shared infrastructure. That makes cloud inference attractive when an application needs frontier-level capabilities, rapidly changing capacity, or occasional access to hardware it would not make sense to own.
Other workloads have different constraints. A factory camera might need to react without relying on a continuous internet connection. An organization may want sensitive documents to stay within its own environment. A home automation service may be useful precisely because it can continue operating when an external service is unavailable.
These are reasons to consider on-premises or edge inference, not proof that local deployment is always better. Local hardware has a purchase cost, a memory and performance ceiling, and an ongoing maintenance burden. Privacy also depends on the complete application, including logs, updates, telemetry, and any external services it calls. Keeping model execution local does not automatically make an entire system private.
A useful architecture may combine both approaches. A small local model can handle frequent, bounded tasks, while a remote model is called for requests that genuinely need more capability. In another application, cloud-only deployment may be the simpler and less expensive choice. Placement should follow the workload and its constraints, not a rule that all AI must move to the edge.
Energy and infrastructure are part of the calculation
The scale of AI deployment is also making power and physical infrastructure harder to ignore. The International Energy Agency’s Energy and AI report estimated that all data centers, including AI and non-AI workloads, used about 415 TWh of electricity in 2024. Its 2025 base-case projection put data-center demand at around 945 TWh by 2030, with AI a major driver of the increase.
Those figures are not measurements of inference alone. They include other data-center activity, and the 2030 number is a projection subject to uncertainty about AI adoption, efficiency, and energy supply. They still show why energy availability, utilization, cooling, and completed work per watt belong in infrastructure planning.
The same care is needed when comparing a small local server with a cloud API. An idle device consumes power even when it does no useful work. A cloud service pools hardware among customers, but its pricing also reflects infrastructure and operating costs. An honest comparison needs the full hardware configuration, measured wall power, actual usage, maintenance, and the quality of the result. There is no universal break-even point.
What a small inference server teaches us
These industry-scale questions become easier to understand when reduced to one real system.
In our Jetson Orin Nano and Ollama deployment on Turing Pi 2.5, an 8GB Orin Nano runs Qwen3 4B with verified GPU acceleration. Ollama serves the model through an API that another trusted computer can call over the local network. The model and service persist across a reboot.
That test also makes the limits visible. In its documented configuration, the server produced a warm generation rate of 16.39 tokens per second. One measured request reported 7.33 seconds of model-load time after the model had been unloaded. These are observations from one model, prompt, runtime, power setting, and hardware configuration, not a general Jetson benchmark or a promise about concurrent users.
The lesson is not that a compact board competes with a cloud inference cluster. It does not have to. The lesson is that inference can be treated as a network service: an application submits a request; a dedicated machine runs the model; the result returns to the application. The client does not need its own GPU or a copy of the model.
Turing Pi 2.5 makes that pattern tangible through its four independent compute-module nodes, integrated Ethernet switch, and management features. A Jetson can serve the accelerated workload while another supported module, if needed, runs a web application, queue, or database. The nodes communicate over the network. Their GPU resources and memory do not merge into one larger accelerator, and a second module is not required just to call the API.
That is a modest, useful example of a much wider infrastructure question: how do we assign each part of an AI application to the resources it actually needs?
The next AI infrastructure question
Training determines what a model can learn. Inference determines whether people can put that capability to work, repeatedly and under real operating constraints.
The important progress will not come from one answer to where AI should run. It will come from better models, more efficient software, hardware suited to particular tasks, and clearer measurements of latency, quality, throughput, power, and total cost. Cloud clusters and small local systems address different parts of that problem, and many applications will use both.
For anyone building an AI product, the useful starting point is no longer just Which model is most capable? It is also Which model is capable enough for this task, and what does it take to serve that model well?