Enterprise AI infrastructure has shifted from optimizing training models to serving them, and that changes the economics. Training is a capital project with an endpoint. Inference is a production workload that runs as long as the service is live, with output measured in tokens. This is the tokenomics problem now facing AI operators: once the













Amazon