Building a model, called training, is a massive upfront expense that can run into the tens or hundreds of millions of dollars, but it happens relatively few times. Inference, the cost of actually answering each request, is smaller per use but never stops and adds up as millions of people use the model. So a company pays a huge sum to create a model, then keeps paying, bit by bit, every time someone uses it.
For example, Training a flagship model might cost a fortune once, while inference is the steady electricity bill of serving it to users every day.