6 Comments
User's avatar
Yves Junqueira's avatar

If inference is commodity, my intuition is to stick to “open” hosted providers like the million providers in OpenRouter.

Would love to hear your thoughts about this path as well.

Julian Harris's avatar

Openrouter is sneaky. They offer free as a taster but it’s severely rate-limited. It’s not really free in any production sense. Even for local stuff it was so slow it was not really usable.

Yves Junqueira's avatar

Sorry I meant OpenRouter paid service as an alternative to buying your hardware.

My use case isn't just coding, but doing inference for clients.

Rather than buying hardware at a time of uncertainty, I curious if I can search the market for providers with the right hardware and good enough setup, then pay for tokens.

It's an alternative between paying anthropic/openai/google and building your own cluster.

OpenRouter has a list of providers that offer open weight models.

You could use OpenRouter for a while, simply to identify high quality hosting providers, then contract with those directly when you are happy — thus avoiding the openrouter tax.

I don't know if this is a good strategy...

Julian Harris's avatar

I would love that to work. When I did it last about 4 months ago I had real reliability problems when trying to use arbitrary models on openrouter. I did a benchmark (as I encourage everyone to do) for a use case and the stddev was often terrible. I hope they get better. Turns out serving reliably at scale is a thing. Almost like you need dedicated specialists to do it. Who knew eh 😉

Yves Junqueira's avatar

What kind of reliability issues did you experience? Latency, throughput, quality or a mix?

Julian Harris's avatar

Quality, latency (time to first token) and throughput.