Discussion about this post

User's avatar
Interesting Engineering ++'s avatar

Worth remembering:

OpenRouter's shared serving amortizes idle GPU time across thousands of customers.

My whole experiment was cheap because I rented the small models “by the sip” (“pay per use”) — like paying per ride on the bus. Someone else owns the vehicle and keeps it full of other passengers, so idle time costs us nothing.

Therefore, I suggest you also read Avi Chawla's article which covers what happens if you buy the bus — run the models on your own rented GPUs. Then you pay for the hardware by the hour whether it's working or sitting idle, and the standard tools force each small model onto its own GPU. Result: five "cheap" models on five mostly-idle cards can cost as much as the expensive model you were escaping.

So the addition is one warning: small models are only cheap if the hardware under them stays busy. Rent “by the sip” while you're experimenting (as we did — 42 cents); if you ever self-host, making the models share GPUs is its own engineering problem, not a detail.

https://blog.dailydoseofds.com/p/why-small-models-alone-dont-reduce?utm_source=profile&utm_medium=reader2&open=false

No posts

Ready for more?