Don't Put All Your Tokens in One Baskettheminimalistdeveloper.com 1 pointrafaelnexus2 months ago5 commentsSaveHideCopy link On HNComments−turtleyacht2moHow do you prepare the fallback if it already took tokens to query the simpler model?−rafaelnexusOP1moI don't know if I understood your question.−plumman12modon't litellm and similar do this already?−santiago-pl1moActually, LiteLLM is the most popular, it was first product of this type, but there is plenty of more efficient and reliable alternate AI Gateways right now. I've written one - GoModel. https://gomodel.enterpilot.io/−rafaelnexusOP1moAbsolutely, a pretty good one.
Comments
How do you prepare the fallback if it already took tokens to query the simpler model?
I don't know if I understood your question.
don't litellm and similar do this already?
Actually, LiteLLM is the most popular, it was first product of this type, but there is plenty of more efficient and reliable alternate AI Gateways right now. I've written one - GoModel. https://gomodel.enterpilot.io/
Absolutely, a pretty good one.