Skip to content

Comment on Adaptive RAG – dynamic retrieval methods adjustmentparent

Comments

It's still much cheaper to run RAG in production (at least if you are using closed models). I'd love to use the entire context of GPT4, but if I do that in production it'll cost much more than using some RAG-dependent implementation.

But this is just current state. Token costs continue to go down and contexts will continue to get larger.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.