I don't think that math will work out. If you are okay with using only 7M tokens per day, and 7M from a small model, then you don't have very demanding needs. If you don't have demanding needs, I think you're unlikely to be willing to space $1200 ( plus $800 for the rest of the system) to run a local model, when for $10/month you can get an open code subscription (or api access) which won't train on your data (or something of open router).
I know that 7M a day is a pittance for my use cases, and given that you have multiple cards, it wasn't enough for your use case either.
Sure, I just add more cards as I need more throughput, and can combine up to four cards when I need 1M context on a smarter model, but in practice since 3.8 27b came out it is all I use.
Also, unless your sessions are end to end encrypted to a secure enclave, then they are living in plain text -somewhere- and privacy policies tend to change when money is left on the table, if blackhats do not get to the data and sell it first.
Comments
I don't think that math will work out. If you are okay with using only 7M tokens per day, and 7M from a small model, then you don't have very demanding needs. If you don't have demanding needs, I think you're unlikely to be willing to space $1200 ( plus $800 for the rest of the system) to run a local model, when for $10/month you can get an open code subscription (or api access) which won't train on your data (or something of open router).
I know that 7M a day is a pittance for my use cases, and given that you have multiple cards, it wasn't enough for your use case either.
Sure, I just add more cards as I need more throughput, and can combine up to four cards when I need 1M context on a smarter model, but in practice since 3.8 27b came out it is all I use.
Also, unless your sessions are end to end encrypted to a secure enclave, then they are living in plain text -somewhere- and privacy policies tend to change when money is left on the table, if blackhats do not get to the data and sell it first.