Not commenting on speed, but it seems models and their harnesses have generally gotten much worse over time. My suspicion is that the "Frontier" labs really dont have a strong handle on good quality evals that equal expectations of their users, so they just churn out new models for marketing to sell the shit out of.
Recently cancelled my Codex subscription because I cannot stand 5.6-sol/terra/luna. On top of that, Codex the harness is just so dammed buggy in VS Code.
I just discovered https://tinfoil.sh, which is fully private and attestable AI which works amazing with Kilo Code. Its costing me at API prices but for cheaper to run models which so far feel more consistent then what I get from frontier models.
As someone working in this space: unless they're a very big player with resources to create custom hardware, afaik there's no way to actually prevent the GPU host from seeing the request content. There's no "secure enclave" that sits between a CPU and the GPU, the decrypted payload must necessarily hit system RAM, as would the output. They are, at best, running in TEE VMs which are also pretty horribly broken and don't provide the security that they say they do.
Specifically, the intel page says "Intel® Xeon® 6 processors with Performance‑cores support Intel® TDX Connect technology, enabling confidential computing across the CPU and connected devices including GPUs, Smart NICs, and storage." which you are claiming do not exist....
Not trying to be an evangelist for them, but they seem to pretty open about their tech stack which is hugely differentiating compared to every other AI company....
Comments
Not commenting on speed, but it seems models and their harnesses have generally gotten much worse over time. My suspicion is that the "Frontier" labs really dont have a strong handle on good quality evals that equal expectations of their users, so they just churn out new models for marketing to sell the shit out of.
Recently cancelled my Codex subscription because I cannot stand 5.6-sol/terra/luna. On top of that, Codex the harness is just so dammed buggy in VS Code.
I just discovered https://tinfoil.sh, which is fully private and attestable AI which works amazing with Kilo Code. Its costing me at API prices but for cheaper to run models which so far feel more consistent then what I get from frontier models.
As someone working in this space: unless they're a very big player with resources to create custom hardware, afaik there's no way to actually prevent the GPU host from seeing the request content. There's no "secure enclave" that sits between a CPU and the GPU, the decrypted payload must necessarily hit system RAM, as would the output. They are, at best, running in TEE VMs which are also pretty horribly broken and don't provide the security that they say they do.
Tinfoil is almost certainly lying to you.
You sure? These links are directly from tinfoil.sh's technology page: https://tinfoil.sh/technology
They use these technologies: - https://www.nvidia.com/en-us/data-center/solutions/confident... - https://www.amd.com/en/developer/sev.html - https://www.intel.com/content/www/us/en/products/details/pro...
Specifically, the intel page says "Intel® Xeon® 6 processors with Performance‑cores support Intel® TDX Connect technology, enabling confidential computing across the CPU and connected devices including GPUs, Smart NICs, and storage." which you are claiming do not exist....
Not trying to be an evangelist for them, but they seem to pretty open about their tech stack which is hugely differentiating compared to every other AI company....