Sounds like you're pushing the limits of each product. Looks cool to see. Do you think it is possible to simplify the stack or would you loose too much?
And another question? Do you use varnish on edge? In case, how do you do this? I normally just use bunnycdn for caching.
I'm not an expert, but I don't think realistically we're anywhere close to the limits of what they can do.
Our "edge" is CF right now, but that's likely to change (coincidentally to bunny) because our traffic model doesn't really fit their offering very well.
This actually goes to answer your first question: we've progressively moved more things into our own origin stack because what CF offers us doesn't work out (and in some cases probably wouldn't be easily solved with nginx either).
As an example: CF and Nginx both have basic rate limiting, but AFAIK, neither can track two separate rates out of the box, if at all. We use a relatively high req/period limit, and then a much lower threshold error/period limit which counts any requests that result in a client-caused error (i.e. 4xx errors).
I'm pretty certain we would lose flexibility if we "simplified" by either offloading some aspects to a service, or even if we tried to consolidate several parts into a kitchen-sink tool like nginx.
We mostly use edge (i.e. CDN) caching the same way we expect a browser cache to work (but across users): stuff that we absolutely know isn't changing (i.e. we reference asset (img, js, css) content using URLs based on a content hash). We use our internal Varnish cache for application content, combined with a job in our queue that can purge that content (selectively) when required. Could we use a commercial CDN for that part? Sure in theory, but it's just another thing we won't have control over, and have to develop specific to the CDN we're using at the time, subject to the rules of their system - and it still wouldn't necessarily to ESI.
I really love your idea of doing content hash for (permanent?) cache. I'll definitely do that. Right now we use a semi-permanent cache, but sometimes we screw up and it really bites us. If it is hashed you know for sure it is permanent. It reminds of IPFS, but for your little corner of the internet :D
We started using content hash addressing with our assets build system (i.e. combining CSS & JS into a single URL), so that we'd get maximum cache life without needing to artificially manage cache busting query strings.
The more substantial use now though, is for the "content" of the site (purchasable ebooks/etc and associated artwork) which is constantly growing and occasionally old content gets updated with newer editions/revisions. We use content hashing to actually store the files (free de-duplication) and obviously it then gives us the same caching benefits.
This is so cool Stephen. I'd love to hear more. This is definitely something I could use in one of my clients projects. You wanna schedule maybe a meeting next week?
I'd be happy to talk about it more, but I'll be out of the office all next week on a short break.
Can you drop me an email (it's in my profile) with a rough idea of what you want to discuss, and we could follow up the following week ( with a call/chat?
Comments
Sounds like you're pushing the limits of each product. Looks cool to see. Do you think it is possible to simplify the stack or would you loose too much?
And another question? Do you use varnish on edge? In case, how do you do this? I normally just use bunnycdn for caching.
I'm not an expert, but I don't think realistically we're anywhere close to the limits of what they can do.
Our "edge" is CF right now, but that's likely to change (coincidentally to bunny) because our traffic model doesn't really fit their offering very well.
This actually goes to answer your first question: we've progressively moved more things into our own origin stack because what CF offers us doesn't work out (and in some cases probably wouldn't be easily solved with nginx either).
As an example: CF and Nginx both have basic rate limiting, but AFAIK, neither can track two separate rates out of the box, if at all. We use a relatively high req/period limit, and then a much lower threshold error/period limit which counts any requests that result in a client-caused error (i.e. 4xx errors).
I'm pretty certain we would lose flexibility if we "simplified" by either offloading some aspects to a service, or even if we tried to consolidate several parts into a kitchen-sink tool like nginx.
We mostly use edge (i.e. CDN) caching the same way we expect a browser cache to work (but across users): stuff that we absolutely know isn't changing (i.e. we reference asset (img, js, css) content using URLs based on a content hash). We use our internal Varnish cache for application content, combined with a job in our queue that can purge that content (selectively) when required. Could we use a commercial CDN for that part? Sure in theory, but it's just another thing we won't have control over, and have to develop specific to the CDN we're using at the time, subject to the rules of their system - and it still wouldn't necessarily to ESI.
That's so cool! I actually made a YouTube video on why I don't use CF anymore. Shameless plug here :D
https://youtu.be/TWtGwW_e56w
I really love your idea of doing content hash for (permanent?) cache. I'll definitely do that. Right now we use a semi-permanent cache, but sometimes we screw up and it really bites us. If it is hashed you know for sure it is permanent. It reminds of IPFS, but for your little corner of the internet :D
We started using content hash addressing with our assets build system (i.e. combining CSS & JS into a single URL), so that we'd get maximum cache life without needing to artificially manage cache busting query strings.
The more substantial use now though, is for the "content" of the site (purchasable ebooks/etc and associated artwork) which is constantly growing and occasionally old content gets updated with newer editions/revisions. We use content hashing to actually store the files (free de-duplication) and obviously it then gives us the same caching benefits.
This is so cool Stephen. I'd love to hear more. This is definitely something I could use in one of my clients projects. You wanna schedule maybe a meeting next week?
https://martinbaun.com/book/
I'd be happy to talk about it more, but I'll be out of the office all next week on a short break.
Can you drop me an email (it's in my profile) with a rough idea of what you want to discuss, and we could follow up the following week ( with a call/chat?
Awesome Steven, I'll msg you once you're back. Happy break