Skip to content

Comment on DuckDB – Data power tools for your laptop, now in Clojure (2023)

Comments

At Cronitor we use ClickHouse, but we're leaving it behind for our next product and building directly on Parquet and DuckDB.

We think the future of observability in the AI age is self-hosted directly on NVMe backed by cheap and limitless object storage. I don't want to send customer conversations and agent thoughts to a giant multi-tenant borg SaaS database like Sentry or BetterStack.

I’m confused as an old school storage guy.

self-hosted directly on NVMe backed by cheap and limitless object storage.

How is that?! NVMe is a protocol for fast PCIE based local storage or NVMe fabric which is PCIe over network. Object storage(in the sense of S3, R2 etc) are usually networked horizontally scaling non posix bucketed storage. They are much much slower because of the network calls..

How would NVMe map to something like S3? And what for?

I'm just guessing but observability is often looking at small slices of hot data and then there's a vast set of cold data that is occasionally needed. Sounds like NVMe is for the hot data cache and object storage is for cold data.

Yes, you nailed it. Old data is important sometimes, like when a problem has been identified and investigated, but most workloads are looking at the current state of the system. So we keep the hot data cached locally (not nas/ebs) and s3 is always the source of truth. DuckDB over parquet files on a local ssd is fast enough you don’t need a traditional database.

We are using DuckLake with a “lakehouse” architecture for the observability agent.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.