But, similarly, if you're setting up a data lake, a co-located collection of Parquet files that share the same schema is often colloquially referred to as a "Parquet table". And it's common put some extra layers beyond just squirreling some files away on a disk somewhere to manage and govern these logical units that people call Parquet tables.
I think that the people elsewhere in this thread who are trying to be pedantic about this are maybe more familiar with some of the individual open source technologies that are used in data lake applications than they are with conventions around how these technologies get assembled into a full-fledged data management system in a business setting?
In other words, it seems like people who are trying to talk about the forest are getting downvoted and picked on by a bunch of folks who seem to maintain that there is no forest, only trees.
…if you're setting up a data lake, a co-located collection of Parquet files that share the same schema is often colloquially referred to as a "Parquet table". And it's common put some extra layers beyond just squirreling some files away on a disk somewhere to manage and govern these logical units that people call Parquet tables.
Absolutely, and standardizing that has really cool benefits. You're clearly knowledgeable enough to know where the author means "Delta table" instead of "Delta Lake," "Parquet tables" instead of "Parquet", etc., but not everyone is.
I can understand if the author feels picked on, but I'm sure he knows the bar for technical correctness for developer content marketing is high (especially on HN). Honestly, if I were Mr. Powers I'd be happy for the "strict mode" feedback!
That's not quite what's annoying me. Maybe the more bothersome thing is that some number of people in here have been regularly downvoting people who have done nothing worse than posting factually correct information.
Almost as if people were trying to use HN's voting system as an ersatz referendum on their preferred big data packages rather than as a way to self-moderate the quality of the discussion.
We're supposed to be a crowd that favors mature discussions about technical topics. We can do better.
Comments
But, similarly, if you're setting up a data lake, a co-located collection of Parquet files that share the same schema is often colloquially referred to as a "Parquet table". And it's common put some extra layers beyond just squirreling some files away on a disk somewhere to manage and govern these logical units that people call Parquet tables.
I think that the people elsewhere in this thread who are trying to be pedantic about this are maybe more familiar with some of the individual open source technologies that are used in data lake applications than they are with conventions around how these technologies get assembled into a full-fledged data management system in a business setting?
In other words, it seems like people who are trying to talk about the forest are getting downvoted and picked on by a bunch of folks who seem to maintain that there is no forest, only trees.
Absolutely, and standardizing that has really cool benefits. You're clearly knowledgeable enough to know where the author means "Delta table" instead of "Delta Lake," "Parquet tables" instead of "Parquet", etc., but not everyone is.
I can understand if the author feels picked on, but I'm sure he knows the bar for technical correctness for developer content marketing is high (especially on HN). Honestly, if I were Mr. Powers I'd be happy for the "strict mode" feedback!
That's not quite what's annoying me. Maybe the more bothersome thing is that some number of people in here have been regularly downvoting people who have done nothing worse than posting factually correct information.
Almost as if people were trying to use HN's voting system as an ersatz referendum on their preferred big data packages rather than as a way to self-moderate the quality of the discussion.
We're supposed to be a crowd that favors mature discussions about technical topics. We can do better.