I'm also curious if there's an increase of diagnosis because they work in a medical setting. Either they recognize the symptoms, or casual conversation with a doctor.
These are benign tumors that were asymptomatic, and they only found them because people in the unit complained about some other health concern, and screenings began
The American Cancer Society says that in order to meet the definition of a cancer cluster, occurrences must be the same type, in the same area, with the same cause, and affecting a number of people that's "greater than expected" when a baseline for occurrences is established.
“Nearly 4 out of 10 people in the United States will develop cancer during their lifetimes," the society said on its cancer clusters webpage. "So, it’s not uncommon for several people in a relatively small area to develop cancer around the same time."
The unstated numbers that matter here are many, how many people were thoroughly investigated here, was it the entire staff of the hospital (as many as a thousand, perhaps)? When X many people are thoroughly scanned how common is it for five people to have benign cancers that aren't doing anything, aren't growing, are just there?
If (for example) twenty percent of the time 500 people were scanned, five at least had benign brain cancers, would this report be unusual or suspicious in itself?
That seems circular, like it's not a cancer cluster until we find out that it's a cancer cluster, or it's not a cancer cluster because we didn't determine a common cause, so don't worry too much about there maybe being a common cause that would make it count as a cancer cluster.
This underlines how stats are no substitute for reasoning about mechanisms.
when there are billions of people in the world it is expected that some where several get cancer at the same time.
That would be the answer if we asked whether such a coincidence ever happens in the world. In this specific case, the question is, 'what are the most likely causes?'
To calculate a p-value (roughly spoken), you need to start with a single hypothesis. Then you gather data and the p-value gives you the probability that your data occurs while your hypothesis is false. When you start with a finite set of multiple hypotheses, you need to take that in to account when calculating your p-value.
When you start with data and come up with a hypothesis afterwards, you would have to find the whole potential space of all hypotheses. So, for example, how many hospitals are there? Do you only consider US? Do you only consider nurses or other employees as well? What about only four nurses would that have made it to the news? What about other forms of cancer? What about time? Do you consider the time period of the last 50 years? As you think about what might have made the news, the set of hypotheses grows bigger and bigger and as it approaches infinity, the p-value for any data would approach one. Because when you have a very large set of unlikely hypotheses, the probability that your data accidentally supports one of them is quite large.
P values are common in science for those that don’t know. It measures what the odds are something you observe would happen in just a random sample. Or something like that.
Comments
No reason to think this is anything other than normal stastical variation - at least at this time. not the same type of cancer even.
when there are billions of people in the world it is expected that some where several get cancer at the same time.
I'm also curious if there's an increase of diagnosis because they work in a medical setting. Either they recognize the symptoms, or casual conversation with a doctor.
These are benign tumors that were asymptomatic, and they only found them because people in the unit complained about some other health concern, and screenings began
So, you're probably not far off
The article closes with that note.
The unstated numbers that matter here are many, how many people were thoroughly investigated here, was it the entire staff of the hospital (as many as a thousand, perhaps)? When X many people are thoroughly scanned how common is it for five people to have benign cancers that aren't doing anything, aren't growing, are just there?If (for example) twenty percent of the time 500 people were scanned, five at least had benign brain cancers, would this report be unusual or suspicious in itself?
Right but five people all getting brain cancer is certainly more suspicious than five people getting any cancer.
That seems circular, like it's not a cancer cluster until we find out that it's a cancer cluster, or it's not a cancer cluster because we didn't determine a common cause, so don't worry too much about there maybe being a common cause that would make it count as a cancer cluster.
This underlines how stats are no substitute for reasoning about mechanisms.
You know, if a bunch of nurses got their heads exposed to a radiation burst in an event, would they all get the same type of cancer? Probably not.
It seems unlikely all 5 would only have their heads exposed. It seems it would be more likely they might develop various other cancers.
If you start with the assumption that brain cancer's occurrence is randomly distributed, then sure I guess.
If you assume it isn't randomly distributed, those odds only go up.
Which odds go up, if you would clarify?
The odds "that some where several get cancer at the same time". (And by odds I mean probability, strictly speaking)
nope. they go down for some people and up for other people. if these nurses are the second kind then that is interesting.
The nurses did not have cancer. They had benign tumors.
That would be the answer if we asked whether such a coincidence ever happens in the world. In this specific case, the question is, 'what are the most likely causes?'
What’s the pvalue ?
You can't really put a useful p-value on that.
To calculate a p-value (roughly spoken), you need to start with a single hypothesis. Then you gather data and the p-value gives you the probability that your data occurs while your hypothesis is false. When you start with a finite set of multiple hypotheses, you need to take that in to account when calculating your p-value.
When you start with data and come up with a hypothesis afterwards, you would have to find the whole potential space of all hypotheses. So, for example, how many hospitals are there? Do you only consider US? Do you only consider nurses or other employees as well? What about only four nurses would that have made it to the news? What about other forms of cancer? What about time? Do you consider the time period of the last 50 years? As you think about what might have made the news, the set of hypotheses grows bigger and bigger and as it approaches infinity, the p-value for any data would approach one. Because when you have a very large set of unlikely hypotheses, the probability that your data accidentally supports one of them is quite large.
That's what parent was talking about.
P values are common in science for those that don’t know. It measures what the odds are something you observe would happen in just a random sample. Or something like that.
https://en.m.wikipedia.org/wiki/P-value
https://community.wolfram.com/groups/-/m/t/1824481