"It’s interesting to note that the distribution differs from the the traditional pattern used in the English lanuage: E,T,A,I,O,N,S,H,R,D,L … Some of this can be explained by the fact that domain names are not just for the consumption of English speaking people. Even though other regions have their own domains, since .com has become the lingua franca, many businesses simply default to .com (For those interested, there is an interesting article on Wikipedia about the differing relative frequencies of letters in other languages)."
That may be part of it, but the author doesn’t recognize at all the likelihood the letter I is used more frequently probably due to Apple’s product naming influence, imitation from other companies pre-pending the letter before their prouducts and services, and the fact that ‘I’ is a strong, powerful pronoun.
The "i" does make a difference, sure, but not as big an influence as you might think. You can only put the "i" infront of so many words, and if you look at the initial letter charts, it's not massively dominant there.
Far more important, for instance are the substrings like "FREE" (which can apply to all things, not just computer related) and this has a couple of "E"s, or anything that has the "%ING%" substrings (which is a very common letter combination in the English language)
I won’t cite any sources, but ETAOINSHRDLU is well-established as a fairly accurate English letter frequency. The point in the article is the frequency found in domain names has a "higher" frequency of I’s and a relatively lower number of T’s, despite ETAOINSHRDLU, but doesn’t really explore why. (-ING endings are already taken into consideration with ETAOINSHRDLU.)
Also, I have to disagree with you; there are thousands of companies that have capitalized on the Apple product ecosystem (iSkin, iLounge, iPodResQ, etc.) and in the commonly associated abbreviation of “Internet” to i. I would say there are many more prefixes with ‘i’ than ‘e’ or any other letter.
Please do cite sources; It's what adds weight to your comments and differentiates them from speculation.
Yes ETAOIN SHRDLU CMFWYP VBGKQJ XZ is an accurate distribution of letters in common English text BUT this is a distribution of letters in written English. As it turns out, however, written English is full of very common little glue words, like THE, OF, AND, A, TO, IN, IS, YOU, THAT, IT ...
One third of all printed English material is made up of the top 25 words, and the most common 100 words account for almost half. Domain names are not typically sentences, and are often just one or two words. Instead the frequency of occurence of letters in distinct words should be used. For all distinct words in the English dictionary, this distribution is a little different: ESIARN TOLCDU PMGHBY FVKWZX QJ
Already "T" is much further down the list.
Interestingly, this distibution varies by length of word. By the time we get to words of length 13, for instance the order has changed to IENTS ... (So already the letter "I" is the most common letter for longer words without any influence of Apple).
There may be "thousands of companies" that have added an "I" to their compnay name (though it would help your arugment if you quote sources). But even so, this is dwarfed by the 102 million names. Even tens of thousands of new "I" companies is a fraction of a percent change against this denominator.
There are millions of companies/organisations that have domain names, and not all are tech related.
I'm happy to continue debating and will gladly run any queries you suggest using the entire domain name database and the English language database to generate numbers.
Comments
"It’s interesting to note that the distribution differs from the the traditional pattern used in the English lanuage: E,T,A,I,O,N,S,H,R,D,L … Some of this can be explained by the fact that domain names are not just for the consumption of English speaking people. Even though other regions have their own domains, since .com has become the lingua franca, many businesses simply default to .com (For those interested, there is an interesting article on Wikipedia about the differing relative frequencies of letters in other languages)."
That may be part of it, but the author doesn’t recognize at all the likelihood the letter I is used more frequently probably due to Apple’s product naming influence, imitation from other companies pre-pending the letter before their prouducts and services, and the fact that ‘I’ is a strong, powerful pronoun.
The "i" does make a difference, sure, but not as big an influence as you might think. You can only put the "i" infront of so many words, and if you look at the initial letter charts, it's not massively dominant there.
Far more important, for instance are the substrings like "FREE" (which can apply to all things, not just computer related) and this has a couple of "E"s, or anything that has the "%ING%" substrings (which is a very common letter combination in the English language)
I won’t cite any sources, but ETAOINSHRDLU is well-established as a fairly accurate English letter frequency. The point in the article is the frequency found in domain names has a "higher" frequency of I’s and a relatively lower number of T’s, despite ETAOINSHRDLU, but doesn’t really explore why. (-ING endings are already taken into consideration with ETAOINSHRDLU.)
Also, I have to disagree with you; there are thousands of companies that have capitalized on the Apple product ecosystem (iSkin, iLounge, iPodResQ, etc.) and in the commonly associated abbreviation of “Internet” to i. I would say there are many more prefixes with ‘i’ than ‘e’ or any other letter.
Please do cite sources; It's what adds weight to your comments and differentiates them from speculation.
Yes ETAOIN SHRDLU CMFWYP VBGKQJ XZ is an accurate distribution of letters in common English text BUT this is a distribution of letters in written English. As it turns out, however, written English is full of very common little glue words, like THE, OF, AND, A, TO, IN, IS, YOU, THAT, IT ...
One third of all printed English material is made up of the top 25 words, and the most common 100 words account for almost half. Domain names are not typically sentences, and are often just one or two words. Instead the frequency of occurence of letters in distinct words should be used. For all distinct words in the English dictionary, this distribution is a little different: ESIARN TOLCDU PMGHBY FVKWZX QJ
Already "T" is much further down the list.
Interestingly, this distibution varies by length of word. By the time we get to words of length 13, for instance the order has changed to IENTS ... (So already the letter "I" is the most common letter for longer words without any influence of Apple).
You can read about a full analysis of the distribution of letters and see a complete table of letter frequency against word length here: http://www.datagenetics.com/blog/april12012/index.html
There may be "thousands of companies" that have added an "I" to their compnay name (though it would help your arugment if you quote sources). But even so, this is dwarfed by the 102 million names. Even tens of thousands of new "I" companies is a fraction of a percent change against this denominator.
There are millions of companies/organisations that have domain names, and not all are tech related.
I'm happy to continue debating and will gladly run any queries you suggest using the entire domain name database and the English language database to generate numbers.