the curse of dimensionality is the worst problem that affects machine learning
customers don't want to create training sets large enough to train text classifiers; often the number of documents they need to sort into a category is too small to fit in a category.
As for semantic indexing, it was hard to do in 2005. In 2011 it's easy. DBpedia and Freebase are a chromosome map for the human memome. With large amounts of instance information, it's possible to do things that a big rulebox can't.
These tools are aiming for the market segment that Cyc aimed for, but will use very different methodologies.
Comments
the curse of dimensionality is the worst problem that affects machine learning
customers don't want to create training sets large enough to train text classifiers; often the number of documents they need to sort into a category is too small to fit in a category.
As for semantic indexing, it was hard to do in 2005. In 2011 it's easy. DBpedia and Freebase are a chromosome map for the human memome. With large amounts of instance information, it's possible to do things that a big rulebox can't.
These tools are aiming for the market segment that Cyc aimed for, but will use very different methodologies.