Skip to content

Comment on Making Text Mining Accessible to Any Developer & Non-Expertparent

Comments

the curse of dimensionality is the worst problem that affects machine learning

customers don't want to create training sets large enough to train text classifiers; often the number of documents they need to sort into a category is too small to fit in a category.

As for semantic indexing, it was hard to do in 2005. In 2011 it's easy. DBpedia and Freebase are a chromosome map for the human memome. With large amounts of instance information, it's possible to do things that a big rulebox can't.

These tools are aiming for the market segment that Cyc aimed for, but will use very different methodologies.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.