Skip to content

Comment on The Big Data Brain Drain: Why Science is in Trouble

Comments

I'm a software developer working with big data, but I think the premise of this article ("the ability to effectively process data is superseding other more classical modes of research") is simply false.

The example problem domain ("automated language translation") is actually a stellar counter-example to the claim. Has anyone actually tried to use Google Translate for anything sophisticated? It's still truly horrible, by human standards. The field needs more research and deeper conceptual understanding, not less.

There may be some problems that can be solved by throwing software/hardware/data at them, but I don't think this is a good paradigm for the big unsolved problems in general.

The example problem domain ("automated language translation") is actually a stellar counter-example to the claim. Has anyone actually tried to use Google Translate for anything sophisticated? It's still truly horrible, by human standards. The field needs more research and deeper conceptual understanding, not less.

Have you compared Google Translate with the previous attempts to do automated translation based on conceptual understanding?

There is a reason why Google Translate would be claimed to be a success.

Hmm,

Being the best of a bad lot isn't enough.

Perhaps it is only a human reflex to believe that some contemplation is needed to solve problems that have resisted mounds of data being thrown on them. But being not-coincidentally human, I happen to find it plausible.

Google Translate is by no means a counter-example. Certainly, it's flawed and imperfect. But Google Translate is the state of the art. You will not find any existing purely automated system that does the task better. Sure, some hypothetical, possibly AI-complete system with rich language understanding would do better. Good luck building that, especially without processing a huge corpus.

Yes you will. You just don't have free access to it online.

Why didn't you name it then?

What about "deep learning" combined with lots of data ?

teams using the technique have won some competitions in kaggle , while doing little feature engineering(which is usually the part which you put domain knowledge) .

And it had shown better results than current systems for big problems like voice recognition and image recognition(the famous google cat experiment).

Check out this essay: http://cacm.acm.org/magazines/2010/9/98038-science-has-only-...

"I believe that science still has only two legs—theory and experimentation. The "four legs" viewpoint seems to imply the scientific method has changed in a fundamental way. I contend it is not the scientific method that has changed, but rather how it is being carried out."

What about Google Search? Its all a black box these days. The only way this changes is the data needs to be opened up.

Giving a few hundred nerds in an ivory tower access to all that hardware and data just makes the black box blacker.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.