My prediction is that there will be a useful "deep learning theory" in the future, but it will look a lot more like physics (such as Kaplan's scaling laws) than early 21st century machine learning mathematics/statistics.
I respectfully but completely disagree.
All the enabling technologies of deep learning come from machine learning and statistical learning theory. Stochastic gradient descent, regularization, dimension reduction, bootstrap, bagging, boosting: these techniques remain the fundamental tools in the deep learning toolbox, and a constant source of inspiration for the latest innovations.
Physics has done next to nothing for deep learning in comparison. It's as marginal as the kernel SVM stuff. Just fiddling around with the same old stat mech / network theory / power law stuff the Complex Systems types been doing for the last few decades.
Stat mech works great for materials because materials are relatively simple things and we have relatively simple questions about them. We want to know how they respond to electricity, magnetism, heat, pressure, etc. When you turn to elaborate gadgets for machine translation or image recognition, sure, you can ask similar questions, but the questions just aren't as interesting. You'll get plenty of plots and histograms and power laws out of it, but it's all going to be very superficial. It's not going to tell you how the gadget works.
All that said, there's no way to be sure where the big advances in deep learning theory are going to come from. Your guess might be as good as mine.
Comments
I respectfully but completely disagree.
All the enabling technologies of deep learning come from machine learning and statistical learning theory. Stochastic gradient descent, regularization, dimension reduction, bootstrap, bagging, boosting: these techniques remain the fundamental tools in the deep learning toolbox, and a constant source of inspiration for the latest innovations.
Physics has done next to nothing for deep learning in comparison. It's as marginal as the kernel SVM stuff. Just fiddling around with the same old stat mech / network theory / power law stuff the Complex Systems types been doing for the last few decades.
Stat mech works great for materials because materials are relatively simple things and we have relatively simple questions about them. We want to know how they respond to electricity, magnetism, heat, pressure, etc. When you turn to elaborate gadgets for machine translation or image recognition, sure, you can ask similar questions, but the questions just aren't as interesting. You'll get plenty of plots and histograms and power laws out of it, but it's all going to be very superficial. It's not going to tell you how the gadget works.
All that said, there's no way to be sure where the big advances in deep learning theory are going to come from. Your guess might be as good as mine.