Skip to content

Comment on What Is Differentiable Programming?

Comments

I have to be honest, I don't think this is a good explanation. I don't know what differential programming is, but I'm fairly sure I have the mathematical background to understand it. But I didn't come away from this article with any confidence that I'm following along.

On a superficial level it seems like it:

1. Generalizes deep learning to an optimization function on decomposable input, and

2. Reduces the number of parameters required to learn the input by exploiting the structure of the input, thereby making learning more efficient.

Is that correct? Is it completely off? What am I missing? Is there any more meat to the article than this?

Could someone who has upvoted this (and ideally understands the topic well) provide a different explanation of the concept? It would be great if I could see a real world example (even a relatively trivial one) represented in both the traditional matrix computation form and the sexy new differentiable form.

cf

Let me try to give a high-level explanation. As we have had success with using neural networks for different tasks, people began to ask what are the properties that make neural networks work so well.

When I mean work well, I mean a family of models which are generally easy to describe as well as be generally easy to train. Nearly all neural network models fit this bill as a they can be described with a loss function. A loss function being something which scores how accurately the current NN performs with a particular set of parameters. This loss function is then usually differentiable, and thanks to autodifferentiation tools, you can obtain the derivative of this function with respect to some parameters in this function. With this derivative, you can go ahead and run SGD to find the optimal parameters. We have now both empirical and theoretical evidence that SGD works fairly well for optimizing functions, particularly the high-dimensional ones neural networks describe.

So, the insight is we need a name to describe all these models which are differentiable and nice to train because of that. So now we can say what really ties CNNs, MLPs, RNNs, etc is not some biological metaphor, but that these are all expressible as some loss function we can find the minimum of using SGD.

Supervised learning is a classic inverse problem -- you know the desired outcome, but not how to get there. (think control systems)

Suppose you had a flexible model in whose parameter space you could search. A naive algorithmic approach would be to treat this as a brute force search problem. Maybe you can be slightly more clever algorithmically (better search algorithms, heuristics, etc)

But can you somehow be really clever and exploit the fact that the parameter space is continuous, with a smooth cost function? That's what gradient descent basically is. Lots of bells and whistles on top, to exploit extra problem structure.

Now, let's go back to our family of models. They just need to be transparent to gradient information -- then you could use simple ideas at the heart of adaptive control theory. Suddenly, the fact that autodifferentiation techniques can differentiate arbitrary computational algorithms (sibling comment explains how compositionality helps this) sounds mouth-watering! That's what makes a lot of people excited -- your model family just expanded to programs performing arbitrary computations on continuous parameters (including loops!) -- not just some silly function approximator like a natural network (exaggerating to make my point)

Last I checked, differential programming involved exploding matrix sizes at 2^n in the number of memory bytes that the differentiable program could access. Is the situation any better now? If not it seems kind of like training a function approximator map the present cycle's memory and register state to the i+1th state.

Differentiable. Differential programming is something else.

How does auto-differentiation apply to loop indices? Is this really any more computationally powerful than techniques used to solve integer linear programmes?

in that regard it’s mostly in the first few paragraphs but it’s subtley stated...

If you have a model and somecparameter space on it your trying to optimise then you can’t do much other than search the whole parameter space or approximate such a search intelligently.

However, if your model is differeniable with respect to its parameters then it means you’ve got gradient ... and if you’ve got gradient then you can cut out a lot of searching because you can use gradient ascent and other such techniques .

So basically: differentiable means can be differentiated which means has gradient which means search less.

The rest of the article basically says that’d be nice because there is a way to think about differentiability modularly and therefore there may be a route to Lego things together from across fields to make uber models which still optimise well but use specialist knowledge.

The parameter space stuff just means that less knowledge means more data and more parameters which is bad.

The basic idea is to have an expressive programming language where all constructs are differentiable. Since the composition of diffeomorphisms is a diffeomorphism, large programs (like ray-tracers) will be differentiable as a result.

I don't think you mean to use a term as strong as diffeomorphism. A diffeomorphism is differentiable, sure, but it is also invertable with a differentiable inverse. They have a lot of properties that machine learning honestly does not want, like preservation of dimension.

The basic idea is to have an expressive programming language where all constructs are differentiable. Since the composition of diffeomorphisms is a diffeomorphism, (...)

No. It has nothing to do with diffeomorphisms (which are necessarily between spaces of the same dimension), but with piecewise differentiable functions.

Thanks, I was also a little confused about that in the parent comment.

Wow, I wish you'd written the article. Thank you. I can only imagine what you'd be able to explain if you had as much space as the author...

It sounds like you're just describing an evolution of functional programming; where each area of computation in the program is contextually generalizable, because its set of inputs and outputs is smooth and differentiable.

One followup question: are you sure the maps need to be (or are generally intended to be) isomorphic? That strikes me as very limiting. I follow your point about compositions of diffeomorphisms being diffeomorphisms themselves, but do you need that invertibility to make this paradigm work?

Differential programming is not tied to functional programming. The Python library PyTorch enables differential programming since you can use almost arbitrary python code and differentiate it, including if-then control flow, allowing you to use gradient descent to optimize the parameters of your model.

Traditionally deep learning just meant a sequence of functions applied compositionally (hence the "deep") where each function (termed a layer) is a matrix multiply followed by some well-behaved non-linear function ("activation function"). These were differentiable by design and optimized using gradient descent.

But now the models we want to build are more complex structurally than this merely sequential composition of functions. We want to be able to use control flow, accept multiple inputs, return multiple outputs, etc but we still want the model to be differentiable so we can use an iterative optimization procedure like gradient descent. So this extension from what deep learning traditionally meant (a fairly restrictive class of sequential function compositions) to complex, branching models are now termed differentiable programs.

Thank you, this added a lot more clarity.

For a concrete example, our recent paper RenderNet learns end to end differentiable ray tracing operations like ambient occlusion and normal maps which can be used in phong shading.

No “coding” required.

https://github.com/thunguyenphuoc/RenderNet https://papers.nips.cc/paper/8014-rendernet-a-deep-convoluti...

Since the composition of diffeomorphisms is a diffeomorphism, large programs (like ray-tracers) will be differentiable as a result.

That's just bs. If you just read the ray-tracer article it clearly states that f(scene parameters) -> image is not differentiable -- simply put because of sharp edges of objects being rendered. Also the word `diffeomorphism` makes no sense at all in this context.

Could someone (...) provide a different explanation of the concept?

Yes. It has nothing to do with deep learning, and nothing to do with optimization (although it is useful for both). It is just a way to write your functions so that they are easy to differentiate.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.