Skip to content

Comment on Energy consumption comparison in machine learning platformsparent

Comments

I'd like to know why there is a difference in the final loss at all. If the two networks had the same architecture, used the same loss function, and had random uniform initialization, then 1000 epochs should have them converging on very similar final loss values. Especially if one was able to converge to 3e-4.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.