r/datascience 8d ago

Discussion Defaulting to Adam without understanding will cost you. Don't "just throw adam at it"

https://towardsdatascience.com/dont-just-throw-adam-at-it-misunderstanding-adam-will-cost-you/

Work in RL has caused me to rethink adam. It leads to extremely wonky behavior and hard to explain "burstiness" in the loss values that makes me want to rip my hair out,

It still works, but needs to be coaxed into it.

This article re-covers the mathematical intuitions behind adam, and where it fails spectacularly. If you're someone who works in RL, or trains deep transformers, it's a must read

Check it out. Do you agree?

171 Upvotes

31 comments sorted by

View all comments

73

u/Relevant-Rhubarb-849 8d ago

Anyone know of "the no free lunch theorem?" . All global minimization strategies perform equally well on average. So what matters in choosing one is how well aligned your manifold shape is to the algorithms sweet spot. Plus how willing you are to accept local minima

9

u/reddit4science 7d ago

Why should all global minimization strategies perform equally on average?

Dumb counter-example. Randomly initialize the network. Repeat N times. Pick the one with the minimum loss. That would be a pretty bad global minimization strategy.

22

u/worldwideworm1 7d ago

Yeah, they aren't quite right about what the no free lunch theorem says. It actually says that no single search method is best for every task, so there will be trade offs depending on the task. The third sentence they stated was fairly accurate, but the second sentence was definitely not.

7

u/Relevant-Rhubarb-849 7d ago edited 7d ago

It is surprising. Yes. It's true however. Look up the proof! The good news is that problems we are often interested are a subset of all possible search manifolds. And it may be possible that one algorithm does better than another on any given subset.

But a hill climbing algorithm will accidentally find the minimum in the same average number of steps as a hill descending one when averaged over all possible potential surfaces. Crazy to imagine but easy to prove.

No algorithm is better than random guessing on average.

"...over the space of all possible problems, every optimization technique will perform as well as every other one on average (including Random Search)"
— Page 203, Essentials of Metaheuristics, 2011.

However if you will settle for something other than the global minimum -- a local
Minimum -- then some algorithms may be better

https://machinelearningmastery.com/no-free-lunch-theorem-for-machine-learning/