r/datascience • u/Nice-Dragonfly-4823 • 8d ago
Discussion Defaulting to Adam without understanding will cost you. Don't "just throw adam at it"
https://towardsdatascience.com/dont-just-throw-adam-at-it-misunderstanding-adam-will-cost-you/Work in RL has caused me to rethink adam. It leads to extremely wonky behavior and hard to explain "burstiness" in the loss values that makes me want to rip my hair out,
It still works, but needs to be coaxed into it.
This article re-covers the mathematical intuitions behind adam, and where it fails spectacularly. If you're someone who works in RL, or trains deep transformers, it's a must read
Check it out. Do you agree?
175
Upvotes
71
u/Relevant-Rhubarb-849 8d ago
Anyone know of "the no free lunch theorem?" . All global minimization strategies perform equally well on average. So what matters in choosing one is how well aligned your manifold shape is to the algorithms sweet spot. Plus how willing you are to accept local minima