Data Skeptic

[MINI] The Vanishing Gradient

Jun 30, 2017 · 15 min · 18.3 MB
0:00-15:16

Streams straight from the publisher. podnod never proxies or re-hosts episode audio.

This episode discusses the vanishing gradient - a problem that arises when training deep neural networks in which nearly all the gradients are very close to zero by the time back-propagation has reached the first hidden layer. This makes learning virtually impossible without some clever trick or improved methodology to help earlier layers begin to learn.