- Published on
Deep Learning
- Authors

- Name
- seren-wib
Contents
- 1. The concept of deep learning
- 2. The vanishing gradient problem
- 1. The concept of the vanishing gradient problem
- Why does this happen
- 2. Solutions to the vanishing gradient problem
- 1. Use the ReLU function instead of sigmoid or tanh
- How does a network represent complex functions using ReLU
- ReLU variants
- 3. Weight initialization
- Improved weight initialization methods
- 1. Uniform distribution initialization
- 2. Xavier initialization
- 3. He initialization
- 4. The overfitting problem
- Mitigating overfitting
- 1. Regularization
- 2. Dropout
- 3. Batch normalization
- 5. Weight training methods
1. The concept of deep learning
- Machine learning using neural networks
- The reason it's called deep learning instead of just neural networks is that the layers got deeper
- A deep learning network stacks multiple hidden layers to learn more complex patterns
- The hidden layers in deep learning are larger than those in a (plain) neural network
- Makes the model learn the features it needs from the data by itself
- As layers increase, feature extraction itself can be included in the training process
- A plain neural network model has people manually provide feature extraction (ears, nose, eyes), but a deep learning model extracts features on its own when given the raw image
2. The vanishing gradient problem
1. The concept of the vanishing gradient problem
In a multilayer perceptron with many hidden layers, the error propagated from the output layer shrinks sharply as it goes down to lower layers, so learning doesn't happen
Why does this happen
Multilayer perceptron: the basic form of neural network, an MLP made of input layer - hidden layer - output layer
Propagated error: the network first makes a prediction. The difference between the correct answer and the actual prediction is the error. So the weights are adjusted in the direction that reduces this error. To fix these weights, you have to ask how responsible each weight is for this error. From the output layer to the hidden layer before it, and to the hidden layer before that. This is backpropagation. So the propagated error means the correction signal passed to earlier weights to reduce the error produced at the output. (gradient)
Why this error shrinks: output error → multiply by the derivative of the activation function → multiply by the weight's influence → pass to the previous layer. The derivatives of activation functions like sigmoid or tanh are mostly less than 1. Even at its maximum, sigmoid's derivative is about 0.25. So every step backward keeps multiplying by a small value like 0.25. For example, even if the gradient starts at 1, it becomes 0.25 after one layer, 0.0625 after two layers, and 0.0156 after three. Go deeper and it gets close to 0. That's the vanishing. It doesn't disappear; it just becomes too small to mean anything.
Why learning fails: learning is updating weights. weight = old weight - learning rate × gradient, but since this gradient gets close to 0 as it moves toward the front, learning doesn't happen.
2. Solutions to the vanishing gradient problem
1. Use the ReLU function instead of sigmoid or tanh
- ReLU has a gradient of 1 when the input is positive, so in the positive range the gradient doesn't get too small.
ReLU(x) = max(0, x)
If the input is less than 0, make it 0; if greater than 0, pass it through as is
- With sigmoid, the derivative is about 0.25 at most, so the gradient kept shrinking through multiple layers, but ReLU's derivative is 1 in the positive range. In other words, for neurons with positive input, the backpropagation signal isn't squashed as it passes through.
- ReLU's derivative is also 0 in the negative range, so those neurons may not update well.
How does a network represent complex functions using ReLU
A ReLU network multiplies the input by weights to form the linear expression Wx+b, then keeps only the neurons where that value is positive and turns the negative ones off to 0.
Using linear combinations of these active neurons, a complex function can be approximated as if split into many straight-line/plane pieces.
ReLU variants
- ReLu
- Leacky ReLU
- ELU(Exponential Linear Unit)
- Maxout
- PReLU(Parameteric ReLU)
- Swish
3. Weight initialization
- A factor with a big impact on neural network performance
- Usually random values close to 0 are used as initial weights
Improved weight initialization methods
1. Uniform distribution initialization
Weights are chosen from a uniform distribution range that accounts for the number of input/output nodes
- variables:
- W: weight
n_i: number of input nodesn_{i+1}: number of output nodes
2. Xavier initialization
Initialize by scaling values drawn from the standard normal distribution to the number of input nodes
- Often used with sigmoid and tanh families
variables:
- W: weight
- Z: value drawn from the standard normal distribution
- n_i: number of input nodes
- N(0,1): standard normal distribution with mean 0 and variance 1
3. He initialization
Initialize by scaling standard normal values with a larger variance to suit ReLU-family activation functions
- Often used with the ReLU family
variables:
- W: weight
- Z: value drawn from the standard normal distribution
- n_i: number of input nodes
- N(0,1): standard normal distribution with mean 0 and variance 1
4. The overfitting problem
Mitigating overfitting
1. Regularization
- The larger the weights, the more complex the model. > Penalize weights with large absolute values
- J = Error + alpha * Model complexity
Final error function = how wrong the predictions are + a penalty for how complex the model is
The larger alpha is, the simpler the model; the smaller, the higher the risk of overfitting, so it needs proper tuning
L1 regularization:
- Uses the sum of absolute weights as the penalty
- Tends to drive some weights to 0
- Good for removing unnecessary features
L2 regularization:
- Uses the sum of squared weights as the penalty
- Strongly suppresses large weights
- Keeps all weights small and stable
L1 → |w| → Lasso → some weights 0 → feature selection
L2 → w² → Ridge → suppresses large weights → smooth model
2. Dropout
- During training, randomly select nodes and remove the weight connections before and after the selected nodes
- Select new nodes to drop out for each mini-batch or training cycle (epoch)
- No dropout during inference
- If 1000 data points are split into 10
- Mini-batch: a bundle of 100 data points,
- epoch: going through all 10 mini-batches is 1 epoch,
- iteration: how many repeats until one epoch is done: 10
3. Batch normalization
Cause of the problem Internal covariate shift: as earlier layers train, their weights change, so the data passed to the current layer changes. This slows down training.
Concept First standardize the distribution that got shaken around by earlier layers, then relearn the mean and scale this layer actually needs with γ and β
Batch normalization first sets the intermediate values x_i in a mini-batch to mean 0, standard deviation 1,
then uses γ and β to readjust them to the scale and position the model wants, producing y_i
x_i = value before normalization
x̂_i = value set to mean 0, standard deviation 1
y_i = final value after applying γ and β
γ = scale
β = shift
5. Weight training methods
- Gradient descent
- Gradient descent with momentum
- NAG(Nesterov accelerated gradient)
- Adagrad (Adaptive Gradient Algorithm)
- Adadelta
- RMSprop
- ADAM (Adaptive Moment Estimation)