logo
Published on

Deep Learning

Authors
  • avatar
    Name
    seren-wib
    Twitter
Contents

1. The concept of deep learning

  • Machine learning using neural networks
  • The reason it's called deep learning instead of just neural networks is that the layers got deeper
  • A deep learning network stacks multiple hidden layers to learn more complex patterns
  • The hidden layers in deep learning are larger than those in a (plain) neural network
  • Makes the model learn the features it needs from the data by itself
  • As layers increase, feature extraction itself can be included in the training process
  • A plain neural network model has people manually provide feature extraction (ears, nose, eyes), but a deep learning model extracts features on its own when given the raw image

2. The vanishing gradient problem

1. The concept of the vanishing gradient problem

In a multilayer perceptron with many hidden layers, the error propagated from the output layer shrinks sharply as it goes down to lower layers, so learning doesn't happen

Why does this happen

  1. Multilayer perceptron: the basic form of neural network, an MLP made of input layer - hidden layer - output layer

  2. Propagated error: the network first makes a prediction. The difference between the correct answer and the actual prediction is the error. So the weights are adjusted in the direction that reduces this error. To fix these weights, you have to ask how responsible each weight is for this error. From the output layer to the hidden layer before it, and to the hidden layer before that. This is backpropagation. So the propagated error means the correction signal passed to earlier weights to reduce the error produced at the output. (gradient)

  3. Why this error shrinks: output error → multiply by the derivative of the activation function → multiply by the weight's influence → pass to the previous layer. The derivatives of activation functions like sigmoid or tanh are mostly less than 1. Even at its maximum, sigmoid's derivative is about 0.25. So every step backward keeps multiplying by a small value like 0.25. For example, even if the gradient starts at 1, it becomes 0.25 after one layer, 0.0625 after two layers, and 0.0156 after three. Go deeper and it gets close to 0. That's the vanishing. It doesn't disappear; it just becomes too small to mean anything.

  4. Why learning fails: learning is updating weights. weight = old weight - learning rate × gradient, but since this gradient gets close to 0 as it moves toward the front, learning doesn't happen.

2. Solutions to the vanishing gradient problem

1. Use the ReLU function instead of sigmoid or tanh

  • ReLU has a gradient of 1 when the input is positive, so in the positive range the gradient doesn't get too small.
ReLU(x) = max(0, x)

If the input is less than 0, make it 0; if greater than 0, pass it through as is
  • With sigmoid, the derivative is about 0.25 at most, so the gradient kept shrinking through multiple layers, but ReLU's derivative is 1 in the positive range. In other words, for neurons with positive input, the backpropagation signal isn't squashed as it passes through.
  • ReLU's derivative is also 0 in the negative range, so those neurons may not update well.

How does a network represent complex functions using ReLU

A ReLU network multiplies the input by weights to form the linear expression Wx+b, then keeps only the neurons where that value is positive and turns the negative ones off to 0.

Using linear combinations of these active neurons, a complex function can be approximated as if split into many straight-line/plane pieces.

ReLU variants

  1. ReLu
  2. Leacky ReLU
  3. ELU(Exponential Linear Unit)
  4. Maxout
  5. PReLU(Parameteric ReLU)
  6. Swish

3. Weight initialization

  • A factor with a big impact on neural network performance
  • Usually random values close to 0 are used as initial weights

Improved weight initialization methods

1. Uniform distribution initialization

Weights are chosen from a uniform distribution range that accounts for the number of input/output nodes

W∼U(−6ni+ni+1,6ni+ni+1)W \sim U\left( -\sqrt{\frac{6}{n_i+n_{i+1}}}, \sqrt{\frac{6}{n_i+n_{i+1}}} \right)
  • variables:
    • W: weight
    • n_i: number of input nodes
    • n_{i+1}: number of output nodes

2. Xavier initialization

Initialize by scaling values drawn from the standard normal distribution to the number of input nodes

  • Often used with sigmoid and tanh families
W=Z⋅1ni,Z∼N(0,1)W = Z \cdot \sqrt{\frac{1}{n_i}}, \quad Z \sim N(0,1)

variables:

  • W: weight
  • Z: value drawn from the standard normal distribution
  • n_i: number of input nodes
  • N(0,1): standard normal distribution with mean 0 and variance 1

3. He initialization

Initialize by scaling standard normal values with a larger variance to suit ReLU-family activation functions

  • Often used with the ReLU family
W=Z⋅2ni,Z∼N(0,1)W = Z \cdot \sqrt{\frac{2}{n_i}}, \quad Z \sim N(0,1)

variables:

  • W: weight
  • Z: value drawn from the standard normal distribution
  • n_i: number of input nodes
  • N(0,1): standard normal distribution with mean 0 and variance 1

4. The overfitting problem

Mitigating overfitting

1. Regularization

  • The larger the weights, the more complex the model. > Penalize weights with large absolute values
  • J = Error + alpha * Model complexity

Final error function = how wrong the predictions are + a penalty for how complex the model is

The larger alpha is, the simpler the model; the smaller, the higher the risk of overfitting, so it needs proper tuning

  • L1 regularization:

    • Uses the sum of absolute weights as the penalty
    • Tends to drive some weights to 0
    • Good for removing unnecessary features
  • L2 regularization:

    • Uses the sum of squared weights as the penalty
    • Strongly suppresses large weights
    • Keeps all weights small and stable

L1 → |w| → Lasso → some weights 0 → feature selection

L2 → w² → Ridge → suppresses large weights → smooth model

2. Dropout

  1. During training, randomly select nodes and remove the weight connections before and after the selected nodes
  2. Select new nodes to drop out for each mini-batch or training cycle (epoch)
  3. No dropout during inference
  • If 1000 data points are split into 10
    • Mini-batch: a bundle of 100 data points,
    • epoch: going through all 10 mini-batches is 1 epoch,
    • iteration: how many repeats until one epoch is done: 10

3. Batch normalization

  1. Cause of the problem Internal covariate shift: as earlier layers train, their weights change, so the data passed to the current layer changes. This slows down training.

  2. Concept First standardize the distribution that got shaken around by earlier layers, then relearn the mean and scale this layer actually needs with γ and β

Batch normalization first sets the intermediate values x_i in a mini-batch to mean 0, standard deviation 1,
then uses γ and β to readjust them to the scale and position the model wants, producing y_i

x_i = value before normalization
x̂_i = value set to mean 0, standard deviation 1
y_i = final value after applying γ and β

γ = scale
β = shift

5. Weight training methods

  1. Gradient descent
  2. Gradient descent with momentum
  3. NAG(Nesterov accelerated gradient)
  4. Adagrad (Adaptive Gradient Algorithm)
  5. Adadelta
  6. RMSprop
  7. ADAM (Adaptive Moment Estimation)