Scientific ML Studio
Learn/ Physics-Informed Neural…/ 2.3
2.3 · Neural networks from scratch

Activation functions

The bend in every neuron is a choice. For PINNs the choice has a consequence the usual advice misses: we will differentiate it, twice.

The activation function $\sigma$ is the only non-linear ingredient in a network, so it decides what the network can express and how well it can be trained. Four are in common use.

Name Formula $\sigma(z)$ Derivative $\sigma'(z)$ Range
Sigmoid $\dfrac{1}{1+e^{-z}}$ $\sigma(z)\,\bigl(1-\sigma(z)\bigr)$ $(0, 1)$
Tanh $\tanh z$ $1-\tanh^2 z$ $(-1, 1)$
ReLU $\max(0, z)$ $0$ for $z<0$, $1$ for $z>0$ $[0, \infty)$
SiLU (swish) $z\,\sigma_\text{sig}(z)$ $\sigma_\text{sig}(z)\,\bigl(1+z(1-\sigma_\text{sig}(z))\bigr)$ $\approx[-0.28, \infty)$

(For SiLU, $\sigma_\text{sig}$ is the sigmoid.) The derivatives are worth knowing by heart for the first two: they are expressed with the function's own value, so once a network has computed $\tanh z$ in the forward pass, the slope costs almost nothing extra. That economy is what makes backpropagation cheap (Chapter 3).

Top row: sigmoid, tanh, ReLU and SiLU curves. Bottom row: their derivatives, with ReLU's derivative a jump from 0 to 1 at the origin.
Figure 2.4. Four activations (top) and their derivatives (bottom). Note the jump in ReLU's derivative.

Checking the derivatives yourself

Never trust a derivative formula you have not tested. A central finite difference, $\sigma'(z)\approx\bigl(\sigma(z+h)-\sigma(z-h)\bigr)/2h$, is accurate to about $h^2$ and takes three lines:

import numpy as np

z = np.linspace(-3, 3, 7)
h = 1e-5

sigmoid = lambda z: 1 / (1 + np.exp(-z))
checks = {
    "sigmoid": (sigmoid,                lambda z: sigmoid(z) * (1 - sigmoid(z))),
    "tanh":    (np.tanh,                lambda z: 1 - np.tanh(z) ** 2),
    "silu":    (lambda z: z * sigmoid(z), lambda z: sigmoid(z) * (1 + z * (1 - sigmoid(z)))),
}
for name, (f, df) in checks.items():
    numeric = (f(z + h) - f(z - h)) / (2 * h)
    print(f"{name:8s} max |formula - finite difference| = {np.abs(df(z) - numeric).max():.1e}")
sigmoid  max |formula - finite difference| = 6.7e-12
tanh     max |formula - finite difference| = 3.3e-11
silu     max |formula - finite difference| = 2.0e-11

All three agree to about ten digits. ReLU is left out on purpose: at $z = 0$ its derivative does not exist, which is the heart of the matter.

The consequence for PINNs

For an ordinary classifier, ReLU is the default activation because it is cheap and trains fast. A PINN is different. The residual of a typical equation contains second derivatives of the network with respect to its inputs, for example $u_{xx}$ in the heat equation $u_t = \alpha u_{xx}$. So we need the network to be at least twice differentiable, and we need those derivatives to carry information.

Look at what ReLU does. A ReLU network is piecewise linear: between the kinks it is a straight line, whose second derivative is exactly zero, and at a kink the derivative jumps. So $u_{xx}$ of a ReLU network is zero almost everywhere and undefined at the kinks. The equation's residual then carries no information about curvature, and training cannot use it. Chapter 5 shows this with a real network and automatic differentiation. For now, remember the rule:

Rule of thumb for PINNs

Use a smooth activation (tanh is the classic choice; SiLU and sine are also used). Avoid ReLU whenever the loss contains second or higher derivatives of the network.

Saturation

Smoothness is not the only consideration. Look at tanh's derivative in Figure 2.4: it is close to 1 near zero and decays towards zero for $|z|$ beyond about 3. A neuron with a large pre-activation is saturated: its output barely changes when its input changes, so the training signal passing back through it is tiny. In a deep network, many small factors multiplied together can make the signal vanish before it reaches the early layers (the vanishing gradient problem). Keeping the initial weights at a sensible scale, as in Section 2.2, keeps neurons out of saturation at the start, and the same concern reappears in Part 4 when we talk about input scaling.

Exercises

  1. Compute $\tanh'(0)$ and $\tanh'(2)$ with the formula $1 - \tanh^2 z$. (Use $\tanh 2 \approx 0.964$.)
  2. Which of sigmoid, tanh, ReLU, SiLU has a second derivative that is zero almost everywhere?
  3. Why does a PINN for the Poisson equation $u''(x) = f(x)$ rule out a plain ReLU network?
Answers
  1. $\tanh'(0) = 1$. $\tanh'(2) = 1 - 0.964^2 \approx 0.071$. It is far smaller, which is saturation.
  2. ReLU. It is piecewise linear.
  3. The residual is $u_\theta'' - f$. For a ReLU network $u_\theta''$ is zero almost everywhere, so the residual is just $-f$ and does not depend on the network's parameters in any useful way. There is nothing to minimise.

Recap

  • The activation supplies the non-linearity. Sigmoid and tanh are smooth and saturate; ReLU is cheap but piecewise linear.
  • Derivatives of tanh and sigmoid are written in terms of the function itself, which keeps training cheap.
  • A PINN differentiates the network twice or more, so it needs a smooth activation; tanh is the standard choice.