Scientific ML Studio
Learn/ Physics-Informed Neural…/ 2.1
2.1 · Neural networks from scratch

The neuron

A weighted sum, a bias, and a bend. Everything else in a neural network is this, repeated.

A neural network looks intimidating in diagrams and is, in fact, very small in code. The whole thing is built from one repeating unit, the neuron, and this section builds it.

The computation

A neuron takes several numbers in, $x_1, \dots, x_n$, and produces one number out. It does two things.

  1. A weighted sum plus a bias. Each input gets a weight $w_j$, the products are added, and a constant $b$ (the bias) is added on top:

$$z = w_1 x_1 + w_2 x_2 + \dots + w_n x_n + b = \mathbf{w}^\top \mathbf{x} + b.$$

  1. An activation. The sum is passed through a fixed non-linear function $\sigma$:

$$a = \sigma(z).$$

The weights and the bias are the neuron's parameters: they are what training changes. The activation function $\sigma$ is chosen by the designer and not trained (Section 2.3 compares the usual choices).

Diagram of a single neuron: three inputs multiplied by weights, summed with a bias to give z, then passed through an activation to give the output a.
Figure 2.1. One neuron. Inputs are scaled by weights, summed with a bias, then bent by the activation function.

A concrete case, with $\sigma = \tanh$:

import numpy as np

x = np.array([0.5, -1.0, 2.0])      # three inputs
w = np.array([0.4, 0.3, -0.2])      # one weight per input
b = 0.1                              # bias

z = w @ x + b                        # 0.2 - 0.3 - 0.4 + 0.1
a = np.tanh(z)
print("z =", round(z, 4))
print("a =", round(a, 4))
z = -0.4
a = -0.3799

That is all a neuron is. The interesting part is what happens when you change the numbers.

What the parameters do

Let us look at a neuron with a single input, $a(x) = \tanh(wx + b)$, because then we can plot it.

  • The activation $\tanh$ is a smooth step: close to $-1$ for very negative arguments, close to $+1$ for very positive ones, with a transition near zero.
  • The bias moves the step. The transition happens where $wx + b = 0$, that is at $x = -b/w$.
  • The weight controls how sharp the step is. Larger $|w|$ means a steeper transition; a negative $w$ flips it.
for w, b in [(1, 0), (5, 0), (5, -2.5), (-5, 2.5)]:
    centre = -b / w
    xs = np.array([centre - 0.5, centre, centre + 0.5])
    print(f"w={w:+d}, b={b:+.1f}: step at x={centre:+.2f},  a at (centre-0.5, centre, centre+0.5) =",
          np.round(np.tanh(w * xs + b), 3))
w=+1, b=+0.0: step at x=+0.00,  a at (centre-0.5, centre, centre+0.5) = [-0.462  0.     0.462]
w=+5, b=+0.0: step at x=+0.00,  a at (centre-0.5, centre, centre+0.5) = [-0.987  0.     0.987]
w=+5, b=-2.5: step at x=+0.50,  a at (centre-0.5, centre, centre+0.5) = [-0.987  0.     0.987]
w=-5, b=+2.5: step at x=+0.50,  a at (centre-0.5, centre, centre+0.5) = [ 0.987  0.    -0.987]
Four tanh curves of one input: a gentle step centred at zero, a steep step centred at zero, a steep step shifted to x equals 0.5, and a steep step flipped and shifted.
Figure 2.2. $\tanh(wx + b)$ for four choices of $(w, b)$. The weight sets the steepness (and sign), the bias slides the step along the axis.

A single such step is a poor model of anything interesting. But steps can be added up, each with its own position and steepness, and a sum of enough steps can approximate a surprisingly wide range of shapes. That is the next two sections, and Section 2.4 makes the claim precise.

Why the bend matters

If the activation were absent ($a = z$), the neuron would be a linear function, and a combination of linear functions is linear. No amount of stacking would ever produce a curve. The non-linearity $\sigma$ is what lets a network represent anything but a straight line; it is the reason it is there.

Where this goes in a PINN

A PINN's network takes the coordinates as inputs, say $(x, t)$, so $n = 2$, and its output is the solution value $u$. Because the equation involves derivatives of $u$, we will differentiate through $\sigma$. That makes the smoothness of the activation a real design constraint (Section 2.3) and not a detail.

Exercises

  1. Compute $a$ for a neuron with $\sigma = \tanh$, $\mathbf{w} = (2, -1)$, $b = 0.5$, input $\mathbf{x} = (0.25, 1)$.
  2. For $a = \tanh(3x - 1.5)$, where is the step centred, and is $a$ increasing or decreasing in $x$?
  3. How many parameters does a neuron with 10 inputs have?
Answers
  1. $z = 2(0.25) - 1(1) + 0.5 = 0$, so $a = \tanh(0) = 0$.
  2. At $x = 1.5/3 = 0.5$, and increasing, because $w = 3 > 0$.
  3. Eleven: ten weights and one bias.

Recap

  • A neuron computes $a = \sigma(\mathbf{w}^\top\mathbf{x} + b)$: a weighted sum, then a fixed non-linear bend.
  • The weights and bias are the learnable parameters; in one dimension they set the steepness and position of a smooth step.
  • Without the non-linearity, stacked neurons would collapse into a single linear function.