The residual as a loss
The idea in one equation.
A physics-informed neural network does not learn a mapping from data. It learns a function that satisfies a differential equation.
Take the one-dimensional heat equation on $x \in [0, 1]$, $t \in [0, T]$:
$$\frac{\partial u}{\partial t} = \alpha \frac{\partial^2 u}{\partial x^2}$$
A network $u_\theta(x, t)$ is trained so the residual
$$r_\theta(x,t) = \partial_t u_\theta - \alpha \, \partial_{xx} u_\theta$$
is close to zero at a set of collocation points. The derivatives come from automatic differentiation, not from a finite-difference stencil, which is why the method needs no mesh.
Why this is different
A supervised network is only as good as the data you show it. A physics-informed network has a second source of truth: the equation itself. Where you have no measurement, the residual still tells the optimiser what a correct answer would look like.
The loss
The total loss is a weighted sum of three terms:
$$\mathcal{L}(\theta) = \lambda_r \mathcal{L}_r + \lambda_b \mathcal{L}_b + \lambda_d \mathcal{L}_d$$
- $\mathcal{L}_r$ — the mean squared residual at interior points
- $\mathcal{L}_b$ — mismatch on the boundary and at $t = 0$
- $\mathcal{L}_d$ — mismatch against measured data, when you have any
Choosing the weights $\lambda$ is most of the practical difficulty, and it is what Chapter 4 is about.