Why the weights matter
The residual and the boundary fight each other.
1 min read
With $\lambda_r = \lambda_b = 1$ the residual term usually dominates by two orders of magnitude, and the network learns a smooth function that ignores the boundary entirely.
$$\mathcal{L} = \lambda_r\mathcal{L}_r + \lambda_b\mathcal{L}_b$$
Fixing it
Gradient-norm balancing rescales each $\lambda$ so that the terms contribute comparable gradient magnitudes.