Choose the optimiser and set up the run
The Optimiser block decides how the weights move; the Train block decides how the run is carried out. Between them they set the epochs, the learning rate, the seed and what gets saved.
Two blocks, both ending at the same place: the Train block, which is the sink the whole graph is assembled into. A block with no path to Train is ignored.
The Optimiser block
It has no inputs. Its one output, Optimiser, goes to the Train block's Optimiser input. Gradient-based training is explained in chapter 3 of the textbook; here is what the fields do.
Stage 1
| Field | Default | Notes |
|---|---|---|
| Optimiser | Adam | Adam, AdamW, SGD with momentum, or RMSprop. Adam is the standard PINN choice and rarely needs changing |
| Learning rate | 0.001 | The step size. Too high and the loss oscillates or diverges; too low and it barely moves. 1e-3 is a good start for Adam |
| Epochs | 5000 | How many optimisation steps stage 1 runs. Each epoch is one step on the full point set |
| Weight decay | 0 | An L2 penalty on the weights; zero turns it off |
| Scheduler | None | How the learning rate changes: Step decay, Cosine annealing, or Reduce on plateau |
| Gradient clip | 0 | Caps the size of the gradient before each step; zero turns it off |
The scheduler adds fields only when you choose it. Step decay asks for a Step size (epochs between drops) and a Decay factor. Cosine annealing needs nothing more and lowers the rate smoothly towards zero over the run. Reduce on plateau asks for a Patience and a Decay factor.
Stage 2: Adam, then L-BFGS
Tick Second stage and the Lab hands the weights that stage 1 reached to a second optimiser, by default L-BFGS, a quasi-Newton method. Adam then L-BFGS is the standard PINN recipe: Adam gets close, and L-BFGS squeezes out the last order or two of magnitude that a first-order method struggles with. The extra fields are Iterations (default 500) and Learning rate (default 1.0, the starting step before L-BFGS's own line search takes over).
Why this matters for the end of a run
Adam on a PINN loss is noisy, and the loss spikes now and then. The generated script saves the weights as they are at the end, not the best ones seen. On the 1D Poisson example we ran the same notebook three times from different starting weights. One run ended on a spike, and its error was roughly ten times that of the other two, although its loss had been as low as theirs a few dozen epochs earlier. With Second stage ticked, the same three runs ended with errors 30 to 300 times lower, and none of them was affected by a spike.
Mind the loss plot: the generated script records the whole second stage as one extra point at the end of the history. You will not see L-BFGS's iterations on loss.png; you will see a last point that is lower than the rest.
The Train block
It takes the Loss and the Optimiser, and gives out a Model, which the Visualisation and Report blocks read.
Run
| Field | Default | Notes |
|---|---|---|
| Device | Auto | Auto uses a GPU if PyTorch can see one, otherwise the CPU. Or force CPU or CUDA |
| Random seed | 0 | Seeds PyTorch's random state at the start of training. It fixes how the points are drawn |
| Log every | 200 | How often, in epochs, a progress line is printed |
| Resample points every | 0 | 0 keeps one fixed set of interior points for the whole run. A positive number draws a fresh set every that many epochs |
| Early stopping | off | Stops stage 1 when the loss has not improved by Minimum improvement for Patience epochs |
About the seed. The seed fixes the random draws made during training, such as the collocation points. In the version of the generator this tour describes, the network's starting weights are drawn when the script is loaded, before the seed is set. Two runs of the same script can therefore start from different weights and finish with different errors. This is normal for PINNs and a good reason to run an important problem several times and report the spread, not one number.
Resampling. Fixed points are fast, but the network can fit those points and nothing between them. Resampling every few hundred epochs removes that risk at some cost in smoothness of the loss curve.
Run target
| Field | Notes |
|---|---|
| Run target | Auto-detect (recommended) works both as a local script and pasted into a Colab cell. Choose Local or Colab to fix the setup lines the script includes |
| Mount Google Drive | Only has an effect on Colab. Saves the checkpoint to your Drive so it survives the session |
Output
| Field | Default | Notes |
|---|---|---|
| Checkpoint file | model.pt |
Where the trained weights are saved |
| Save loss history | on | Writes the per-term loss history next to the checkpoint |
Your screenshot · the Optimiser block with Second stage ticked, and the Train block beside it.
Choosing numbers
For a smooth problem the defaults are a fine start: Adam at 1e-3 for 5000 epochs. If the loss is still falling at the end, raise the epochs. If it is jumping about, lower the learning rate or add the second stage. If you do not yet know which, run the defaults once and let the loss curve tell you.
Next: plots and tables.