Scientific ML Studio
Learn/ Studio Tour/ 4.2
4.2 · Getting the code out

Run the script

A generated script needs Python, PyTorch and, for the figures, Matplotlib. It runs from the command line or as a pasted Colab cell, and it writes its results next to itself.

The Lab does not train anything. It writes a script, and the script trains. That is deliberate: it means you own the result, you can read every line, and you can change it. This page is about running it.

What you need

  • Python 3.10 or newer. The script uses type hints such as list[str] | None, which older versions do not understand.
  • PyTorch. The network and the automatic differentiation come from it.
  • Matplotlib, only if you want the figures. Without it the script still trains, saves the model and writes the table; it just skips the plots.
pip install torch matplotlib

PyTorch has several variants (CPU only, various GPU builds). If that plain command gives you something that does not suit your machine, the selector on pytorch.org gives the exact command.

Run it on your computer

Put the downloaded file in an empty folder, open a terminal in that folder and run it. Suppose the file is called poisson-1d.py:

python poisson-1d.py

A progress line is printed at the first epoch and then every 200:

epoch       1  total=4.823119e+01  poisson residual=4.820e+01  BC1 dirichlet=1.380e-02  BC2 dirichlet=1.380e-02
epoch     200  total=1.075312e-03  poisson residual=1.063e-03  BC1 dirichlet=1.010e-05  BC2 dirichlet=2.262e-06
epoch     400  total=1.837956e-04  poisson residual=1.831e-04  BC1 dirichlet=5.140e-07  BC2 dirichlet=1.950e-07

Each line gives the epoch, the total loss (the weighted sum from the Loss block), and one number for each term, with the labels that came from your blocks: poisson residual, BC1 dirichlet, and so on. These are the unweighted values. The lines above are real output from the 1D Poisson example of the textbook.

A healthy run shows the total falling. The first line is large because the untrained network does not satisfy the equation. The default network for a 1D problem, run for 5000 epochs, took about a minute on a plain CPU in our tests; a 2D problem takes longer. Your machine will differ. A GPU is not needed for the problems in this tour.

Options

The script reads a few command-line options. All are optional; the defaults are the numbers you set in the blocks.

Option What it does
--epochs N Run N epochs instead of the number in the Optimiser block. Good for a quick test: --epochs 300
--device cpu or --device cuda Choose where it runs. Without it, a GPU is used if PyTorch can see one
--checkpoint PATH Where to save the trained weights, instead of the file named in the Train block
--log-every N How often to print a progress line
--quiet Print nothing
--no-plots Skip the figures. The checkpoint and the table are still written

Test before you commit to a long run

Run python poisson-1d.py --epochs 200 first. If it starts and the first lines look sensible, the setup works; then run the full number. A mistake in a formula shows up in seconds this way instead of after an hour.

Run it in Google Colab

Colab has PyTorch and Matplotlib already installed.

  1. Open a new notebook at colab.research.google.com.
  2. Press Copy code in the Generate code window.
  3. Paste it into one cell and run the cell.

The script works unchanged, because it checks at start-up whether it is inside Colab and adjusts. The same file runs both places. The files are written to the session's folder (open the folder icon on the left to see them), and that folder is cleared when the session ends, so download what you want to keep. Tick Open windows on the Visualisation block if you also want each figure shown in the cell's output.

If you set Mount Google Drive on the Train block, the script asks permission to use your Drive when it starts, and then writes the checkpoint, the figures and the table into a folder called omega-lab in My Drive. That keeps them after the session closes.

For a GPU, choose Runtime → Change runtime type and select a GPU before you run.

What comes out

Everything is written to the folder you ran the script in.

File What it is
model.pt The trained weights (and a little information about the problem)
model.history.json The loss history: the total and every term at every epoch
figures/field.png The trained field. For a time-dependent problem there is one file per time slice, like field_t0.5.png
figures/loss.png The convergence plot: total loss and each term, against epoch
figures/residual.png Only if you ticked Residual map
results.csv The table of the trained field at the points you chose

The names model.pt, figures and results.csv are the defaults from the Train, Visualisation and Report blocks. If you changed those, your files carry your names.

Using the trained model again

Because the script is an ordinary Python module that does nothing when imported, you can load the saved weights in another script or a notebook without training again. Keep the downloaded file and model.pt in the same folder. (Python module names cannot contain a hyphen, so rename the file first: poisson_1d.py.)

import torch
import poisson_1d as p

checkpoint = torch.load("model.pt")      # the weights, plus some metadata
p.net.load_state_dict(checkpoint["model"])
p.net.eval()

x = torch.tensor([[0.25], [0.50], [0.75]])
with torch.no_grad():
    print(p.net(x).squeeze())

For the 1D Poisson example, with exact answer $\sin(\pi x)$, this prints values near 0.7071, 1.0000 and 0.7071. The module also holds the network's input scaling, the residual functions and the plotting code, so you can reuse any of them: p.residual_pde1(points) is the equation's residual at any points you give it.

The weights saved are the last ones

The script saves the network as it is when training stops, not the best one it passed along the way. Adam's loss can jump late in a run, and then the saved model is a little worse than the one a few hundred epochs earlier. If the loss curve ends on a spike, run again, or add the second stage (Adam, then L-BFGS) on the Optimiser block.

If something goes wrong

Symptom Likely cause
ModuleNotFoundError: No module named 'torch' PyTorch is not installed in the Python you ran. Check with python -m pip list
SyntaxError on a line with list[str] \| None Python older than 3.10
Training runs, but no figures folder Matplotlib is not installed. Install it, or ignore the plots
Loss is nan from the start A formula produced an infinity or undefined value somewhere in the domain, such as log(x) where $x \le 0$ or division by zero. Check the formulas and the domain's range
It is slow Normal for a CPU and 5000 epochs. Use --epochs for tests, fewer or narrower layers, or a GPU

Next: reading the script.