Raster Modeller

Build continuous-value raster predictions, train reusable machine-learning and deep-learning models, tune hyperparameters, save trained models as project assets, and apply an existing model to new predictor data without retraining.

Regression & Raster PredictionReusable Trained ModelsCNN + U-NetProcessing & Model Reports

1. What Raster Modeller Does

Raster Modeller learns the relationship between raster predictor variables and a numeric target, then produces a continuous prediction surface. The current workflow can stop after a one-off raster prediction, save the trained model for later reuse, or train a model without producing a raster immediately.

Core idea: the training samples provide examples of the target variable; the predictor raster supplies the explanatory variables. The model learns target = f(predictors), then evaluates that learned function across valid raster pixels.
ŷ = f(x1, x2, …, xp; θ)ŷ is the predicted target value, x are predictor bands/features, and θ represents the fitted model parameters.

Typical applications include biomass, canopy metrics, environmental variables, soil or water parameters, continuous risk surfaces, and other numeric targets that have representative reference samples.

2. Workflow Modes

Choose the workflow from the top of Raster Modeller. The form changes so that only the inputs and outputs relevant to the selected workflow are shown.

Process Once

Fit the selected method from the current training samples and immediately create a prediction raster. The fitted model is not retained as a reusable project model.

Output: Raster
Process Once + Save Model

Train the model, select a tuned configuration, create the raster prediction, and save the same fitted model as a reusable model asset.

Output: RasterOutput: Model
In the current UI, Manual is disabled for this workflow; Hyperparameter Tuning is used to select the saved-model configuration.
Train & Save Model

Train and validate a reusable model without generating a raster prediction in the same run. Use this when the main goal is to build a model asset for later inference.

Output: Model only
Use Trained Model

Select a previously saved Raster Modeller model and apply it to a compatible predictor raster. The saved model is loaded directly; it is not trained again.

Output: New Raster
Training workflow Predictor Raster + Training Samples + Numeric Target ↓ Choose Algorithm + Validation + Parameters ↓ Optional Hyperparameter Tuning ↓ Fit Final Model ↓ ┌───────────────────┬────────────────────────┐ │ Create Raster │ Save Reusable Model │ │ (when requested) │ (when requested) │ └───────────────────┴────────────────────────┘ Reuse workflow Saved Model + Compatible Predictor Raster ↓ Load Model → Validate Compatibility → Predict Raster

3. Training, Sampling & Validation

Training samples

Training samples may be point or area-based reference features. Each sample must provide a numeric target field. Raster Modeller extracts the selected predictor bands/features at the sample locations and builds a learning table.

Area samples: polygons may contribute many raster samples. Validation should be separated at the feature level where possible so pixels from the same reference feature do not leak into both training and validation sets.

Validation methods

MethodHow it worksWhen to use
Auto SplitAutomatically divides usable samples into training and validation subsets using the configured training percentage and random state.Good default when one representative reference dataset is available.
Manual Validation LayerUses a separate validation dataset and target field that were not used to fit the model.Preferred when an independent reference dataset is available.

Why validation matters

Training accuracy describes fit to samples the model has already seen. Validation accuracy estimates generalization to unseen samples. A model can have excellent training performance and still generalize poorly.

Generalization gap = R²training − R²validationA large positive gap may indicate overfitting. Interpret the size in the context of sample quality, target variability and the application.

4. Algorithms & Theory

Algorithm availability depends on workflow. Simple linear methods are intended primarily for one-off statistical regression; in workflows that create reusable trained-model assets, the current UI disables Single Linear Regression and Multiple Linear Regression.

Single Linear Regression (SLR)

Models one predictor band with a straight-line relationship.

ŷ = β0 + β1x

Use when one predictor has a clearly interpretable, approximately linear relationship with the target.

Multiple Linear Regression (MLR)

Combines several predictor bands using an additive linear model.

ŷ = β0 + Σβjxj

Useful for interpretable multiband relationships. Check collinearity and VIF diagnostics.

Random Forest

Ensembles many decision trees trained from randomized samples and predictor subsets.

ŷ(x)= (1/T)Σ ft(x)

Strong general-purpose nonlinear baseline; handles interactions and usually requires limited feature scaling.

SVR (Support Vector Regression)

Fits a regularized regression function with an ε-insensitive error region; kernels can model nonlinear relationships.

min ½||w||² + CΣ(ξi+ξi*)

Effective for moderate sample sizes; scaling and C/ε/kernel settings matter.

MLP (Multi-Layer Perceptron)

A feed-forward neural network that learns nonlinear transformations of predictor values.

h(l)=φ(W(l)h(l−1)+b(l))

Useful when relationships are nonlinear and the training set is sufficiently representative.

KAN (Kolmogorov–Arnold Network)

Uses learnable univariate functions on network edges, inspired by the Kolmogorov–Arnold representation theorem.

f(x) ≈ Σ Φq(Σ φqp(xp))

A flexible alternative for nonlinear regression and scientific experimentation.

Transformer

Uses attention to learn cross-band or feature interactions instead of treating every predictor independently.

Attention(Q,K,V)=softmax(QKT/√dk)V

Best suited to complex feature interactions when training data and computation are adequate.

CNN (Convolutional Neural Network)

Uses convolution kernels over local raster neighborhoods. Unlike a per-pixel tabular model, a CNN can learn texture, edges and spatial context around the prediction location.

H(l) = φ(W(l) * H(l−1) + b(l))* denotes convolution over a spatial neighborhood.

Useful when nearby pixels contain information that a single-pixel spectrum cannot capture. It typically needs more samples and computation than RF/SVR.

U-Net

A convolutional encoder–decoder with skip connections. The encoder learns multiscale context; the decoder reconstructs a dense per-pixel prediction while skip connections preserve fine spatial detail.

Decoderl = Up(Decoderl+1) ⊕ Encoderl⊕ denotes feature fusion through the U-Net skip connection.

Useful for spatially structured continuous surfaces where both broad context and local boundaries matter. It is more compute- and data-intensive.

Algorithm choice is not a ranking. The best method depends on target physics, spatial scale, sample size, predictor quality and validation performance. Prefer the simplest model that generalizes adequately unless the application benefits from spatial/deep-learning context.

5. Hyperparameter Tuning

Hyperparameters control model structure or learning behavior but are not directly fitted like ordinary model coefficients. Examples include tree depth, number of trees, regularization, learning rate, network depth, batch size, convolution settings or attention dimensions.

Manual

You specify the algorithm parameters directly. This is useful for controlled experiments, reproducing a known configuration, or quick one-off processing.

Hyperparameter Tuning

The system evaluates multiple candidate configurations, compares them using the selected objective/validation design, and keeps the best configuration for the final fit.

Search space → Candidate 1 → validation score Candidate 2 → validation score Candidate 3 → validation score ↓ Best hyperparameters ↓ Fit final model ↓ Save model / create raster as requested
Important: the model saved after tuning is the final model fitted with the selected best hyperparameters. Tuning results should be interpreted together with validation metrics—not by the winning parameter values alone.

6. Inputs, Parameters & Outputs

Input / ControlPurpose
Predictor RasterRaster bands/features used to explain the numeric target. A trained model can only be reused with compatible predictor structure.
Training SamplesPoint or area reference features used for training.
Target FieldNumeric attribute representing the variable to predict.
Input Bands / PredictorsSelects which raster bands/features are used by the model.
Validation MethodAuto Split or an independent Manual Validation Layer.
Algorithm MethodSLR, MLR, Random Forest, SVR, MLP, KAN, Transformer, CNN or U-Net, depending on workflow.
Hyperparameter ModeManual configuration or Hyperparameter Tuning. Process Once + Save Model uses tuning in the current workflow.
Model Output NameName of the reusable model asset when a model is saved.
Raster Output NameName of the generated prediction raster when the workflow produces raster output.

Output matrix

WorkflowRaster OutputReusable ModelRetraining?
Process OnceYesNoYes, for this run
Process Once + Save ModelYesYesYes, then the fitted model is saved
Train & Save ModelNoYesYes
Use Trained ModelYesExisting model reusedNo

7. How to Read the Reports

Raster Modeller can expose two complementary reports. They answer different questions and should not be interpreted as duplicates.

Processing Report

Question answered: “What happened during this analysis job?”

  • Inputs and selected parameters
  • Method and validation configuration
  • Run-level model diagnostics
  • ROI & sampling details
  • Raster/model outputs
  • Execution stages and duration
  • CPU/RAM resource usage
  • Context, warnings and technical metadata

Model Report

Question answered: “What is this saved model, how well was it trained, and can I reuse it?”

  • Algorithm / model family
  • Target and predictor bands/features
  • Training & validation metrics
  • Final hyperparameters
  • Hyperparameter tuning result
  • Artifact/contract information
  • Compatibility and reuse information
  • Link back to the originating processing report

Key model metrics

Training R²How much target variance is explained on the fitting data. High training R² alone is not proof of good generalization.
Validation R²Variance explained on held-out or independent samples. Compare this with Training R².
RMSERoot Mean Squared Error. Same units as the target and more sensitive to large errors. Lower is better for the same target/dataset.
MAEMean Absolute Error. Typical absolute deviation in target units. More robust to large outliers than RMSE.
BiasMean signed error (prediction − observed). Values near zero indicate little systematic over- or under-prediction.
NRMSENormalized RMSE. Helps compare error relative to the target range/scale when normalization is defined consistently.
Generalization GapDifference between training and validation performance. A large gap can indicate overfitting or dataset shift.
Adjusted R²For linear models, penalizes adding predictors that do not provide enough explanatory value.
OOB Metrics (Random Forest)Out-of-bag evaluation from training samples not used by individual trees. Useful as an additional internal diagnostic.
Predictor ImportanceShows which predictors influence predictive performance. Importance is not proof of causation.

Recommended reading order

  1. Confirm inputs. Make sure predictor raster, bands, training layer, target field and validation method are the intended ones.
  2. Check sample counts. Very small or unbalanced training/validation sets make metrics unstable.
  3. Compare training vs validation. Look for good validation performance without an excessive generalization gap.
  4. Read RMSE/MAE in target units. Decide whether the error magnitude is acceptable for the application—not against a universal threshold.
  5. Check bias and residuals. Residuals should not show strong systematic structure; bias should be small relative to target scale.
  6. Review predictor diagnostics. Look for unstable dependence on one predictor, collinearity in MLR, or plausible feature importance patterns.
  7. Inspect execution/resources. CNN/U-Net/Transformer/KAN may need more CPU/GPU/RAM/time than RF/SVR/linear methods.
  8. Validate spatially. Compare the prediction raster with independent reference information and check for artifacts, edge effects, NoData problems or extrapolation.
No universal “good” R²/RMSE threshold exists. Acceptance depends on the target variable, measurement uncertainty, spatial scale, sample design and intended decision. Compare models on the same validation design.

8. Reusing a Trained Model

A saved Raster Modeller model is a project model asset that contains the fitted estimator plus the information required to interpret its expected predictors. Reuse avoids retraining when the model is already approved and the new raster is compatible.

  1. Open Raster Modeller.
  2. Set Workflow to Use Trained Model.
  3. Select the saved model from Trained Model. Newest models are listed first and can be searched.
  4. Select the new Predictor Raster.
  5. Verify that predictor bands/features, preprocessing, scale and units are compatible with the original training data.
  6. Enter an output raster name and run prediction.
The model is not trained again. The stored artifact is loaded, its compatibility is checked, and inference is applied to the new raster.
Model transfer warning: a technically compatible raster can still be scientifically out-of-domain. Sensor differences, season, atmospheric processing, band order, resolution or target population changes may reduce accuracy.

9. Best Practices

  • Use analysis-ready predictor rasters with consistent units, masks and preprocessing.
  • Design training samples to cover the full target range and spatial variability; avoid clustering all samples in one easy area.
  • Prefer independent validation data when available.
  • Do not choose an algorithm from training R² alone; compare validation error and residual behavior.
  • Use Random Forest/SVR as strong conventional baselines before concluding that a deeper model is necessary.
  • Use CNN/U-Net when spatial neighborhood and multiscale context are scientifically relevant and sufficient training data exist.
  • When tuning, record the search space and validation design; a best trial is only meaningful within that search space.
  • Before reusing a model, verify predictor identity, band order, preprocessing, resolution and valid numeric ranges.
  • Review both the Processing Report and Model Report when a reusable model will be used operationally.