Hyperspectral Analysis
Technical & Scientific Documentation

A complete single-page reference for the planned hyperspectral toolbox: spectral preprocessing, dimensionality reduction, endmember discovery, spectral-library classification, and sub-pixel unmixing.

Surface ReflectanceVNIR / SWIRWavelength-awareSAM / SIDEndmembersAbundance Mapping

1. Scope & Final Hyperspectral Toolbox

This module starts after radiometric and atmospheric preprocessing. Its preferred input is a calibrated reflectance cube with reliable wavelength metadata. The tools are intentionally independent: users do not have to run every tool in sequence.

Hyperspectral
├─ Spectral Smoothing
├─ Spectral Derivative
├─ Continuum Removal
├─ Minimum Noise Fraction (MNF)
├─ Endmember Extraction
├─ Spectral Library Matching
│    ├─ SAM — Spectral Angle Mapper
│    └─ SID — Spectral Information Divergence
└─ Spectral Unmixing
Key design decision: Spectral Library Matching and Spectral Unmixing are parallel analytical outcomes. Library Matching asks “which material is this pixel most similar to?”, while Unmixing asks “what fraction of each material is present in this pixel?”.
ToolMain PurposeInput BandsMain OutputOutput Dimension
Spectral SmoothingReduce spectral noise while preserving spectral shapeN wavelength bandsSmoothed spectral cubeUsually N bands
Spectral DerivativeEmphasize slope, edge, and absorption-shape changesN wavelength bandsDerivative spectral cubeN recommended; mathematically N−1/N−2 is also possible
Continuum RemovalNormalize absorption features relative to an upper continuumN wavelength bandsContinuum-removed cubeN bands
MNFSeparate signal from noise and reduce dimensionalityN bandsM MNF componentsM ≤ N, commonly M ≪ N
Endmember ExtractionFind spectrally pure/extreme signaturesN bands or selected MNF componentsK spectral signatures + locationsK rows × N wavelengths in the preferred physical library
Spectral Library MatchingWhole-pixel material identification/classificationN spectral bands + reference spectraClass/material raster + score1 class band + optional score/rule bands
Spectral UnmixingEstimate sub-pixel material fractionsN bands + K endmembersAbundance rasterK abundance bands (+ residual/RMSE optional)

2. Scientific Foundation

2.1 Hyperspectral pixel as a vector

A hyperspectral pixel is not a single brightness value. It is a vector containing measurements across many narrow spectral bands. For a pixel with N bands:

x = [x1, x2, …, xN]Txᵢ is the radiometric quantity measured at wavelength λᵢ; for this toolbox, surface reflectance is preferred.

If the image contains 200 bands, each pixel is a point in a 200-dimensional spectral space. A material such as vegetation, water, soil, kaolinite, calcite, or iron oxide can have a characteristic spectral shape or “spectral fingerprint”.

2.2 Reflectance domain

Reflectance is dimensionless. Physically, a reflectance of 0.20 corresponds to 20% reflectance under the measurement definition. In real corrected remote-sensing products, valid values can occasionally fall slightly below 0 or above 1 because of noise, correction artifacts, adjacency effects, scaling, or retrieval uncertainty. Therefore the analysis engine should not automatically clip all values to [0,1] unless the selected method specifically requires it.

2.3 Wavelength alignment is fundamental

For spectral matching or endmember-based analysis, the numerical band index is not enough. The engine should know each band's center wavelength and preferably its spectral response or FWHM. Two spectra can only be compared meaningfully when they represent the same wavelength space. Reference libraries therefore normally require spectral resampling to the image wavelengths.

Important: “Band 40” from one sensor is not automatically equivalent to “Band 40” from another sensor. The correct relationship is established by wavelength and spectral response, not by band number.

3. Recommended Workflow

The seven tools form a toolbox, not a mandatory linear pipeline.

Radiometric / Atmospheric Preprocessing ↓ Surface Reflectance Cube │ ├── Spectral Smoothing ───────────────┐ │ │ ├── Spectral Derivative │ │ ├──→ Spectral Library Matching ├── Continuum Removal │ ├─ SAM │ │ └─ SID ├── MNF ──→ Endmember Extraction ─────┤ ↓ │ │ │ Material/Class Map │ │ │ │ └─────────────────┴──→ Spectral Unmixing │ ↓ │ Abundance Maps │ └── Direct ML / Deep Learning / other analysis

3.1 Two important final branches

Whole-pixel spectral classification

SAM/SID compare each image spectrum to a reference spectrum and can assign one dominant material/class to each pixel.

Pixel spectrum
     +
Reference library
     ↓
SAM / SID
     ↓
1-band class map

Sub-pixel decomposition

Spectral Unmixing estimates the fractional contribution of multiple endmembers to the same pixel.

Pixel spectrum
     +
K endmembers
     ↓
Unmixing
     ↓
K abundance bands
A spectral-unmixing result should normally not be passed to SAM/SID spectral-library matching, because its bands are no longer wavelengths. They represent endmember abundances.

4. Recommended Input Data Contract

4.1 Required or strongly recommended metadata

MetadataRequirementReason
Band countRequiredDefines spectral vector dimension.
Center wavelength λStrongly recommendedNeeded for derivatives, continuum analysis, spectral resampling and library matching.
Wavelength unitsRequired for matchingnm and μm must not be mixed silently.
FWHM / spectral responseRecommendedImproves library-to-sensor resampling.
Scale / offsetRequired if encodedConverts stored integer values to physical reflectance.
NoData / bad-band maskRecommendedPrevents atmospheric absorption/noisy bands from corrupting analysis.
Product levelRecommendedAllows the system to distinguish DN/radiance/TOA/surface reflectance.

4.2 Preferred data state

Preferred:
Surface Reflectance
+ wavelength metadata
+ bad-band mask
+ valid scale/offset already applied or declared

Accepted with warning:
Radiance / TOA Reflectance
Unknown wavelength metadata
Generic band-index-only raster

Band-index-only mode can still support some mathematical processing, but physically meaningful spectral matching should be marked as limited or rejected when wavelength alignment cannot be established.

5. Spectral Smoothing

Spectral smoothing reduces high-frequency noise along the spectral axis of each pixel. It does not blur neighboring image pixels spatially. The implementation preserves the selected band count and wavelength metadata, so a 200-band input normally remains a 200-band spectral cube.

N → NBand count preserved
4 methodsSG / MA / Gaussian / Median
Spectral axisNo spatial mixing
WindowedCube processed in blocks

5.1 Available smoothing methods

MethodCore ideaBest suited forMain parameter
Savitzky–GolayLocal polynomial least-squares fitPreserving spectral shape, peaks and absorption featuresWindow length + polynomial order
Moving AverageUniform local meanSimple low-pass suppression of random high-frequency noiseWindow length
GaussianDistance-weighted local meanSmooth low-pass filtering with softer weighting than a box averageσ + truncate
MedianLocal order-statistic medianSpike / impulse / isolated-band noiseWindow length

5.2 Moving Average

Moving Average replaces each spectral sample by the arithmetic mean of an odd spectral neighborhood.

R̃(λi) = (1 / (2m+1)) Σj=-m…m R(λi+j)The window width is 2m+1 bands. Uniform weights make this method fast, but a large window can flatten narrow absorption features.

5.3 Gaussian Smoothing

Gaussian smoothing gives the largest weight to bands near the target wavelength and progressively smaller weights to more distant bands.

wj = exp[−j²/(2σ²)] / Σk exp[−k²/(2σ²)]
R̃i = Σj wj Ri-jσ controls spectral smoothing strength. The implementation limits the effective kernel radius with the truncate parameter, approximately radius ≈ truncate × σ bands.

Compared with a moving average, Gaussian weights reduce abrupt changes at the edge of the smoothing window.

5.4 Median Filter

The median filter is nonlinear. Instead of averaging values, it selects the median value inside the spectral neighborhood. It is especially useful when a spectrum contains isolated spikes that would strongly influence a mean.

R̃(λi) = median { R(λi-m), …, R(λi), …, R(λi+m) }The window is an odd number of spectral bands. Because it is an order-statistic filter, it can reject isolated spectral spikes without averaging them into nearby bands.

5.5 Savitzky–Golay Smoothing

Savitzky–Golay smoothing fits a low-order polynomial to a moving spectral window by least squares. The polynomial value at the center of the window becomes the smoothed spectral value. It is widely used in spectroscopy because it can preserve local spectral shape better than a simple moving average.

mina₀…aₚ Σj=-m…m [Ri+j − Σk=0…pakjk]2p = polynomial order; 2m+1 = window length. The window must be larger than the polynomial order.

5.6 Parameter contract in iT Sensing

MethodParameters shownImplementation note
Savitzky–GolayWindow Length; Polynomial OrderDefault method. Odd window ≥ 3 and polynomial order < window length.
Moving AverageWindow LengthUniform spectral filter; edge handling uses nearest-value extension.
GaussianGaussian Sigma; Gaussian TruncateGaussian convolution only along the band axis.
MedianWindow LengthMedian neighborhood has shape [spectral window × 1 spatial sample].

5.7 Output and QC

Input : 200-band hyperspectral reflectance
      ↓
Spectral Smoothing
      ↓
Output: 200-band smoothed cube

Band 1   → same λ1
Band 2   → same λ2
...
Band 200 → same λ200
  • The selected smoothing window must not exceed the selected spectral-band count.
  • All four methods operate independently for each spatial pixel.
  • Wavelength tags and band descriptions should be preserved.
  • Bad atmospheric/noisy bands should preferably be excluded or handled before smoothing.
Interpretation warning: stronger smoothing is not automatically better. A wide window or large Gaussian σ can suppress narrow absorption features that may be the scientific target of hyperspectral analysis.

6. Spectral Derivative

Spectral derivatives emphasize changes in the shape of a spectrum. They can suppress broad baseline effects and reveal absorption edges, inflection points, and subtle spectral differences. The output is not reflectance; its units depend on wavelength units.

6.1 First derivative

R′(λi) ≈ [R(λi+1) − R(λi-1)] / [λi+1 − λi-1]Centered finite difference for interior bands.

6.2 Second derivative

For approximately uniform spectral spacing Δλ:

R″(λi) ≈ [Ri+1 − 2Ri + Ri-1] / (Δλ)2

6.3 Savitzky–Golay derivative

The same local polynomial used for Savitzky–Golay smoothing can be differentiated analytically, often producing a more stable derivative than applying a finite difference directly to noisy spectra.

6.4 Why smoothing often precedes derivative

Differentiation amplifies high-frequency noise. Therefore:

Surface Reflectance
      ↓
Spectral Smoothing
      ↓
1st / 2nd Derivative

6.5 Output dimension

Mathematically, naïve differencing may yield N−1 or N−2 samples. For a raster-analysis platform, the recommended contract is to preserve N output bands by using centered differences for interior wavelengths, one-sided differences or declared NoData at edges, and retaining original wavelength metadata.

Recommended platform behavior: 200-band input → 200-band derivative output, with metadata indicating derivative order and wavelength units.

7. Continuum Removal

Continuum removal normalizes a spectrum relative to an upper continuum, usually represented by a convex hull over the spectrum. It is especially useful for comparing the position, shape and depth of absorption features.

7.1 Basic equation

RCR(λ) = R(λ) / C(λ)R(λ) = observed reflectance; C(λ) = continuum curve; RCR = continuum-removed reflectance.

At wavelengths where the measured spectrum touches the continuum, RCR is close to 1. Inside an absorption feature, it is typically less than 1.

7.2 Absorption depth

D(λ) = 1 − RCR(λ)Greater D indicates a deeper feature relative to the local continuum.

7.3 Output

Input : 200-band reflectance
Output: 200-band continuum-removed cube

Original wavelength dimension is preserved.
Output values are normalized spectral response, not original surface reflectance.

7.4 Spectral subset matters

The continuum depends on the wavelength interval used to construct it. The UI should therefore allow all valid bands, a wavelength range, or selected spectral windows.

If SAM/SID is applied to a continuum-removed image, the reference spectral library should undergo the same continuum-removal operation over the same wavelength range.

8. Minimum Noise Fraction (MNF)

MNF is a noise-aware linear transform that orders transformed components by image quality / signal-to-noise information. Unlike ordinary PCA, MNF explicitly requires a model of the noise covariance. The implementation keeps signal-covariance estimation separate from the selectable noise-covariance estimator.

N → MDimensional reduction
4 estimatorsNoise covariance options
Noise-awareGeneralized eigenproblem
ComponentsNot wavelengths

8.1 Signal + noise model

x = s + nx = observed spectral vector; s = underlying signal; n = noise. Let Σx be the image covariance and Σn the noise covariance.

8.2 MNF generalized eigenproblem

A convenient formulation is to solve a generalized eigenproblem that compares total image covariance against noise covariance:

Σx vi = λi Σn viEigenvectors are ordered by λ. Components with stronger signal relative to the estimated noise are retained first.

An equivalent two-stage interpretation first whitens noise and then rotates the whitened data.

z = Σn−1/2(x − μ)
y = VT zAfter whitening, the estimated noise covariance is approximately identity. V rotates the data into ordered MNF components.

8.3 Why the output can be smaller than the input

Input: 200 spectral bands
      ↓
MNF transform
      ↓
MNF 1  — high image quality / useful structure
MNF 2  — high image quality / useful structure
...
MNF 20 — retained
MNF 21+ — increasingly noise-dominated
      ↓
Output M = 20 components

MNF components are mathematical component axes. They are not wavelengths and cannot be directly matched against an ordinary laboratory spectral library unless the same fitted transform is applied to the reference spectra.

8.4 Noise Estimation Method

The four available options change only the estimate of Σn. Signal covariance is still estimated from spatially distributed valid image samples.

8.4.1 Spatial Shift Difference — default

Adjacent pixels are assumed to have similar local signal, so their difference emphasizes noise. Horizontal and vertical valid neighbor pairs are sampled.

dp,δ = xp+δ − xp
Σ̂n = ½ Cov(d)The factor ½ is equivalent to using d/√2 when the two neighboring noise realizations are independent with the same covariance.

This is a practical general-purpose default when no dedicated noise calibration is available.

8.4.2 Homogeneous ROI Difference

The same shift-difference principle is applied only inside a user-selected polygon that is expected to be spectrally homogeneous. Restricting the estimate to a homogeneous region reduces contamination by genuine scene transitions.

dp,δ = xp+δ − xp,    p,p+δ ∈ Ωh
Σ̂n,ROI = ½ Cov { dp,δ | p ∈ Ωh }Ωh is the homogeneous-noise ROI. The ROI should contain enough valid adjacent pixels and should avoid strong material boundaries.

8.4.3 Local Residual Estimation

A local spatial predictor is used to estimate the slowly varying signal around each pixel. The residual is treated as a noise sample. The current implementation uses a local mean excluding the center pixel.

x̂p = (1 / |N(p) ∖ {p}|) Σq∈N(p), q≠p xq
rp = xp − x̂p
Σ̂n = Cov(r)N(p) is an odd spatial neighborhood such as 3×3, 5×5, etc. A larger window models smoother local signal but can mix genuine spatial variation into the residual.

8.4.4 External Noise Covariance

If sensor calibration, dark-current measurements, laboratory characterization, or another trusted process already provides a noise covariance matrix, the matrix can be supplied directly.

Σn = [ σij ]B×BB must equal the selected input-band count. The implementation checks numeric values, finite entries, matrix shape, symmetry by averaging with its transpose, and nonnegative diagonal variances.

This option bypasses image-derived noise sampling and makes the external covariance the authoritative noise model.

8.5 Noise-estimator comparison

EstimatorAdditional inputStrengthMain caution
Spatial Shift DifferenceNoneFast, general-purpose, automaticScene edges can contribute real signal differences
Homogeneous ROI DifferenceHomogeneous polygon ROIBetter isolates noise when a good homogeneous region existsResult depends on ROI quality and representativeness
Local Residual EstimationSpatial window sizeUseful in heterogeneous scenes without a dedicated ROITrue fine spatial structure can enter the residual
External Noise CovarianceB×B covariance matrixBest when a trusted sensor/noise model is availableMatrix must correspond to the selected bands and data scaling

8.6 Component selection

  • Auto: lets the analysis choose an informative component count from the fitted component statistics.
  • Manual: user specifies M explicitly.
  • Covariance Sample Pixels: bounded distributed samples used for covariance estimation; full-raster transformation remains windowed.
  • Regularization: a small diagonal stabilization is applied when the noise covariance is ill-conditioned.
MNF is only as good as its noise model. Different noise estimators can produce different component ordering. The selected estimator, sample count, ROI/window parameters, covariance regularization and fitted transform should therefore be recorded in provenance.

9. Endmember Extraction

An endmember is a spectral signature representing a relatively pure material or an extreme spectral component in the scene. Endmember Extraction produces a reusable spectral-library dataset, not a classification raster. The current toolbox implements three classical pure-pixel / geometric methods: PPI, N-FINDR and VCA.

3 methodsPPI / N-FINDR / VCA
K spectraRequested endmembers
Pure pixelsReal sampled source pixels
JSONReusable spectral library

9.1 Geometric basis

Under a linear mixture model, observed spectra approximately occupy a simplex in spectral space. Pure materials tend to occur near simplex vertices, while mixtures occur inside or along simplex edges.

x ≈ M a + e
ak ≥ 0,    Σk=1…Kak ≈ 1M contains endmember spectra as columns. Pure-pixel extraction methods search for extreme observations that approximate the vertices of this simplex.

Extraction can be performed from a sampled subset for efficiency, but the selected indices should resolve back to real source pixels so that coordinates and original spectra can be preserved.

9.2 PPI — Pixel Purity Index

PPI projects the spectral cloud onto many random directions (“skewers”). Pixels repeatedly appearing at projection minima or maxima receive high purity counts.

zjr = urT yj
PPI(j) = Σr=1…R I[j = arg min z·r or j = arg max z·r]yj is the centered/reduced spectrum of pixel j; ur is a random unit direction; R is the number of skewers.

The implementation first projects the sampled spectra to a principal signal subspace and scales those axes before random projection. A minimum spectral-angle criterion can prevent nearly duplicate PPI endmembers.

θ(x,s) = cos−1[(x·s)/(||x||₂||s||₂)]Candidates separated from an already selected endmember by less than the chosen minimum angle can be skipped.

9.3 N-FINDR — Maximum-volume simplex

N-FINDR seeks the set of K observed pixels that maximizes the volume of the simplex formed by those candidate endmembers. The sampled spectra are projected into K−1 affine signal dimensions and augmented by a constant coordinate.

A = [ỹ1, ỹ2, …, ỹK]
V(A) = |det(A)| / (K−1)!Each ỹk is an augmented candidate vertex. Larger |det(A)| means a larger spectral simplex.

For a current simplex A, replacing vertex j by candidate v changes the determinant by a column-replacement factor:

det(Aj←v) = det(A) · ejTA−1vThe implementation accepts replacements only when the simplex volume increases, iterating until no useful replacement is found or the configured maximum number of passes is reached.

N-FINDR is computationally more expensive as K grows; the current server implementation limits N-FINDR to 32 endmembers per run.

9.4 VCA — Vertex Component Analysis

VCA exploits the fact that linear mixtures live in a low-dimensional affine subspace. For K endmembers, sampled spectra are projected to K−1 signal dimensions and augmented by a constant coordinate. Vertices are then identified through successive projections in directions constrained to be independent of previously selected vertices.

ỹj = [ UT(xj−μ) ; 1 ]
fk ⟂ span{ỹj₁, …, ỹjk−1}
jk = arg maxj | fkT ỹj |U spans the K−1 dimensional signal subspace. Each selected jk is an actual sampled pixel used as an endmember.

VCA is generally fast because it avoids the repeated full simplex-volume search of N-FINDR.

9.5 Method comparison and parameters

MethodGeometric criterionParameters exposedUse case
PPI DefaultFrequency of extreme random projectionsEndmember Count; Sample Pixels; Random Skewers; Projection Components; Minimum Separation Angle; Random SeedRobust pure-pixel ranking, especially after MNF
N-FINDRMaximum simplex volumeEndmember Count; Sample Pixels; Maximum Iterations; Random SeedStrong geometric pure-pixel assumption and moderate K
VCASuccessive independent vertex projectionsEndmember Count; Sample Pixels; Random SeedFast geometric extraction for many practical scenes

9.6 Output object and project dataset

Hyperspectral Raster
      ↓
PPI / N-FINDR / VCA
      ↓
K selected source pixels
      ↓
Spectral Library JSON in object storage
      ↓
public.layers
data_type = endmember
      ↓
Reusable by SAM / SID / Spectral Unmixing

Each signature should retain its source row/column and geospatial coordinate, spectral-axis metadata, extraction method, quality/selection score and preprocessing provenance.

{
  "analysis": "endmember_extraction",
  "method": "VCA",
  "source_raster": "...",
  "wavelength_unit": "nm",
  "wavelengths": [400.0, 410.0, "..."],
  "endmembers": [
    {
      "id": 1,
      "name": "Endmember 1",
      "pixel": {"row": 812, "col": 455},
      "coordinate": {"x": "...", "y": "..."},
      "spectrum": [0.081, 0.084, "..."]
    }
  ]
}
Pure-pixel assumption: PPI, N-FINDR and VCA are most reliable when sufficiently pure observations exist in the scene. If every pixel is strongly mixed, extracted “vertices” may be extreme mixtures rather than physically pure materials.

10. Spectral Library Matching

Spectral Library Matching is the hyperspectral material-identification / whole-pixel classification tool. The user supplies or selects a reference spectral library, and each image pixel is compared against candidate reference spectra.

Image Spectrum + Reference Spectral Library ↓ Wavelength overlap check ↓ Library resampling to image spectral grid ↓ Optional identical preprocessing ↓ Matching Method ├─ SAM └─ SID ↓ Best matching material/class ↓ Class Raster + Match Score

10.1 Library preprocessing requirements

  • Reference wavelengths must overlap image wavelengths.
  • Wavelength units must be known and harmonized.
  • Library spectra should be resampled using wavelength/FWHM information when possible.
  • If image spectra are continuum-removed, derivative-transformed, or otherwise normalized, the reference library must receive the equivalent transformation.
  • Bad bands should be excluded from both image and reference spectra.

10.2 SAM — Spectral Angle Mapper

SAM treats each spectrum as a vector in N-dimensional band space and computes the angle between image spectrum x and reference spectrum s.

θ = cos−1 [ (x · s) / (||x||₂ ||s||₂) ]
x · s = Σi=1…Nxisi,    ||x||₂ = √(Σi=1…Nxi2)
  • θ ≈ 0: highly similar spectral direction.
  • Larger θ: less similar.
  • Best reference is typically the class with the smallest valid angle.
  • A maximum-angle threshold should allow pixels to remain unclassified.

Because SAM focuses on vector direction rather than magnitude, it can be relatively insensitive to overall albedo/illumination scaling when used with calibrated data; this does not remove the need for correct radiometric/atmospheric preprocessing.

Recommended SAM outputs

classification.tif
  0 = Unclassified
  1 = Material A
  2 = Material B
  ...

sam_best_angle.tif
  pixel value = minimum spectral angle

optional rule images:
  one angle band per reference material

10.3 SID — Spectral Information Divergence

SID converts nonnegative spectral vectors into probability-like distributions and measures their divergence using information theory.

Step 1 — normalize spectra

pi = xi / Σj=1…Nxj,    qi = si / Σj=1…Nsj

Step 2 — directional divergence

D(p || q) = Σi=1…N pi ln(pi/qi)
D(q || p) = Σi=1…N qi ln(qi/pi)

Step 3 — symmetric SID

SID(x,s) = D(p || q) + D(q || p)
  • SID close to 0: highly similar.
  • Larger SID: greater spectral divergence.
  • A maximum SID threshold should support an unclassified class.
Because the logarithm requires positive probability terms, implementations must define robust handling for zero, negative, NoData, or invalid reflectance values. A small ε floor may be used after a scientifically documented validity mask, rather than silently converting arbitrary negative values.

Recommended SID outputs

classification.tif
  best material/reference class

sid_best_score.tif
  minimum SID divergence

optional rule images:
  one SID score band per reference material

10.4 SAM vs SID

AspectSAMSID
Core ideaAngle between spectral vectorsInformation divergence between normalized spectral distributions
Best matchSmallest angleSmallest divergence
Output class mapYesYes
Reference libraryRequiredRequired
Zero/negative handlingVector norm must be validRequires careful positive-value handling for logarithms
ThresholdMaximum angleMaximum divergence

11. Spectral Unmixing

Spectral Unmixing estimates how much each endmember contributes to an observed pixel. The implementation is supervised abundance estimation: the endmember matrix is supplied by a project Endmember dataset or a selected public reference library. The toolbox now provides six methods: three linear and three nonlinear.

6 methods3 linear + 3 nonlinear
Pixel-wiseIndependent spectra
K bandsOne abundance band/endmember
WindowedChunk raster execution

11.1 Common notation

For B spectral bands and K endmembers:

x ∈ ℝB,    M = [m1 … mK] ∈ ℝB×K,    a ∈ ℝKx = observed pixel spectrum; M = endmember matrix resampled to the raster spectral grid; a = abundance vector.

The linear mixing model (LMM) is:

x = M a + e

Nonlinear methods extend this relationship when multiple scattering, material interaction, intimate mixing, canopy structure or other effects make a purely linear model inadequate.

x = M a + Ψ(M,a) + eΨ represents an additional nonlinear interaction term. Different nonlinear methods make different assumptions about Ψ.

11.2 UCLS — Unconstrained Least Squares

UCLS minimizes reconstruction error without abundance constraints.

â = arg mina ||x − Ma||22

When M has suitable rank, the least-squares solution can be written with the Moore–Penrose pseudoinverse:

â = M+x ≈ (MTM)−1MTx
  • Fast baseline linear solution.
  • Abundances may be negative.
  • Abundances are not required to sum to one.

11.3 NNLS — Non-Negative Least Squares

NNLS adds the abundance non-negativity constraint (ANC).

â = arg mina ||x − Ma||22
subject to   ak ≥ 0

Negative material fractions are prevented, but the abundance sum is not forced to one.

11.4 FCLS — Fully Constrained Least Squares

FCLS adds both abundance non-negativity (ANC) and abundance sum-to-one (ASC). It remains the default method for fractional abundance mapping.

â = arg mina ||x − Ma||22
subject to   ak ≥ 0,    Σk=1…Kak = 1

The feasible abundance vectors lie on the probability simplex:

ΔK = { a ∈ ℝK : a ≥ 0, 1Ta = 1 }

The implementation uses projected-gradient iterations with simplex projection.

11.5 Why nonlinear unmixing?

A linear model assumes photons interact with only one material before reaching the sensor. Real scenes can contain second-order scattering, canopy–soil interaction, intimate mineral mixtures, complex urban surfaces or other mechanisms that distort this assumption. Nonlinear unmixing introduces a structured or flexible nonlinear term rather than forcing all residual structure into e.

11.6 GBM — Generalized Bilinear Model

GBM explicitly models pairwise second-order interactions between endmembers using Hadamard (band-by-band) products.

x = M a + Σi=1…K−1 Σj=i+1…K γijaiaj(mi ⊙ mj) + e
a ∈ ΔK,    0 ≤ γij ≤ 1γij controls the strength of interaction between endmembers i and j. The implementation alternates abundance updates with bounded bilinear-interaction updates.

The number of pairwise terms is:

Q = K(K−1)/2

Because Q grows quadratically, the current server implementation limits GBM to 16 endmembers per run.

11.7 PPNM — Polynomial Post-Nonlinear Model

PPNM first forms the linear mixture and then applies a polynomial post-nonlinearity. In the implemented quadratic form:

z = M a
x = z + b(z ⊙ z) + e
a ∈ ΔK,    |b| ≤ bmaxb is estimated per pixel and controls the strength and sign of the quadratic post-nonlinearity. The UI parameter “PPNM Maximum |b|” bounds this coefficient.

When b ≈ 0, PPNM approaches the linear mixing model.

11.8 K-Hype — Kernel-Based Semiparametric Unmixing

K-Hype models a linear abundance term plus a flexible nonlinear residual function living in a reproducing kernel Hilbert space (RKHS).

xℓ = mℓTa + ψ(mℓ) + eℓ,    ψ ∈ Hmℓ is the vector of endmember responses at spectral band ℓ. ψ models nonlinear fluctuation without fixing a single explicit bilinear or polynomial mechanism.

The implementation uses an RBF kernel across the band-wise endmember vectors:

Kℓr = exp[ −||mℓ − mr||² / (2σ²) ]

After eliminating the RKHS residual analytically, abundance estimation is performed with a kernel-weighted quadratic objective:

W = λ(K + λI)−1
â = arg mina∈ΔK (x − Ma)T W (x − Ma)λ is the RKHS regularization parameter. σ is the RBF bandwidth; a value of 0 in the UI invokes the automatic median-distance rule.

K-Hype constructs a dense band×band kernel matrix; the current implementation therefore limits a run to 512 selected spectral bands.

11.9 Method comparison

MethodFamilyAbundance constraintsNonlinear assumptionTypical role
UCLSLinearNoneNoneFast baseline / diagnostic
NNLSLineara ≥ 0NoneNonnegative linear abundance
FCLS DefaultLineara ≥ 0; Σa = 1NoneDefault fractional abundance mapping
GBMNonlineara ≥ 0; Σa = 1Pairwise bilinear interactionsSecond-order scattering / material interactions
PPNMNonlineara ≥ 0; Σa = 1Quadratic post-nonlinearitySmooth post-linear spectral distortion
K-HypeNonlinear / kernela ≥ 0; Σa = 1Flexible RKHS nonlinear residualUnknown or difficult-to-parametrize nonlinear effects

11.10 Output band count

Critical rule: output dimensionality follows the number of selected endmembers/materials, not the number of hyperspectral wavelengths.
Input hyperspectral raster = 285 bands
Selected endmembers = 6
        ↓
Spectral Unmixing
        ↓
Output abundance raster = 6 bands

Band 1 = abundance Endmember 1
Band 2 = abundance Endmember 2
...
Band 6 = abundance Endmember 6

For UCLS the output values are unconstrained coefficients. For NNLS they are nonnegative coefficients. FCLS, GBM, PPNM and K-Hype produce simplex-constrained fractions under the current implementation.

11.11 Residual and reconstruction error

e = x − x̂
RMSE = √[(1/B) Σℓ=1…B(xℓ − x̂ℓ)²]

For nonlinear methods, x̂ must include the fitted nonlinear contribution. High residuals can indicate missing endmembers, poor wavelength alignment, atmospheric/bad-band contamination, spectral variability, or a mixing mechanism not represented by the selected method.

11.12 Method-selection guidance

  • Start with FCLS when the goal is interpretable material fractions and linear mixing is plausible.
  • Compare UCLS/NNLS to diagnose whether the constraints materially change the solution.
  • Use GBM when pairwise material interactions / multiple scattering are scientifically plausible.
  • Use PPNM when a post-linear nonlinear distortion is plausible.
  • Use K-Hype when the nonlinear mechanism is not well known and a flexible kernel residual is preferred.
  • Nonlinear models have more degrees of freedom and computational cost; better fit does not automatically mean better physical interpretation.
Model-selection warning: nonlinear unmixing should not be selected simply because it is more complex. Compare reconstruction error, abundance plausibility, spatial coherence and available validation data against a linear baseline.

12. Relationship Between the Seven Tools

FromCan FeedNotes
Spectral SmoothingDerivative, Continuum Removal, MNF, Endmember Extraction, Matching, Unmixing, ML/DLPreserves wavelength domain.
Spectral DerivativeMatching, ML/DL, feature analysisReference library must use equivalent derivative for matching.
Continuum RemovalMatching, feature analysis, ML/DLUse the same spectral window on reference spectra.
MNFEndmember Extraction, ML/DL, advanced unmixingComponents are not wavelengths.
Endmember ExtractionLibrary Matching, Spectral UnmixingEndmembers may first need material identification.
Library MatchingFinal class/material analysis, GIS statisticsNot a prerequisite for Unmixing.
Spectral UnmixingAbundance analysis, thresholding, ML/DL, dominant-material derivationOutput bands are abundance dimensions.

12.1 Valid workflows

Surface Reflectance → Smoothing → SAM/SID

Surface Reflectance → Smoothing → Continuum Removal → SAM/SID

Surface Reflectance → Smoothing → Derivative → Random Forest

Surface Reflectance → MNF → Random Forest

Surface Reflectance → MNF → Endmember Extraction
                    → sample original spectra
                    → Spectral Unmixing

Surface Reflectance → Endmember Extraction → SAM/SID

External Spectral Library → resample to image wavelengths → SAM/SID

12.2 Workflow to avoid by default

Spectral Unmixing → SAM/SID spectral-library matching   ✕

After unmixing, the bands represent abundance of endmembers rather than wavelength samples.

13. Recommended Platform Output Contract

ToolMain FileSidecar / AuxiliaryLayer Type
SmoothingFloat32 COG, N bands.statistics.jsonRaster
DerivativeFloat32 COG, N bands.statistics.json + derivative metadataRaster
Continuum RemovalFloat32 COG, N bands.statistics.json + spectral rangeRaster
MNFFloat32 COG, M bands.statistics.json + MNF transform/statistics metadataRaster
Endmember Extractionendmembers.jsonCSV, optional point layer, optional purity rasterEndmember dataset (public.layers data_type=endmember)
Library MatchingInteger class COGconfidence/score COG, class legend JSONRaster classification
UnmixingFloat32 COG, K bands.statistics.json + endmember legend + RMSE/residual optionalRaster abundance

13.1 Suggested endmember JSON fields

  • analysis/method/version;
  • source raster identifier;
  • spectral domain and wavelength units;
  • wavelength array;
  • endmember identifier and optional material label;
  • source pixel row/column and geospatial coordinate;
  • spectral vector;
  • purity/quality score where available;
  • preprocessing history;
  • bad-band list;
  • creation timestamp.

13.2 Suggested classification legend

{
  "classes": [
    {"value": 0, "name": "Unclassified"},
    {"value": 1, "name": "Kaolinite", "reference_id": "lib-001"},
    {"value": 2, "name": "Calcite", "reference_id": "lib-021"},
    {"value": 3, "name": "Vegetation", "reference_id": "lib-104"}
  ],
  "method": "SAM",
  "score_semantics": "lower_is_better"
}

13.3 Suggested abundance metadata

{
  "bands": [
    {"band": 1, "endmember": "Vegetation", "unit": "fraction"},
    {"band": 2, "endmember": "Soil", "unit": "fraction"},
    {"band": 3, "endmember": "Shadow", "unit": "fraction"}
  ],
  "method": "FCLS",
  "family": "linear",
  "sum_to_one": true,
  "nonnegative": true
}

14. Validation & Quality Control

14.1 Common validation

  • Reject empty rasters and invalid spectral dimensions.
  • Verify NoData behavior and finite-value proportion.
  • Verify wavelength-array length equals raster band count.
  • Check wavelengths are monotonically ordered or reorder data and metadata together.
  • Detect duplicate wavelengths.
  • Apply declared scale/offset before physics-based spectral analysis.
  • Exclude known bad bands and strong atmospheric-absorption regions when appropriate.

14.2 Smoothing QC

  • Savitzky–Golay, Moving Average and Median windows must be odd and cannot exceed the selected band count.
  • Savitzky–Golay polynomial order must be smaller than the window length.
  • Gaussian σ must be positive and truncate must define a finite practical kernel radius.
  • Confirm filtering occurs only along the spectral axis; neighboring spatial pixels must not be mixed.
  • Inspect absorption-feature preservation before choosing aggressive smoothing parameters.

14.3 Derivative QC

  • Reject zero wavelength spacing.
  • Handle irregular wavelength spacing explicitly.
  • Record derivative order and units.

14.4 MNF QC

  • Record the selected noise estimator: Spatial Shift Difference, Homogeneous ROI Difference, Local Residual, or External Covariance.
  • For homogeneous ROI estimation, verify the polygon is sufficiently large, valid and spectrally homogeneous.
  • For local residual estimation, verify the odd spatial window is appropriate for scene texture.
  • For external covariance, require a finite B×B matrix matching the selected bands and data scaling.
  • Noise covariance must be numerically valid; regularize near-singular matrices.
  • Return eigenvalues/component-quality information and never label MNF components as wavelengths.

14.5 Endmember QC

  • K must be consistent with spectral dimensionality and available pure-pixel structure.
  • PPI: inspect purity counts and minimum-angle duplicate suppression.
  • N-FINDR: inspect simplex-volume convergence; avoid excessive K and remember the server limit of 32 endmembers.
  • VCA: verify selected vertices are unique and source pixels are valid.
  • Report duplicate/highly similar extracted endmembers, source pixels and validity masks.
  • Allow user interpretation/renaming after extraction; geometric extremeness is not automatically a material identity.

14.6 Matching QC

  • Require spectral overlap and harmonized wavelength units.
  • Resample the selected reference library to the image spectral grid.
  • Use thresholds to avoid forced classification.
  • Return best score and, where useful, second-best score/margin.
  • For SID, handle nonpositive values explicitly.

14.7 Unmixing QC

  • Endmember and image spectral dimensions must match after resampling/transformation.
  • Use UCLS as a useful unconstrained baseline but do not interpret negative coefficients as physical fractions.
  • NNLS must satisfy non-negativity; FCLS, GBM, PPNM and K-Hype should satisfy both non-negativity and sum-to-one under the implemented solver.
  • GBM: K ≤ 16 in the current implementation because interaction count grows as K(K−1)/2.
  • PPNM: inspect the fitted nonlinearity magnitude and chosen maximum |b|.
  • K-Hype: inspect RKHS regularization and kernel bandwidth; selected band count is currently limited to 512.
  • Return reconstruction residual/RMSE and compare nonlinear methods against a linear baseline.

15. Implementation Notes for an Analysis Server

15.1 Processing strategy

A hyperspectral cube can be hundreds of bands and several gigabytes. Algorithms should use block/window processing where mathematically possible. Do not read a full cube into RAM unless required and size-checked.

ToolWindow/Block Friendly?Global Preparation Needed?
SmoothingYesNo, except metadata
DerivativeYesNo, except wavelength array
Continuum RemovalYes, pixel blocksNo
MNFTransform yesNoise/covariance statistics first
Endmember ExtractionMethod dependentOften sampling/reduction/global search
SAM/SIDYesReference resampling/preprocessing first
UnmixingYesEndmember matrix/solver setup first

15.2 Dtype

  • Use Float32 for transformed spectral/abundance outputs unless higher precision is specifically required.
  • Use integer type for class raster, choosing UInt8/UInt16/UInt32 based on class count.
  • Do not encode transformed physical values into integers unless scale/offset is explicit.

15.3 Raster statistics

A statistics sidecar should describe each output band correctly. For a 200-band cube this can be large; the implementation may store per-band min/max/mean/std plus sampled histograms. MNF and abundance outputs should have semantic band descriptions.

15.4 Provenance

{
  "input_domain": "surface_reflectance",
  "operations": [
    {"tool": "spectral_smoothing", "method": "savitzky_golay", "window": 7, "order": 2},
    {"tool": "mnf", "noise_estimation": "homogeneous_roi_difference", "components": 20},
    {"tool": "continuum_removal", "range_nm": [2000, 2400]}
  ]
}

Provenance is essential because a reference library must receive compatible preprocessing before matching.

15.5 Do not silently change spectral semantics

OutputBand semantics
Smoothed cubeWavelength
Derivative cubeDerivative at wavelength
Continuum-removed cubeNormalized value at wavelength
MNF rasterMNF component
SAM/SID class rasterClass ID
Unmixing rasterEndmember abundance

16. Example Analytical Workflows

16.1 Mineral identification from reflectance

Surface Reflectance
      ↓
Remove bad bands
      ↓
Spectral Smoothing
(SG / Gaussian / Moving Average / Median)
      ↓
Continuum Removal (target SWIR range)
      ↓
Resample mineral spectral library to image wavelengths
      ↓
Apply same continuum removal to library
      ↓
Spectral Library Matching — SAM or SID
      ↓
Mineral Class Map + Match Score

16.2 Endmember discovery and linear abundance mapping

Surface Reflectance (200 bands)
      ↓
MNF
Noise estimator:
Spatial / Homogeneous ROI / Local Residual / External Covariance
      ↓
retain informative components
      ↓
PPI / N-FINDR / VCA
      ↓
K pure-pixel locations
      ↓
sample original reflectance
      ↓
K endmember spectra
      ↓
FCLS (default) / NNLS / UCLS
      ↓
K abundance bands + reconstruction QC

16.3 Nonlinear abundance mapping

Surface Reflectance
      +
Focused Endmember Library
      ↓
Start with FCLS baseline
      ↓
Evidence of nonlinear residual?
      │
      ├── pairwise material interaction → GBM
      ├── post-linear quadratic effect  → PPNM
      └── nonlinear form uncertain      → K-Hype
      ↓
Compare RMSE + abundance plausibility + spatial pattern

16.4 Vegetation stress feature engineering

Surface Reflectance
      ↓
Spectral Smoothing
      ↓
First Derivative
      ↓
Selected spectral features / ML
      ↓
Random Forest / SVM / Deep Learning

16.5 Direct spectral-library classification

Surface Reflectance
      +
Reference Library / Project Endmember Dataset
      ↓
Spectral resampling
      ↓
SAM / SID
      ↓
Class Raster
      +
Best-Match Score Raster

16.6 Why Unmixing and Matching are parallel

Matching result

Pixel 100,200
→ Kaolinite

One dominant material identity, subject to threshold and library design.

Unmixing result

Pixel 100,200
Kaolinite = 0.58
Calcite   = 0.27
Soil      = 0.15

Fractional composition. Multiple materials can coexist in the same pixel.

17. Scientific & Technical References

Primary papers and official technical documentation used to ground the definitions and implementation notes.

  1. Savitzky, A. & Golay, M. J. E. (1964). Smoothing and Differentiation of Data by Simplified Least Squares Procedures. Analytical Chemistry 36(8), 1627–1639. ACS.
  2. NV5 Geospatial / ENVI. Continuum Removal. Official documentation.
  3. NV5 Geospatial / ENVI. Minimum Noise Fraction Transform. Official documentation.
  4. NV5 Geospatial / ENVI. Forward MNF Transform Task. Official documentation.
  5. NV5 Geospatial / ENVI. Pixel Purity Index. Official documentation.
  6. Nascimento, J. M. P. & Bioucas-Dias, J. M. (2005). Vertex Component Analysis: A Fast Algorithm to Unmix Hyperspectral Data. IEEE Transactions on Geoscience and Remote Sensing 43(4), 898–910. IEEE.
  7. NV5 Geospatial / ENVI. SMACC Endmember Extraction. Official documentation.
  8. NV5 Geospatial / ENVI. Spectral Angle Mapper. Official documentation.
  9. Chang, C.-I. (2000). An Information-Theoretic Approach to Spectral Variability, Similarity, and Discrimination for Hyperspectral Image Analysis. IEEE Transactions on Information Theory 46(5), 1927–1932. IEEE.
  10. NV5 Geospatial / ENVI. Material Identification Tool — SAM, SID and spectral-library resampling. Official documentation.
  11. NV5 Geospatial / ENVI. Linear Spectral Unmixing. Official documentation.
  12. U.S. Geological Survey. USGS Spectral Library. Official USGS resource.
  13. U.S. Geological Survey. USGS Spectral Library Version 7 Data. Official data release.
  14. NV5 Geospatial / ENVI. Spectral Hourglass Workflow. Official documentation.
  15. Green, A. A., Berman, M., Switzer, P. & Craig, M. D. (1988). A Transformation for Ordering Multispectral Data in Terms of Image Quality with Implications for Noise Removal. IEEE Transactions on Geoscience and Remote Sensing 26(1), 65–74. DOI.
  16. Winter, M. E. (1999). N-FINDR: An Algorithm for Fast Autonomous Spectral End-member Determination in Hyperspectral Data. Proceedings of SPIE 3753, 266–275. DOI.
  17. Heinz, D. C. & Chang, C.-I. (2001). Fully Constrained Least Squares Linear Spectral Mixture Analysis Method for Material Quantification in Hyperspectral Imagery. IEEE Transactions on Geoscience and Remote Sensing 39(3), 529–545. DOI.
  18. Halimi, A., Altmann, Y., Dobigeon, N. & Tourneret, J.-Y. (2011). Unmixing Hyperspectral Images Using the Generalized Bilinear Model. IEEE IGARSS 2011, 1886–1889. DOI.
  19. Altmann, Y., Halimi, A., Dobigeon, N. & Tourneret, J.-Y. (2012). Supervised Nonlinear Spectral Unmixing Using a Postnonlinear Mixing Model for Hyperspectral Imagery. IEEE Transactions on Image Processing 21(6), 3017–3025. DOI.
  20. Chen, J., Richard, C. & Honeine, P. (2013). Nonlinear Unmixing of Hyperspectral Data Based on a Linear-Mixture/Nonlinear-Fluctuation Model. IEEE Transactions on Signal Processing 61(2), 480–492. DOI.
  21. Dobigeon, N., Tourneret, J.-Y., Richard, C., Bermudez, J. C. M., McLaughlin, S. & Hero, A. O. (2014). Nonlinear Unmixing of Hyperspectral Images: Models and Algorithms. IEEE Signal Processing Magazine 31(1), 82–94. DOI.
Implementation principle: the formulas in this document define the scientific contract. Exact numerical behavior still depends on choices such as edge handling, noise estimation, spectral resampling, regularization, thresholds, constraints and invalid-value policy. Those choices should be explicit in analysis parameters and provenance.