Hyperspectral Analysis
Technical & Scientific Documentation
A complete single-page reference for the planned hyperspectral toolbox: spectral preprocessing, dimensionality reduction, endmember discovery, spectral-library classification, and sub-pixel unmixing.
1. Scope & Final Hyperspectral Toolbox
This module starts after radiometric and atmospheric preprocessing. Its preferred input is a calibrated reflectance cube with reliable wavelength metadata. The tools are intentionally independent: users do not have to run every tool in sequence.
Hyperspectral ├─ Spectral Smoothing ├─ Spectral Derivative ├─ Continuum Removal ├─ Minimum Noise Fraction (MNF) ├─ Endmember Extraction ├─ Spectral Library Matching │ ├─ SAM — Spectral Angle Mapper │ └─ SID — Spectral Information Divergence └─ Spectral Unmixing
| Tool | Main Purpose | Input Bands | Main Output | Output Dimension |
|---|---|---|---|---|
| Spectral Smoothing | Reduce spectral noise while preserving spectral shape | N wavelength bands | Smoothed spectral cube | Usually N bands |
| Spectral Derivative | Emphasize slope, edge, and absorption-shape changes | N wavelength bands | Derivative spectral cube | N recommended; mathematically N−1/N−2 is also possible |
| Continuum Removal | Normalize absorption features relative to an upper continuum | N wavelength bands | Continuum-removed cube | N bands |
| MNF | Separate signal from noise and reduce dimensionality | N bands | M MNF components | M ≤ N, commonly M ≪ N |
| Endmember Extraction | Find spectrally pure/extreme signatures | N bands or selected MNF components | K spectral signatures + locations | K rows × N wavelengths in the preferred physical library |
| Spectral Library Matching | Whole-pixel material identification/classification | N spectral bands + reference spectra | Class/material raster + score | 1 class band + optional score/rule bands |
| Spectral Unmixing | Estimate sub-pixel material fractions | N bands + K endmembers | Abundance raster | K abundance bands (+ residual/RMSE optional) |
2. Scientific Foundation
2.1 Hyperspectral pixel as a vector
A hyperspectral pixel is not a single brightness value. It is a vector containing measurements across many narrow spectral bands. For a pixel with N bands:
If the image contains 200 bands, each pixel is a point in a 200-dimensional spectral space. A material such as vegetation, water, soil, kaolinite, calcite, or iron oxide can have a characteristic spectral shape or “spectral fingerprint”.
2.2 Reflectance domain
Reflectance is dimensionless. Physically, a reflectance of 0.20 corresponds to 20% reflectance under the measurement definition. In real corrected remote-sensing products, valid values can occasionally fall slightly below 0 or above 1 because of noise, correction artifacts, adjacency effects, scaling, or retrieval uncertainty. Therefore the analysis engine should not automatically clip all values to [0,1] unless the selected method specifically requires it.
2.3 Wavelength alignment is fundamental
For spectral matching or endmember-based analysis, the numerical band index is not enough. The engine should know each band's center wavelength and preferably its spectral response or FWHM. Two spectra can only be compared meaningfully when they represent the same wavelength space. Reference libraries therefore normally require spectral resampling to the image wavelengths.
3. Recommended Workflow
The seven tools form a toolbox, not a mandatory linear pipeline.
3.1 Two important final branches
Whole-pixel spectral classification
SAM/SID compare each image spectrum to a reference spectrum and can assign one dominant material/class to each pixel.
Pixel spectrum
+
Reference library
↓
SAM / SID
↓
1-band class mapSub-pixel decomposition
Spectral Unmixing estimates the fractional contribution of multiple endmembers to the same pixel.
Pixel spectrum
+
K endmembers
↓
Unmixing
↓
K abundance bands4. Recommended Input Data Contract
4.1 Required or strongly recommended metadata
| Metadata | Requirement | Reason |
|---|---|---|
| Band count | Required | Defines spectral vector dimension. |
| Center wavelength λ | Strongly recommended | Needed for derivatives, continuum analysis, spectral resampling and library matching. |
| Wavelength units | Required for matching | nm and μm must not be mixed silently. |
| FWHM / spectral response | Recommended | Improves library-to-sensor resampling. |
| Scale / offset | Required if encoded | Converts stored integer values to physical reflectance. |
| NoData / bad-band mask | Recommended | Prevents atmospheric absorption/noisy bands from corrupting analysis. |
| Product level | Recommended | Allows the system to distinguish DN/radiance/TOA/surface reflectance. |
4.2 Preferred data state
Preferred: Surface Reflectance + wavelength metadata + bad-band mask + valid scale/offset already applied or declared Accepted with warning: Radiance / TOA Reflectance Unknown wavelength metadata Generic band-index-only raster
Band-index-only mode can still support some mathematical processing, but physically meaningful spectral matching should be marked as limited or rejected when wavelength alignment cannot be established.
5. Spectral Smoothing
Spectral smoothing reduces high-frequency noise along the spectral axis of each pixel. It does not blur neighboring image pixels spatially. The implementation preserves the selected band count and wavelength metadata, so a 200-band input normally remains a 200-band spectral cube.
5.1 Available smoothing methods
| Method | Core idea | Best suited for | Main parameter |
|---|---|---|---|
| Savitzky–Golay | Local polynomial least-squares fit | Preserving spectral shape, peaks and absorption features | Window length + polynomial order |
| Moving Average | Uniform local mean | Simple low-pass suppression of random high-frequency noise | Window length |
| Gaussian | Distance-weighted local mean | Smooth low-pass filtering with softer weighting than a box average | σ + truncate |
| Median | Local order-statistic median | Spike / impulse / isolated-band noise | Window length |
5.2 Moving Average
Moving Average replaces each spectral sample by the arithmetic mean of an odd spectral neighborhood.
5.3 Gaussian Smoothing
Gaussian smoothing gives the largest weight to bands near the target wavelength and progressively smaller weights to more distant bands.
R̃i = Σj wj Ri-jσ controls spectral smoothing strength. The implementation limits the effective kernel radius with the truncate parameter, approximately radius ≈ truncate × σ bands.
Compared with a moving average, Gaussian weights reduce abrupt changes at the edge of the smoothing window.
5.4 Median Filter
The median filter is nonlinear. Instead of averaging values, it selects the median value inside the spectral neighborhood. It is especially useful when a spectrum contains isolated spikes that would strongly influence a mean.
5.5 Savitzky–Golay Smoothing
Savitzky–Golay smoothing fits a low-order polynomial to a moving spectral window by least squares. The polynomial value at the center of the window becomes the smoothed spectral value. It is widely used in spectroscopy because it can preserve local spectral shape better than a simple moving average.
5.6 Parameter contract in iT Sensing
| Method | Parameters shown | Implementation note |
|---|---|---|
| Savitzky–Golay | Window Length; Polynomial Order | Default method. Odd window ≥ 3 and polynomial order < window length. |
| Moving Average | Window Length | Uniform spectral filter; edge handling uses nearest-value extension. |
| Gaussian | Gaussian Sigma; Gaussian Truncate | Gaussian convolution only along the band axis. |
| Median | Window Length | Median neighborhood has shape [spectral window × 1 spatial sample]. |
5.7 Output and QC
Input : 200-band hyperspectral reflectance
↓
Spectral Smoothing
↓
Output: 200-band smoothed cube
Band 1 → same λ1
Band 2 → same λ2
...
Band 200 → same λ200
- The selected smoothing window must not exceed the selected spectral-band count.
- All four methods operate independently for each spatial pixel.
- Wavelength tags and band descriptions should be preserved.
- Bad atmospheric/noisy bands should preferably be excluded or handled before smoothing.
6. Spectral Derivative
Spectral derivatives emphasize changes in the shape of a spectrum. They can suppress broad baseline effects and reveal absorption edges, inflection points, and subtle spectral differences. The output is not reflectance; its units depend on wavelength units.
6.1 First derivative
6.2 Second derivative
For approximately uniform spectral spacing Δλ:
6.3 Savitzky–Golay derivative
The same local polynomial used for Savitzky–Golay smoothing can be differentiated analytically, often producing a more stable derivative than applying a finite difference directly to noisy spectra.
6.4 Why smoothing often precedes derivative
Differentiation amplifies high-frequency noise. Therefore:
Surface Reflectance
↓
Spectral Smoothing
↓
1st / 2nd Derivative6.5 Output dimension
Mathematically, naïve differencing may yield N−1 or N−2 samples. For a raster-analysis platform, the recommended contract is to preserve N output bands by using centered differences for interior wavelengths, one-sided differences or declared NoData at edges, and retaining original wavelength metadata.
7. Continuum Removal
Continuum removal normalizes a spectrum relative to an upper continuum, usually represented by a convex hull over the spectrum. It is especially useful for comparing the position, shape and depth of absorption features.
7.1 Basic equation
At wavelengths where the measured spectrum touches the continuum, RCR is close to 1. Inside an absorption feature, it is typically less than 1.
7.2 Absorption depth
7.3 Output
Input : 200-band reflectance Output: 200-band continuum-removed cube Original wavelength dimension is preserved. Output values are normalized spectral response, not original surface reflectance.
7.4 Spectral subset matters
The continuum depends on the wavelength interval used to construct it. The UI should therefore allow all valid bands, a wavelength range, or selected spectral windows.
8. Minimum Noise Fraction (MNF)
MNF is a noise-aware linear transform that orders transformed components by image quality / signal-to-noise information. Unlike ordinary PCA, MNF explicitly requires a model of the noise covariance. The implementation keeps signal-covariance estimation separate from the selectable noise-covariance estimator.
8.1 Signal + noise model
8.2 MNF generalized eigenproblem
A convenient formulation is to solve a generalized eigenproblem that compares total image covariance against noise covariance:
An equivalent two-stage interpretation first whitens noise and then rotates the whitened data.
y = VT zAfter whitening, the estimated noise covariance is approximately identity. V rotates the data into ordered MNF components.
8.3 Why the output can be smaller than the input
Input: 200 spectral bands
↓
MNF transform
↓
MNF 1 — high image quality / useful structure
MNF 2 — high image quality / useful structure
...
MNF 20 — retained
MNF 21+ — increasingly noise-dominated
↓
Output M = 20 components
MNF components are mathematical component axes. They are not wavelengths and cannot be directly matched against an ordinary laboratory spectral library unless the same fitted transform is applied to the reference spectra.
8.4 Noise Estimation Method
The four available options change only the estimate of Σn. Signal covariance is still estimated from spatially distributed valid image samples.
8.4.1 Spatial Shift Difference — default
Adjacent pixels are assumed to have similar local signal, so their difference emphasizes noise. Horizontal and vertical valid neighbor pairs are sampled.
Σ̂n = ½ Cov(d)The factor ½ is equivalent to using d/√2 when the two neighboring noise realizations are independent with the same covariance.
This is a practical general-purpose default when no dedicated noise calibration is available.
8.4.2 Homogeneous ROI Difference
The same shift-difference principle is applied only inside a user-selected polygon that is expected to be spectrally homogeneous. Restricting the estimate to a homogeneous region reduces contamination by genuine scene transitions.
Σ̂n,ROI = ½ Cov { dp,δ | p ∈ Ωh }Ωh is the homogeneous-noise ROI. The ROI should contain enough valid adjacent pixels and should avoid strong material boundaries.
8.4.3 Local Residual Estimation
A local spatial predictor is used to estimate the slowly varying signal around each pixel. The residual is treated as a noise sample. The current implementation uses a local mean excluding the center pixel.
rp = xp − x̂p
Σ̂n = Cov(r)N(p) is an odd spatial neighborhood such as 3×3, 5×5, etc. A larger window models smoother local signal but can mix genuine spatial variation into the residual.
8.4.4 External Noise Covariance
If sensor calibration, dark-current measurements, laboratory characterization, or another trusted process already provides a noise covariance matrix, the matrix can be supplied directly.
This option bypasses image-derived noise sampling and makes the external covariance the authoritative noise model.
8.5 Noise-estimator comparison
| Estimator | Additional input | Strength | Main caution |
|---|---|---|---|
| Spatial Shift Difference | None | Fast, general-purpose, automatic | Scene edges can contribute real signal differences |
| Homogeneous ROI Difference | Homogeneous polygon ROI | Better isolates noise when a good homogeneous region exists | Result depends on ROI quality and representativeness |
| Local Residual Estimation | Spatial window size | Useful in heterogeneous scenes without a dedicated ROI | True fine spatial structure can enter the residual |
| External Noise Covariance | B×B covariance matrix | Best when a trusted sensor/noise model is available | Matrix must correspond to the selected bands and data scaling |
8.6 Component selection
- Auto: lets the analysis choose an informative component count from the fitted component statistics.
- Manual: user specifies M explicitly.
- Covariance Sample Pixels: bounded distributed samples used for covariance estimation; full-raster transformation remains windowed.
- Regularization: a small diagonal stabilization is applied when the noise covariance is ill-conditioned.
9. Endmember Extraction
An endmember is a spectral signature representing a relatively pure material or an extreme spectral component in the scene. Endmember Extraction produces a reusable spectral-library dataset, not a classification raster. The current toolbox implements three classical pure-pixel / geometric methods: PPI, N-FINDR and VCA.
9.1 Geometric basis
Under a linear mixture model, observed spectra approximately occupy a simplex in spectral space. Pure materials tend to occur near simplex vertices, while mixtures occur inside or along simplex edges.
ak ≥ 0, Σk=1…Kak ≈ 1M contains endmember spectra as columns. Pure-pixel extraction methods search for extreme observations that approximate the vertices of this simplex.
Extraction can be performed from a sampled subset for efficiency, but the selected indices should resolve back to real source pixels so that coordinates and original spectra can be preserved.
9.2 PPI — Pixel Purity Index
PPI projects the spectral cloud onto many random directions (“skewers”). Pixels repeatedly appearing at projection minima or maxima receive high purity counts.
PPI(j) = Σr=1…R I[j = arg min z·r or j = arg max z·r]yj is the centered/reduced spectrum of pixel j; ur is a random unit direction; R is the number of skewers.
The implementation first projects the sampled spectra to a principal signal subspace and scales those axes before random projection. A minimum spectral-angle criterion can prevent nearly duplicate PPI endmembers.
9.3 N-FINDR — Maximum-volume simplex
N-FINDR seeks the set of K observed pixels that maximizes the volume of the simplex formed by those candidate endmembers. The sampled spectra are projected into K−1 affine signal dimensions and augmented by a constant coordinate.
V(A) = |det(A)| / (K−1)!Each ỹk is an augmented candidate vertex. Larger |det(A)| means a larger spectral simplex.
For a current simplex A, replacing vertex j by candidate v changes the determinant by a column-replacement factor:
N-FINDR is computationally more expensive as K grows; the current server implementation limits N-FINDR to 32 endmembers per run.
9.4 VCA — Vertex Component Analysis
VCA exploits the fact that linear mixtures live in a low-dimensional affine subspace. For K endmembers, sampled spectra are projected to K−1 signal dimensions and augmented by a constant coordinate. Vertices are then identified through successive projections in directions constrained to be independent of previously selected vertices.
fk ⟂ span{ỹj₁, …, ỹjk−1}
jk = arg maxj | fkT ỹj |U spans the K−1 dimensional signal subspace. Each selected jk is an actual sampled pixel used as an endmember.
VCA is generally fast because it avoids the repeated full simplex-volume search of N-FINDR.
9.5 Method comparison and parameters
| Method | Geometric criterion | Parameters exposed | Use case |
|---|---|---|---|
| PPI Default | Frequency of extreme random projections | Endmember Count; Sample Pixels; Random Skewers; Projection Components; Minimum Separation Angle; Random Seed | Robust pure-pixel ranking, especially after MNF |
| N-FINDR | Maximum simplex volume | Endmember Count; Sample Pixels; Maximum Iterations; Random Seed | Strong geometric pure-pixel assumption and moderate K |
| VCA | Successive independent vertex projections | Endmember Count; Sample Pixels; Random Seed | Fast geometric extraction for many practical scenes |
9.6 Output object and project dataset
Hyperspectral Raster
↓
PPI / N-FINDR / VCA
↓
K selected source pixels
↓
Spectral Library JSON in object storage
↓
public.layers
data_type = endmember
↓
Reusable by SAM / SID / Spectral Unmixing
Each signature should retain its source row/column and geospatial coordinate, spectral-axis metadata, extraction method, quality/selection score and preprocessing provenance.
{
"analysis": "endmember_extraction",
"method": "VCA",
"source_raster": "...",
"wavelength_unit": "nm",
"wavelengths": [400.0, 410.0, "..."],
"endmembers": [
{
"id": 1,
"name": "Endmember 1",
"pixel": {"row": 812, "col": 455},
"coordinate": {"x": "...", "y": "..."},
"spectrum": [0.081, 0.084, "..."]
}
]
}
10. Spectral Library Matching
Spectral Library Matching is the hyperspectral material-identification / whole-pixel classification tool. The user supplies or selects a reference spectral library, and each image pixel is compared against candidate reference spectra.
10.1 Library preprocessing requirements
- Reference wavelengths must overlap image wavelengths.
- Wavelength units must be known and harmonized.
- Library spectra should be resampled using wavelength/FWHM information when possible.
- If image spectra are continuum-removed, derivative-transformed, or otherwise normalized, the reference library must receive the equivalent transformation.
- Bad bands should be excluded from both image and reference spectra.
10.2 SAM — Spectral Angle Mapper
SAM treats each spectrum as a vector in N-dimensional band space and computes the angle between image spectrum x and reference spectrum s.
- θ ≈ 0: highly similar spectral direction.
- Larger θ: less similar.
- Best reference is typically the class with the smallest valid angle.
- A maximum-angle threshold should allow pixels to remain unclassified.
Because SAM focuses on vector direction rather than magnitude, it can be relatively insensitive to overall albedo/illumination scaling when used with calibrated data; this does not remove the need for correct radiometric/atmospheric preprocessing.
Recommended SAM outputs
classification.tif 0 = Unclassified 1 = Material A 2 = Material B ... sam_best_angle.tif pixel value = minimum spectral angle optional rule images: one angle band per reference material
10.3 SID — Spectral Information Divergence
SID converts nonnegative spectral vectors into probability-like distributions and measures their divergence using information theory.
Step 1 — normalize spectra
Step 2 — directional divergence
Step 3 — symmetric SID
- SID close to 0: highly similar.
- Larger SID: greater spectral divergence.
- A maximum SID threshold should support an unclassified class.
Recommended SID outputs
classification.tif best material/reference class sid_best_score.tif minimum SID divergence optional rule images: one SID score band per reference material
10.4 SAM vs SID
| Aspect | SAM | SID |
|---|---|---|
| Core idea | Angle between spectral vectors | Information divergence between normalized spectral distributions |
| Best match | Smallest angle | Smallest divergence |
| Output class map | Yes | Yes |
| Reference library | Required | Required |
| Zero/negative handling | Vector norm must be valid | Requires careful positive-value handling for logarithms |
| Threshold | Maximum angle | Maximum divergence |
11. Spectral Unmixing
Spectral Unmixing estimates how much each endmember contributes to an observed pixel. The implementation is supervised abundance estimation: the endmember matrix is supplied by a project Endmember dataset or a selected public reference library. The toolbox now provides six methods: three linear and three nonlinear.
11.1 Common notation
For B spectral bands and K endmembers:
The linear mixing model (LMM) is:
Nonlinear methods extend this relationship when multiple scattering, material interaction, intimate mixing, canopy structure or other effects make a purely linear model inadequate.
11.2 UCLS — Unconstrained Least Squares
UCLS minimizes reconstruction error without abundance constraints.
When M has suitable rank, the least-squares solution can be written with the Moore–Penrose pseudoinverse:
- Fast baseline linear solution.
- Abundances may be negative.
- Abundances are not required to sum to one.
11.3 NNLS — Non-Negative Least Squares
NNLS adds the abundance non-negativity constraint (ANC).
subject to ak ≥ 0
Negative material fractions are prevented, but the abundance sum is not forced to one.
11.4 FCLS — Fully Constrained Least Squares
FCLS adds both abundance non-negativity (ANC) and abundance sum-to-one (ASC). It remains the default method for fractional abundance mapping.
subject to ak ≥ 0, Σk=1…Kak = 1
The feasible abundance vectors lie on the probability simplex:
The implementation uses projected-gradient iterations with simplex projection.
11.5 Why nonlinear unmixing?
A linear model assumes photons interact with only one material before reaching the sensor. Real scenes can contain second-order scattering, canopy–soil interaction, intimate mineral mixtures, complex urban surfaces or other mechanisms that distort this assumption. Nonlinear unmixing introduces a structured or flexible nonlinear term rather than forcing all residual structure into e.
11.6 GBM — Generalized Bilinear Model
GBM explicitly models pairwise second-order interactions between endmembers using Hadamard (band-by-band) products.
The number of pairwise terms is:
Because Q grows quadratically, the current server implementation limits GBM to 16 endmembers per run.
11.7 PPNM — Polynomial Post-Nonlinear Model
PPNM first forms the linear mixture and then applies a polynomial post-nonlinearity. In the implemented quadratic form:
x = z + b(z ⊙ z) + e
When b ≈ 0, PPNM approaches the linear mixing model.
11.8 K-Hype — Kernel-Based Semiparametric Unmixing
K-Hype models a linear abundance term plus a flexible nonlinear residual function living in a reproducing kernel Hilbert space (RKHS).
The implementation uses an RBF kernel across the band-wise endmember vectors:
After eliminating the RKHS residual analytically, abundance estimation is performed with a kernel-weighted quadratic objective:
â = arg mina∈ΔK (x − Ma)T W (x − Ma)λ is the RKHS regularization parameter. σ is the RBF bandwidth; a value of 0 in the UI invokes the automatic median-distance rule.
K-Hype constructs a dense band×band kernel matrix; the current implementation therefore limits a run to 512 selected spectral bands.
11.9 Method comparison
| Method | Family | Abundance constraints | Nonlinear assumption | Typical role |
|---|---|---|---|---|
| UCLS | Linear | None | None | Fast baseline / diagnostic |
| NNLS | Linear | a ≥ 0 | None | Nonnegative linear abundance |
| FCLS Default | Linear | a ≥ 0; Σa = 1 | None | Default fractional abundance mapping |
| GBM | Nonlinear | a ≥ 0; Σa = 1 | Pairwise bilinear interactions | Second-order scattering / material interactions |
| PPNM | Nonlinear | a ≥ 0; Σa = 1 | Quadratic post-nonlinearity | Smooth post-linear spectral distortion |
| K-Hype | Nonlinear / kernel | a ≥ 0; Σa = 1 | Flexible RKHS nonlinear residual | Unknown or difficult-to-parametrize nonlinear effects |
11.10 Output band count
Input hyperspectral raster = 285 bands
Selected endmembers = 6
↓
Spectral Unmixing
↓
Output abundance raster = 6 bands
Band 1 = abundance Endmember 1
Band 2 = abundance Endmember 2
...
Band 6 = abundance Endmember 6
For UCLS the output values are unconstrained coefficients. For NNLS they are nonnegative coefficients. FCLS, GBM, PPNM and K-Hype produce simplex-constrained fractions under the current implementation.
11.11 Residual and reconstruction error
For nonlinear methods, x̂ must include the fitted nonlinear contribution. High residuals can indicate missing endmembers, poor wavelength alignment, atmospheric/bad-band contamination, spectral variability, or a mixing mechanism not represented by the selected method.
11.12 Method-selection guidance
- Start with FCLS when the goal is interpretable material fractions and linear mixing is plausible.
- Compare UCLS/NNLS to diagnose whether the constraints materially change the solution.
- Use GBM when pairwise material interactions / multiple scattering are scientifically plausible.
- Use PPNM when a post-linear nonlinear distortion is plausible.
- Use K-Hype when the nonlinear mechanism is not well known and a flexible kernel residual is preferred.
- Nonlinear models have more degrees of freedom and computational cost; better fit does not automatically mean better physical interpretation.
12. Relationship Between the Seven Tools
| From | Can Feed | Notes |
|---|---|---|
| Spectral Smoothing | Derivative, Continuum Removal, MNF, Endmember Extraction, Matching, Unmixing, ML/DL | Preserves wavelength domain. |
| Spectral Derivative | Matching, ML/DL, feature analysis | Reference library must use equivalent derivative for matching. |
| Continuum Removal | Matching, feature analysis, ML/DL | Use the same spectral window on reference spectra. |
| MNF | Endmember Extraction, ML/DL, advanced unmixing | Components are not wavelengths. |
| Endmember Extraction | Library Matching, Spectral Unmixing | Endmembers may first need material identification. |
| Library Matching | Final class/material analysis, GIS statistics | Not a prerequisite for Unmixing. |
| Spectral Unmixing | Abundance analysis, thresholding, ML/DL, dominant-material derivation | Output bands are abundance dimensions. |
12.1 Valid workflows
Surface Reflectance → Smoothing → SAM/SID
Surface Reflectance → Smoothing → Continuum Removal → SAM/SID
Surface Reflectance → Smoothing → Derivative → Random Forest
Surface Reflectance → MNF → Random Forest
Surface Reflectance → MNF → Endmember Extraction
→ sample original spectra
→ Spectral Unmixing
Surface Reflectance → Endmember Extraction → SAM/SID
External Spectral Library → resample to image wavelengths → SAM/SID12.2 Workflow to avoid by default
Spectral Unmixing → SAM/SID spectral-library matching ✕
After unmixing, the bands represent abundance of endmembers rather than wavelength samples.
13. Recommended Platform Output Contract
| Tool | Main File | Sidecar / Auxiliary | Layer Type |
|---|---|---|---|
| Smoothing | Float32 COG, N bands | .statistics.json | Raster |
| Derivative | Float32 COG, N bands | .statistics.json + derivative metadata | Raster |
| Continuum Removal | Float32 COG, N bands | .statistics.json + spectral range | Raster |
| MNF | Float32 COG, M bands | .statistics.json + MNF transform/statistics metadata | Raster |
| Endmember Extraction | endmembers.json | CSV, optional point layer, optional purity raster | Endmember dataset (public.layers data_type=endmember) |
| Library Matching | Integer class COG | confidence/score COG, class legend JSON | Raster classification |
| Unmixing | Float32 COG, K bands | .statistics.json + endmember legend + RMSE/residual optional | Raster abundance |
13.1 Suggested endmember JSON fields
- analysis/method/version;
- source raster identifier;
- spectral domain and wavelength units;
- wavelength array;
- endmember identifier and optional material label;
- source pixel row/column and geospatial coordinate;
- spectral vector;
- purity/quality score where available;
- preprocessing history;
- bad-band list;
- creation timestamp.
13.2 Suggested classification legend
{
"classes": [
{"value": 0, "name": "Unclassified"},
{"value": 1, "name": "Kaolinite", "reference_id": "lib-001"},
{"value": 2, "name": "Calcite", "reference_id": "lib-021"},
{"value": 3, "name": "Vegetation", "reference_id": "lib-104"}
],
"method": "SAM",
"score_semantics": "lower_is_better"
}13.3 Suggested abundance metadata
{
"bands": [
{"band": 1, "endmember": "Vegetation", "unit": "fraction"},
{"band": 2, "endmember": "Soil", "unit": "fraction"},
{"band": 3, "endmember": "Shadow", "unit": "fraction"}
],
"method": "FCLS",
"family": "linear",
"sum_to_one": true,
"nonnegative": true
}14. Validation & Quality Control
14.1 Common validation
- Reject empty rasters and invalid spectral dimensions.
- Verify NoData behavior and finite-value proportion.
- Verify wavelength-array length equals raster band count.
- Check wavelengths are monotonically ordered or reorder data and metadata together.
- Detect duplicate wavelengths.
- Apply declared scale/offset before physics-based spectral analysis.
- Exclude known bad bands and strong atmospheric-absorption regions when appropriate.
14.2 Smoothing QC
- Savitzky–Golay, Moving Average and Median windows must be odd and cannot exceed the selected band count.
- Savitzky–Golay polynomial order must be smaller than the window length.
- Gaussian σ must be positive and truncate must define a finite practical kernel radius.
- Confirm filtering occurs only along the spectral axis; neighboring spatial pixels must not be mixed.
- Inspect absorption-feature preservation before choosing aggressive smoothing parameters.
14.3 Derivative QC
- Reject zero wavelength spacing.
- Handle irregular wavelength spacing explicitly.
- Record derivative order and units.
14.4 MNF QC
- Record the selected noise estimator: Spatial Shift Difference, Homogeneous ROI Difference, Local Residual, or External Covariance.
- For homogeneous ROI estimation, verify the polygon is sufficiently large, valid and spectrally homogeneous.
- For local residual estimation, verify the odd spatial window is appropriate for scene texture.
- For external covariance, require a finite B×B matrix matching the selected bands and data scaling.
- Noise covariance must be numerically valid; regularize near-singular matrices.
- Return eigenvalues/component-quality information and never label MNF components as wavelengths.
14.5 Endmember QC
- K must be consistent with spectral dimensionality and available pure-pixel structure.
- PPI: inspect purity counts and minimum-angle duplicate suppression.
- N-FINDR: inspect simplex-volume convergence; avoid excessive K and remember the server limit of 32 endmembers.
- VCA: verify selected vertices are unique and source pixels are valid.
- Report duplicate/highly similar extracted endmembers, source pixels and validity masks.
- Allow user interpretation/renaming after extraction; geometric extremeness is not automatically a material identity.
14.6 Matching QC
- Require spectral overlap and harmonized wavelength units.
- Resample the selected reference library to the image spectral grid.
- Use thresholds to avoid forced classification.
- Return best score and, where useful, second-best score/margin.
- For SID, handle nonpositive values explicitly.
14.7 Unmixing QC
- Endmember and image spectral dimensions must match after resampling/transformation.
- Use UCLS as a useful unconstrained baseline but do not interpret negative coefficients as physical fractions.
- NNLS must satisfy non-negativity; FCLS, GBM, PPNM and K-Hype should satisfy both non-negativity and sum-to-one under the implemented solver.
- GBM: K ≤ 16 in the current implementation because interaction count grows as K(K−1)/2.
- PPNM: inspect the fitted nonlinearity magnitude and chosen maximum |b|.
- K-Hype: inspect RKHS regularization and kernel bandwidth; selected band count is currently limited to 512.
- Return reconstruction residual/RMSE and compare nonlinear methods against a linear baseline.
15. Implementation Notes for an Analysis Server
15.1 Processing strategy
A hyperspectral cube can be hundreds of bands and several gigabytes. Algorithms should use block/window processing where mathematically possible. Do not read a full cube into RAM unless required and size-checked.
| Tool | Window/Block Friendly? | Global Preparation Needed? |
|---|---|---|
| Smoothing | Yes | No, except metadata |
| Derivative | Yes | No, except wavelength array |
| Continuum Removal | Yes, pixel blocks | No |
| MNF | Transform yes | Noise/covariance statistics first |
| Endmember Extraction | Method dependent | Often sampling/reduction/global search |
| SAM/SID | Yes | Reference resampling/preprocessing first |
| Unmixing | Yes | Endmember matrix/solver setup first |
15.2 Dtype
- Use Float32 for transformed spectral/abundance outputs unless higher precision is specifically required.
- Use integer type for class raster, choosing UInt8/UInt16/UInt32 based on class count.
- Do not encode transformed physical values into integers unless scale/offset is explicit.
15.3 Raster statistics
A statistics sidecar should describe each output band correctly. For a 200-band cube this can be large; the implementation may store per-band min/max/mean/std plus sampled histograms. MNF and abundance outputs should have semantic band descriptions.
15.4 Provenance
{
"input_domain": "surface_reflectance",
"operations": [
{"tool": "spectral_smoothing", "method": "savitzky_golay", "window": 7, "order": 2},
{"tool": "mnf", "noise_estimation": "homogeneous_roi_difference", "components": 20},
{"tool": "continuum_removal", "range_nm": [2000, 2400]}
]
}Provenance is essential because a reference library must receive compatible preprocessing before matching.
15.5 Do not silently change spectral semantics
| Output | Band semantics |
|---|---|
| Smoothed cube | Wavelength |
| Derivative cube | Derivative at wavelength |
| Continuum-removed cube | Normalized value at wavelength |
| MNF raster | MNF component |
| SAM/SID class raster | Class ID |
| Unmixing raster | Endmember abundance |
16. Example Analytical Workflows
16.1 Mineral identification from reflectance
Surface Reflectance
↓
Remove bad bands
↓
Spectral Smoothing
(SG / Gaussian / Moving Average / Median)
↓
Continuum Removal (target SWIR range)
↓
Resample mineral spectral library to image wavelengths
↓
Apply same continuum removal to library
↓
Spectral Library Matching — SAM or SID
↓
Mineral Class Map + Match Score
16.2 Endmember discovery and linear abundance mapping
Surface Reflectance (200 bands)
↓
MNF
Noise estimator:
Spatial / Homogeneous ROI / Local Residual / External Covariance
↓
retain informative components
↓
PPI / N-FINDR / VCA
↓
K pure-pixel locations
↓
sample original reflectance
↓
K endmember spectra
↓
FCLS (default) / NNLS / UCLS
↓
K abundance bands + reconstruction QC
16.3 Nonlinear abundance mapping
Surface Reflectance
+
Focused Endmember Library
↓
Start with FCLS baseline
↓
Evidence of nonlinear residual?
│
├── pairwise material interaction → GBM
├── post-linear quadratic effect → PPNM
└── nonlinear form uncertain → K-Hype
↓
Compare RMSE + abundance plausibility + spatial pattern
16.4 Vegetation stress feature engineering
Surface Reflectance
↓
Spectral Smoothing
↓
First Derivative
↓
Selected spectral features / ML
↓
Random Forest / SVM / Deep Learning
16.5 Direct spectral-library classification
Surface Reflectance
+
Reference Library / Project Endmember Dataset
↓
Spectral resampling
↓
SAM / SID
↓
Class Raster
+
Best-Match Score Raster
16.6 Why Unmixing and Matching are parallel
Matching result
Pixel 100,200 → Kaolinite
One dominant material identity, subject to threshold and library design.
Unmixing result
Pixel 100,200 Kaolinite = 0.58 Calcite = 0.27 Soil = 0.15
Fractional composition. Multiple materials can coexist in the same pixel.
17. Scientific & Technical References
Primary papers and official technical documentation used to ground the definitions and implementation notes.
- Savitzky, A. & Golay, M. J. E. (1964). Smoothing and Differentiation of Data by Simplified Least Squares Procedures. Analytical Chemistry 36(8), 1627–1639. ACS.
- NV5 Geospatial / ENVI. Continuum Removal. Official documentation.
- NV5 Geospatial / ENVI. Minimum Noise Fraction Transform. Official documentation.
- NV5 Geospatial / ENVI. Forward MNF Transform Task. Official documentation.
- NV5 Geospatial / ENVI. Pixel Purity Index. Official documentation.
- Nascimento, J. M. P. & Bioucas-Dias, J. M. (2005). Vertex Component Analysis: A Fast Algorithm to Unmix Hyperspectral Data. IEEE Transactions on Geoscience and Remote Sensing 43(4), 898–910. IEEE.
- NV5 Geospatial / ENVI. SMACC Endmember Extraction. Official documentation.
- NV5 Geospatial / ENVI. Spectral Angle Mapper. Official documentation.
- Chang, C.-I. (2000). An Information-Theoretic Approach to Spectral Variability, Similarity, and Discrimination for Hyperspectral Image Analysis. IEEE Transactions on Information Theory 46(5), 1927–1932. IEEE.
- NV5 Geospatial / ENVI. Material Identification Tool — SAM, SID and spectral-library resampling. Official documentation.
- NV5 Geospatial / ENVI. Linear Spectral Unmixing. Official documentation.
- U.S. Geological Survey. USGS Spectral Library. Official USGS resource.
- U.S. Geological Survey. USGS Spectral Library Version 7 Data. Official data release.
- NV5 Geospatial / ENVI. Spectral Hourglass Workflow. Official documentation.
- Green, A. A., Berman, M., Switzer, P. & Craig, M. D. (1988). A Transformation for Ordering Multispectral Data in Terms of Image Quality with Implications for Noise Removal. IEEE Transactions on Geoscience and Remote Sensing 26(1), 65–74. DOI.
- Winter, M. E. (1999). N-FINDR: An Algorithm for Fast Autonomous Spectral End-member Determination in Hyperspectral Data. Proceedings of SPIE 3753, 266–275. DOI.
- Heinz, D. C. & Chang, C.-I. (2001). Fully Constrained Least Squares Linear Spectral Mixture Analysis Method for Material Quantification in Hyperspectral Imagery. IEEE Transactions on Geoscience and Remote Sensing 39(3), 529–545. DOI.
- Halimi, A., Altmann, Y., Dobigeon, N. & Tourneret, J.-Y. (2011). Unmixing Hyperspectral Images Using the Generalized Bilinear Model. IEEE IGARSS 2011, 1886–1889. DOI.
- Altmann, Y., Halimi, A., Dobigeon, N. & Tourneret, J.-Y. (2012). Supervised Nonlinear Spectral Unmixing Using a Postnonlinear Mixing Model for Hyperspectral Imagery. IEEE Transactions on Image Processing 21(6), 3017–3025. DOI.
- Chen, J., Richard, C. & Honeine, P. (2013). Nonlinear Unmixing of Hyperspectral Data Based on a Linear-Mixture/Nonlinear-Fluctuation Model. IEEE Transactions on Signal Processing 61(2), 480–492. DOI.
- Dobigeon, N., Tourneret, J.-Y., Richard, C., Bermudez, J. C. M., McLaughlin, S. & Hero, A. O. (2014). Nonlinear Unmixing of Hyperspectral Images: Models and Algorithms. IEEE Signal Processing Magazine 31(1), 82–94. DOI.