Langford Analytic · Knowledge Base

Surrogate Models & Response Surfaces

Inexpensive algebraic approximations built from a limited set of high-fidelity runs — how they are constructed, where interpolation ends and extrapolation begins, and why a simple polynomial is often the right answer.

Article 03Computational Reduction & Approximation14 min read
surrogate modelsresponse surfacespolynomial response surfaceradial basis functionsGaussian processeskrigingneural-network surrogatesinterpolationextrapolationvalidationdesign of experiments

What a surrogate is

A surrogate model is an inexpensive approximation to an expensive analysis. The expensive analysis — a nonlinear finite-element run, a CFD solve, a coupled thermal-structural simulation — is treated as a black-box function that maps a vector of inputs to an output. A small number of high-fidelity runs are performed at chosen input points; from those known input-output pairs a cheaper mathematical model is constructed that can be evaluated anywhere in the design space at a fraction of the cost. The surrogate is then used wherever many evaluations are needed: design optimisation, uncertainty propagation, tolerance studies, trade studies. The surrogate never replaces the high-fidelity model for the final verification of a chosen design; it makes the search for that design affordable.

Output of expensive model:   y = f(x)
Surrogate approximation:     ŷ = f̂(x)
where:
  x  = vector of inputs (design variables, loads, material properties, etc.)
  y  = true output of the high-fidelity model at x
  f̂  = surrogate model (cheap algebraic approximation)
  ŷ  = surrogate prediction at x

Families of surrogate, conceptually

Several functional forms are used to build surrogates, and the choice should not become an end in itself. Polynomial response surfaces fit a low-order polynomial to the data; they are transparent, cheap, and often perfectly adequate for smooth, mildly nonlinear responses over a bounded region. Radial basis function models build the surface from radially symmetric basis functions centred at the training points; they interpolate the training data and behave well for moderate-dimensional problems. Gaussian-process models (often called kriging in the engineering literature) provide not only a prediction but also a measure of predicted uncertainty that grows with distance from the training data — a property that is useful but whose statistical assumptions must be understood, not taken on trust. Neural-network surrogates can represent highly nonlinear responses but require more data, are harder to interrogate, and are easy to overfit. Interpolation methods such as inverse-distance or nearest-neighbour schemes are simple but can produce rough surfaces. The engineering question is not "which algorithm is most sophisticated?" but "which algorithm describes this response adequately with the data I can afford?"

  • Polynomial response surfaces — low-order polynomial fit; transparent and cheap.
  • Radial basis functions — radially symmetric bases centred at training points; interpolate the data.
  • Gaussian processes / kriging — prediction with a distance-based uncertainty measure.
  • Neural-network surrogates — flexible for highly nonlinear responses; data-hungry and harder to interrogate.
  • Interpolation methods — inverse-distance, nearest-neighbour; simple but can produce rough surfaces.

The training set and what it contains

A surrogate is built from a training set: a collection of high-fidelity runs at chosen input points, each with its known output. The information content of the surrogate is exactly the information content of that training set. If the training set does not sample a region of the design space, the surrogate has no knowledge of the response there. If the training set is clustered, the surrogate is well informed near the cluster and poorly informed elsewhere. This is why the design of the training set — the choice of where to run the expensive model — is as important as the choice of surrogate form, and is the subject of a separate article. A surrogate built from an arbitrary or convenient set of runs inherits the biases of that set.

A SURROGATE CAN ONLY LEARN THE BEHAVIOUR PRESENT IN THE DATA USED TO BUILD IT.

From sparse simulation points to a prediction surface

The diagram shows the construction. A sparse set of simulation points, each with a known high-fidelity response, is used to fit a surrogate surface. Predictions at new points inside the region spanned by the training data are interpolations; predictions outside that region are extrapolations. The distinction is not a technicality: interpolation is supported by data on either side, extrapolation is supported by data on one side only and by the assumed form of the surrogate. Confidence in a surrogate prediction should be tied to where the prediction sits relative to the training data.

SURROGATE CONSTRUCTION AND PREDICTION

  INPUT x (design / load / material variables)
  │
  │      ●           ●                ●
  │       \         / \              /
  │        \       /   \            /         ◆ new prediction
  │         ●─────●     ●──────────●              (interpolation —
  │        /       \   /              \            inside training
  │       /         \ /                \           region)
  │      ●           ●                  ●
  │
  │  ◆ new prediction (extrapolation — outside training region)
  │
  └────────────────────────────────────────────────► x

  ● = training point (high-fidelity run, known response)
  ◆ = surrogate prediction at a new point

  ┌──────────────────────────────────────────────────────┐
  │  INTERPOLATION REGION  │  EXTRAPOLATION REGION        │
  │  (between training     │  (outside training data;     │
  │   data; supported)     │   confidence should fall)    │
  └──────────────────────────────────────────────────────┘

CONFIDENCE SHOULD GENERALLY FALL AS THE MODEL MOVES OUTSIDE THE DOMAIN REPRESENTED BY ITS TRAINING DATA.

Validation: checking the surrogate against the truth

A surrogate must be checked against the high-fidelity model it replaces, and that check must use data that was not used to build the surrogate. A validation set of high-fidelity runs at points independent of the training set is compared against the surrogate predictions at those points. The error is mapped across the design space, not only reported as a single number: a low average error can conceal a high-error region exactly where the engineering decision is sensitive. Boundary regions, where extrapolation begins, deserve particular attention. If a region shows unacceptable error, the high-fidelity model must be used there — or the training set must be enriched and the surrogate rebuilt. The surrogate is a tool with a known, mapped accuracy; it is not a replacement for the high-fidelity model everywhere.

  • Training-data coverage adequate for the design space? — No large unsampled regions in the area of interest.
  • Validation points independent of training points? — Not re-using training data to report accuracy.
  • Errors mapped across the design space, not just averaged? — A single global error can hide a local high-error region.
  • Boundary regions checked? — Extrapolation begins at the boundary of the training data.
  • Extrapolation detected and flagged in use? — Predictions outside the training domain should be visible to the user.
  • Discontinuities or sharp gradients present? — Smooth surrogates smear discontinuities such as contact onset or buckling.
  • Physical monotonicity expected? — A response that must be monotone in a variable should be checked for it.
  • Predicted trends physically plausible? — The surrogate should not predict behaviour the physics forbids.
  • High-error region returned to high-fidelity model? — Do not rely on the surrogate where it is known to be wrong.

Surrogate methods and their characteristics

The table compares the common surrogate forms. The "right" choice is the simplest form that adequately describes the response over the region of interest with the data available. Surrogate selection is an engineering judgement informed by the character of the response, the affordable number of high-fidelity runs and the required accuracy — not by the sophistication of the algorithm.

MethodApproachData requirementsInterpolation characterExtrapolation riskWhen appropriateLimitations
Polynomial response surfaceFit low-order polynomial to training dataLow; works with modest point countsSmooths over data; does not generally interpolateHigh; polynomial grows unbounded outside dataSmooth, mildly nonlinear response over a bounded regionCannot capture sharp curvature, discontinuities or high nonlinearity
Radial basis functionSum of radially symmetric bases centred at training pointsModerate; scales with dimensionInterpolates training data exactlyModerate; behaviour depends on basis and shape parameterModerate-dimensional smooth responses where interpolation is wantedShape-parameter selection affects quality; can be ill-conditioned
Gaussian process / krigingProbabilistic model with distance-based uncertaintyModerate; covariance estimation needs careInterpolates training data (with suitable covariance)Moderate; uncertainty grows away from dataWhere a predicted uncertainty is useful for adaptive samplingCovariance assumptions must be understood, not assumed; cost scales with data count
Neural-network surrogateFlexible nonlinear function fit by trainingHigh; needs substantial data and validationDoes not generally interpolate; can overfitHigh; flexible form can extrapolate wildlyHighly nonlinear responses with ample dataOpaque, data-hungry, easy to overfit; hard to interrogate

Where surrogates are used

Surrogates are most valuable in tasks that require many evaluations of an expensive model. In optimisation, the search moves on the surrogate and the high-fidelity model is called only to confirm or refine promising candidates. In uncertainty propagation, the surrogate replaces the expensive model inside a Monte Carlo loop so that thousands of samples become affordable. In tolerance and trade studies, the surface can be interrogated interactively to see how the response moves with the inputs. In every case the surrogate is a means of searching an otherwise unaffordable space; the final design is still verified with the high-fidelity model.

The classic misuse: complexity without value

The most common misuse is building a sophisticated surrogate — a neural network, a deep Gaussian process — when a simple polynomial response surface already describes the problem adequately. The added complexity does not improve the engineering decision; it adds data requirements, validation burden, opacity and maintenance. If a second-order polynomial captures the response to within the tolerance the decision requires, that is the right surrogate. Sophistication is justified only when the response genuinely cannot be represented by a simpler form and the data to support the complex form is available.

BUILDING A NEURAL-NETWORK SURROGATE WHEN A SIMPLE POLYNOMIAL RESPONSE SURFACE ALREADY DESCRIBES THE PROBLEM ADEQUATELY ADDS COMPLEXITY WITHOUT ENGINEERING VALUE.