Langford Analytic · Knowledge Base

Machine Learning in Engineering Analysis

Machine learning can approximate relationships in data, but approximating data is not the same as learning physics. This article covers where ML can genuinely help in engineering analysis, where it cannot, the distinction between physics-based and data-driven models, the interpretability required for high-consequence decisions, and why validating a model on the same dataset it was trained on demonstrates interpolation, not physical validation.

Article 13Automation & Computational Scale15 min read
machine-learningsurrogate-modellingdata-drivenphysics-informedblack-boxinterpretabilityout-of-distributiondata-leakageanomaly-detectionvalidation

A Balanced View

Machine learning has a legitimate and growing role in engineering analysis, but the role is specific and conditional, not universal. Machine learning can build fast surrogates from simulation data, classify large datasets, detect anomalies in sensor streams, interpret non-destructive testing images, fuse multi-sensor data and solve difficult inverse mappings. These are real capabilities, and where the data supports them and the decision tolerates the method's limitations, they can deliver engineering value. But machine learning is not a substitute for physical understanding. A model that approximates the relationship in a dataset has learned a statistical mapping, not the governing equations. It can interpolate within the training data; it cannot reliably extrapolate beyond it. It can be fast; it cannot be assumed to be physically correct. This article takes a technically sceptical position: machine learning should be used where it genuinely helps, understood where it does not, and never used to replace a well-understood physical equation merely because the black box sounds more advanced. The central principle is that advanced methods are valuable when they reduce engineering uncertainty or computational cost without losing the physics required by the decision. For machine learning, that means the method must be validated against independent evidence, its domain of applicability must be known, and its failure modes must be identifiable.

MACHINE LEARNING CAN APPROXIMATE RELATIONSHIPS IN DATA. THAT DOES NOT AUTOMATICALLY MEAN IT HAS LEARNED THE UNDERLYING PHYSICS. A data-driven model that fits the training data has learned a statistical mapping — a correlation, a pattern — within the region of the input space covered by the training data. Whether that mapping reflects the governing physics outside the training data is a separate question, and the answer is not guaranteed to be yes.

Supervised and Unsupervised Learning — At a High Level

Machine learning methods are broadly categorised by the type of data they use and the type of output they produce. Supervised learning uses labelled training data — input–output pairs where the correct output is known — to learn a mapping from inputs to outputs. A surrogate model trained on simulation results is a supervised model: the simulation provides the input–output pairs, and the model learns to predict the output for new inputs. Unsupervised learning uses unlabelled data to discover structure — clusters, patterns, anomalies — without a predefined output. An anomaly-detection model trained on sensor streams is an unsupervised model: it learns the normal pattern of sensor data and flags deviations. The distinction matters because the validation requirements are different. A supervised model is validated by comparing its predictions against held-out test data — inputs and outputs that were not used in training. An unsupervised model is validated by examining whether the discovered structure is physically meaningful and whether the flagged anomalies correspond to real engineering events. In both cases, the quality of the model depends on the quality, coverage and representativeness of the training data. If the training data does not cover the region of the input space where the model will be used, the model is being used out of distribution, and its predictions are unreliable.

Training Data, Extrapolation and Feature Distribution

The behaviour of a machine learning model is determined by the data it was trained on. The training data defines a region of the input space — the distribution of features, the range of each parameter, the combinations of parameters that appear — within which the model has been taught to predict. Inside this region, the model interpolates: it predicts outputs for inputs that lie between training points, and if the underlying function is smooth and the training data is dense, the interpolation may be accurate. Outside this region, the model extrapolates: it predicts outputs for inputs that lie beyond the training data, and the prediction is unconstrained by data. Different model types extrapolate differently — some flat, some linear, some divergent — but none can be assumed to extrapolate correctly, because the data that would constrain the prediction in that region does not exist. The feature distribution of the training data — which regions of the input space are densely sampled and which are sparse — determines where the model can be trusted. A model trained on simulation data from one configuration cannot be assumed to generalise to a different configuration. A model trained on data from one operating regime cannot be assumed to generalise to a different regime. The engineer must know the training data distribution, must know where the model is being used relative to that distribution, and must treat out-of-distribution predictions as unreliable until validated.

NEVER REPLACE A WELL-UNDERSTOOD PHYSICAL EQUATION WITH A BLACK-BOX MODEL MERELY BECAUSE THE BLACK BOX SOUNDS MORE ADVANCED. If the physical equation is known, validated and understood, it provides predictions with known assumptions, known limitations and traceable reasoning. A black-box model that replaces it provides predictions without traceable reasoning, without known assumptions and without known limitations outside the training data. The black box is justified only when it offers a genuine advantage — speed, dimensionality, a problem the equation cannot solve — that outweighs the loss of traceability.

Physics-Based, Data-Driven and Hybrid Approaches

The diagram below compares three modelling approaches at a conceptual level: a physics-based model, which uses governing equations to predict response from inputs; a data-driven model, which uses training examples to learn a statistical relationship; and a hybrid or physics-informed approach, which combines physical constraints with data. The diagram is deliberately restrained: it does not imply that the hybrid approach is automatically superior. Each approach has domains where it is appropriate and domains where it is not. The choice depends on the problem, the available data, the required accuracy, the consequence of error and the need for interpretability.

THREE MODELLING APPROACHES — CONCEPTUAL COMPARISON

  PHYSICS-BASED MODEL
  ┌──────────┐    ┌──────────────────┐    ┌──────────┐
  │  Inputs  │───▶│  Governing       │───▶│ Response │
  │          │    │  equations       │    │          │
  │ Geometry │    │  (FE, CFD,       │    │ Stress,  │
  │ Loads    │    │   thermal,       │    │ temp,    │
  │ Material │    │   dynamics)      │    │ force    │
  └──────────┘    └──────────────────┘    └──────────┘
  • Mechanism: known physics, explicit equations
  • Extrapolation: governed by equations (within their validity)
  • Interpretability: high — each step is physical
  • Data requirement: material properties, BCs — not training data

  DATA-DRIVEN MODEL
  ┌──────────┐    ┌──────────────────┐    ┌──────────┐
  │  Input   │───▶│  Learned          │───▶│ Response │
  │  examples│    │  statistical      │    │ estimate │
  │          │    │  relationship     │    │          │
  │ Training │    │  (neural net,     │    │ Predicted│
  │ data     │    │   Gaussian        │    │ output   │
  │ pairs    │    │   process, etc.)  │    │          │
  └──────────┘    └──────────────────┘    └──────────┘
  • Mechanism: statistical mapping from training data
  • Extrapolation: unconstrained beyond training data — unreliable
  • Interpretability: low (black box) to moderate
  • Data requirement: large, representative training set

  HYBRID / PHYSICS-INFORMED APPROACH
  ┌──────────┐    ┌──────────────────┐    ┌──────────┐
  │  Inputs  │───▶│  Physics          │───▶│ Response │
  │  +       │    │  constraints      │    │          │
  │  Data    │    │  +                │    │ Physics- │
  │          │    │  data-driven      │    │ informed │
  │          │    │  correction       │    │ estimate │
  └──────────┘    └──────────────────┘    └──────────┘
  • Mechanism: physics provides structure; data provides correction
  • Extrapolation: constrained by physics, corrected by data
  • Interpretability: moderate — physics is traceable, correction may not be
  • Data requirement: physics model + data for correction

  NONE OF THESE IS AUTOMATICALLY SUPERIOR. The choice depends on the
  problem, the data, the consequence and the need for traceability.

When Machine Learning May Be Useful

Machine learning is not universally useful, but there are specific engineering tasks where it can provide genuine value. The common characteristic of these tasks is that the physical model is too expensive to run at the required scale, or the problem is one of pattern recognition that a physical model does not naturally address. In each case below, the value of ML is conditional on the availability of representative training data and on the decision being tolerant of the method's limitations.

  • Rapid surrogate: a model trained on simulation data can predict the response in milliseconds where the simulation takes hours, enabling optimisation and uncertainty quantification at scale.
  • Anomaly detection: a model trained on normal sensor or operational data can flag deviations that may indicate damage, malfunction or changing conditions — triggering investigation rather than providing a diagnosis.
  • Large-image or data classification: a model can classify non-destructive testing images, satellite imagery or inspection data at a scale and speed that manual review cannot match.
  • Sensor fusion: a model can combine data from multiple sensors — strain, acceleration, temperature, pressure — into a coherent estimate of the system state that no single sensor provides.
  • Feature recognition: a model can identify patterns in high-dimensional data — modal parameters, frequency content, spatial distributions — that are difficult to extract by manual inspection.
  • Difficult inverse mapping: a model can learn the inverse mapping from response to parameters — strain to load, vibration to damage — that is computationally expensive or ill-posed by conventional methods.

When Machine Learning May Not Be Useful

Equally important is recognising where machine learning does not help or is actively misleading. The common characteristic of these situations is that the conditions for reliable ML — representative data, in-distribution operation, tolerance of opacity — are not met. Using ML in these situations does not produce an advanced analysis; it produces an unreliable prediction dressed in the language of sophistication.

  • Sparse data: if the training data is too few, too narrow or too noisy, the learned mapping is unreliable. No model type compensates for insufficient data.
  • Poor coverage: if the training data does not cover the region of the input space where the model will be used, the model extrapolates and the predictions are unconstrained.
  • Safety-critical extrapolation: if the decision depends on predictions outside the training data — a structural assessment at a load case not represented in training — the model is unreliable and the consequence of error is high.
  • Simple well-understood equation already exists: if the physical equation is known, cheap and validated, replacing it with a black box loses traceability without gaining anything.
  • Training and test configurations differ materially: if the model is trained on one configuration and applied to a different one — different geometry, different material, different boundary conditions — the model is out of distribution and the predictions are unreliable.

Failure Cases

Machine learning models fail in characteristic ways, and the engineer must be aware of these failure modes before deploying a model in an engineering context. Out-of-distribution input: the model is given an input that lies outside the training data distribution, and it produces a prediction that is unconstrained by data — a prediction that may be physically implausible but that the model has no mechanism to flag as unreliable. Sparse training region: the model is given an input that lies within the training data bounds but in a region where the training data is sparse, and the prediction is unreliable because the local density of training points is insufficient. Biased test data: the test data used to validate the model is drawn from the same distribution as the training data, so the validation demonstrates interpolation within that distribution, not generalisation to new conditions. Data leakage: information from the test set leaks into the training process — through feature selection, preprocessing or cross-validation design — and the reported accuracy is inflated because the model has indirectly seen the test data. Confounding parameters: the model learns a correlation between an input and an output that is actually driven by a third, unmeasured parameter, and the learned mapping breaks when the confound changes. These failure modes are not rare; they are the dominant causes of ML models that perform well in development and fail in deployment.

The Black-Box Problem and Interpretability

A machine learning model — particularly a deep neural network — is a black box in the sense that the mapping from input to output is encoded in thousands or millions of parameters that do not correspond to identifiable physical quantities. The model produces a prediction, but it does not explain why. This does not mean that every ML model must be fully interpretable; the level of interpretability required depends on the decision the model supports. For a low-consequence decision — flagging an image for manual review, screening a dataset for anomalies — a black box may be acceptable because the human review provides the interpretability. For a high-consequence decision — a structural margin, a certification result, a safety assessment — the opacity of the model is a serious concern, because the engineer cannot trace the prediction to its physical basis. For high-consequence decisions, the engineer must be able to answer: can the failure modes of the model be identified? Is extrapolation — operation outside the training data — detectable? Can the sensitivity of the prediction to each input be understood? Is the uncertainty in the prediction quantified? Can the result be independently checked against a physics-based model or test data? If the answer to these questions is "no", the engineer must decide whether other evidence is sufficient to trust the model's prediction, or whether the prediction must be treated as one input among several rather than as the basis for the decision.

IF THE MODEL CANNOT EXPLAIN WHY IT PRODUCED AN ANSWER, THE ENGINEER MUST DECIDE WHETHER OTHER EVIDENCE IS SUFFICIENT TO TRUST THAT ANSWER. A black-box prediction is not engineering evidence by itself; it is a claim that requires corroboration. For high-consequence decisions, the prediction must be checked against a physics-based model, test data or an independent method. The consequence of the decision determines how much opacity is tolerable.

Machine Learning Engineering Checklist

Before a machine learning model is used to support an engineering decision, the following questions should be answered. The checklist is not a formality; each question addresses a specific failure mode that can cause the model to produce confident but incorrect predictions. If any question cannot be answered satisfactorily, the model is not ready to support the decision.

  • What physical problem is the ML model solving? — State the engineering question, not the ML task. "Predict stress at the fillet" is the question; "regression on a neural network" is the method.
  • Why is conventional physics insufficient or inefficient? — If a physics-based model can answer the question adequately, the ML model needs a specific advantage — speed, scale, dimensionality — to justify the loss of traceability.
  • Is the training data representative of the application conditions? — The training data must cover the region of the input space where the model will be used. Representativeness is not about size alone; it is about coverage.
  • Is the validation data independent of the training data? — Validation data must be drawn from conditions not used in training. Validation on the training data demonstrates interpolation, not generalisation.
  • Has data leakage been prevented? — Ensure no information from the validation or test set has influenced training — through preprocessing, feature selection or cross-validation design.
  • Can out-of-distribution input be detected? — The model or the pipeline must flag inputs that lie outside the training data distribution, so that out-of-distribution predictions are not presented as reliable.
  • Is the prediction uncertainty estimated? — A point prediction without uncertainty is incomplete. The model should provide an estimate of confidence, or the pipeline should provide one by other means.
  • Are the important features physically plausible? — If the model's sensitivity to inputs contradicts physical understanding — it predicts that stress decreases with load — the model has learned a spurious correlation.
  • Have the failure cases been identified? — Test the model on edge cases, boundary cases and physically extreme cases. Identify where it fails and whether those failures are detectable.
  • Are predictions independently checked? — For high-consequence decisions, ML predictions should be checked against a physics-based model, test data or an independent method.
  • Does the consequence of the decision justify the method's opacity? — A high-consequence decision requires higher interpretability. If the model cannot explain its answer and the consequence is high, other evidence must be sufficient.

Physics Model vs Data-Driven Model vs Hybrid

The table below compares the three modelling approaches across the dimensions that matter for engineering decisions: the input mechanism, the extrapolation behaviour, the interpretability, the data requirements, the treatment of physical constraints, when each is appropriate and the primary risk. The comparison is not a ranking; it is a guide to selecting the approach that fits the problem.

DimensionPhysics-based modelData-driven (ML) modelHybrid / physics-informed
Input mechanismGoverning equations solved numerically (FE, CFD, thermal, dynamics)Statistical mapping learned from training input–output pairsPhysics equations provide structure; data provides correction or residual learning
Extrapolation behaviourGoverned by the equations — extrapolates according to the physics within the equations' validityUnconstrained beyond training data — extrapolation behaviour depends on model type and is generally unreliableConstrained by the physics; correction extrapolates with the physics but may degrade if the correction model extrapolates
InterpretabilityHigh — each step is physical and traceable to governing equationsLow to moderate — prediction is a function of learned weights; sensitivity analysis may provide partial insightModerate — physics is traceable; the data-driven correction may be less so
Data requirementsMaterial properties, geometry, loads, boundary conditions — not training dataLarge, representative training dataset covering the input space; cost depends on data sourcePhysics model inputs plus data for the correction; typically less data than pure data-driven
Physical constraintsEnforced by the equations — equilibrium, conservation, constitutive lawNot enforced unless explicitly imposed as constraints or penaltiesEnforced by the physics component; data component may violate if not constrained
When appropriateWhen the physics is known, the model is tractable and the cost is acceptableWhen the physics model is too expensive for the required scale, or the problem is pattern recognitionWhen the physics captures most of the behaviour but a correction is needed for accuracy or speed
Primary riskComputational cost; model-form error if the physics is incompleteOut-of-distribution failure; spurious correlations; opacity for high-consequence decisionsCorrection model may extrapolate; complexity of two coupled components; validation of the correction

The Most Common ML Validation Mistake

The most common mistake in applying machine learning to engineering analysis is training a model on simulation data and then validating it on the same simulation dataset — or on a split of the same dataset drawn from the same distribution — and presenting the validation accuracy as evidence that the model is physically valid. This is not physical validation; it is interpolation within the training distribution. The model has demonstrated that it can approximate the simulation data it was trained on and tested on, which is the minimum requirement for a surrogate, not evidence that it captures the underlying physics. True physical validation requires the model to predict the response for conditions that were not used in training and that differ from the training conditions in a physically meaningful way — a different geometry, a different load case, a different material, or a comparison against experimental test data. If the model generalises to these independent conditions, there is evidence that it has learned something physically meaningful. If it only performs well on held-out data from the same distribution, it has learned to interpolate the simulation data — which may be valuable as a surrogate, but should not be presented as physical validation.

TRAINING A MACHINE-LEARNING MODEL ON SIMULATION DATA AND THEN VALIDATING IT ON THE SAME SIMULATION DATASET WITHOUT INDEPENDENT TEST EVIDENCE DEMONSTRATES INTERPOLATION, NOT PHYSICAL VALIDATION. The model has learned to approximate the simulation data within its distribution. Whether it has learned the physics — whether it generalises to conditions not represented in the training data — is a separate question that requires independent evidence. Surrogate accuracy within the training distribution is necessary but not sufficient for engineering use.

Key Takeaways

  • Machine learning can approximate relationships in data — that does not mean it has learned the underlying physics
  • ML is useful for surrogates, anomaly detection, classification, sensor fusion, feature recognition and inverse mapping — where data is representative and the decision tolerates the limitations
  • ML is not useful with sparse data, poor coverage, safety-critical extrapolation, when a simple equation exists, or when training and application configurations differ materially
  • Never replace a well-understood physical equation with a black box merely because the black box sounds more advanced
  • The level of interpretability required depends on the consequence of the decision — high consequence requires higher interpretability or independent corroboration
  • Out-of-distribution input, sparse training regions, data leakage, biased test data and confounding parameters are the dominant failure modes
  • Validation on the same training distribution demonstrates interpolation, not physical validation — independent evidence is required
  • The ML engineering checklist should be answered before any model supports an engineering decision