High-Performance Computing & Large-Scale Simulation
High-performance computing makes large-scale simulation possible, but more processors do not guarantee proportionally faster engineering. This article covers parallelism, strong and weak scaling, model decomposition, I/O bottlenecks, the types of analysis that benefit from HPC, and the scaling tests that must be performed before assuming that additional compute will deliver additional engineering value.
What HPC Is — Conceptually
High-performance computing is the use of parallel computational resources — many processors, many cores, large memory, specialised accelerators — to solve problems that are too large or too expensive for a single workstation. The premise is straightforward: if a problem can be divided into parts that are solved concurrently, the wall-clock time to solution can be reduced by using more processors. The reality is more complex. The speed-up from parallelism is limited by the fraction of the problem that cannot be parallelised, by the communication overhead between processors, by memory bandwidth, by I/O bottlenecks and by the architecture of the solver. HPC is not a guarantee of faster results; it is an opportunity for faster results that must be assessed and exploited with understanding. The components of an HPC system include: CPUs with multiple cores, which execute instructions concurrently; distributed memory across multiple nodes, which allows larger models to be solved; GPUs, which provide massive parallelism for suitable workloads; batch scheduling systems, which manage the allocation of compute resources to jobs; and high-speed interconnects, which reduce the communication overhead between nodes. No specific vendor or architecture is promoted here; the suitability of a given platform depends on the solver, the problem and the organisation's computing infrastructure.
Strong Scaling vs Weak Scaling
The two fundamental measures of parallel performance are strong scaling and weak scaling, and they answer different questions. Strong scaling asks: for a fixed problem size, how does the wall-clock time decrease as more processors are added? If a problem that takes 100 hours on 1 core takes 50 hours on 2 cores, 25 hours on 4 cores and so on, the scaling is ideal — the speed-up is proportional to the processor count. In practice, strong scaling always departs from ideal at some processor count, because the communication overhead between processors grows as the per-processor workload shrinks, and the fraction of the work that is inherently serial becomes a larger fraction of the total. Strong scaling tells the engineer how many processors are worth using for a given problem size. Weak scaling asks: for a fixed problem size per processor, how does the wall-clock time change as the problem size and processor count are increased together? If a problem of size N on 1 core takes the same time as a problem of size 2N on 2 cores, the weak scaling is ideal. Weak scaling tells the engineer whether the solver can handle larger problems at the same per-problem cost by using more processors. Both measures are problem- and solver-dependent; neither can be assumed. They must be tested.
MORE PROCESSORS DO NOT GUARANTEE PROPORTIONALLY FASTER ENGINEERING. Parallel speed-up is limited by communication overhead, memory bandwidth, I/O, serial fractions and solver architecture. A problem that scales ideally to 16 cores may show diminishing returns at 64 and no benefit at 256. The scaling behaviour must be tested for the specific problem and solver before committing computational resources.
Model Decomposition and I/O
Parallel solvers decompose the model into domains — one per processor or core — and each processor solves its domain while exchanging boundary information with the processors that hold adjacent domains. The decomposition strategy affects both the computational load balance and the communication overhead: a decomposition that gives each processor an equal amount of work but creates long boundaries between domains balances computation but increases communication. The efficiency of the solver on a parallel platform depends on the quality of the decomposition, the ratio of computation to communication and the bandwidth and latency of the interconnect. I/O is a frequently overlooked bottleneck. A large model may generate gigabytes or terabytes of output — field variables at every node, time-history data, restart files. If the I/O system cannot write this data as fast as the solver produces it, the solver spends time waiting for I/O, and the parallel speed-up is lost. I/O bottlenecks are particularly severe for transient and explicit dynamics analyses, which may write results at every timestep. The engineer must understand where the I/O bottleneck lies — disk bandwidth, network filesystem, serial output — and address it if the parallel efficiency is to be realised. Not all analyses scale efficiently: a small linear static model may run faster on a single core than on a cluster, because the communication overhead exceeds the computation. The problem must be large enough to benefit from parallelism.
Types of Analysis That Benefit from HPC
HPC is most valuable for analyses that are computationally expensive either because the model is large, the physics is complex, or many evaluations are needed. The following types of analysis are typical HPC beneficiaries. In each case, the value of HPC depends on the specific problem — a small CFD model may not need HPC, and a large linear static model may scale poorly — but these categories are where HPC most often delivers a step change in capability.
- CFD: high-fidelity turbulence-resolving simulations — large eddy simulation, detached eddy simulation — require meshes of millions to billions of cells and many thousands of timesteps.
- Explicit dynamics: crash, impact and blast simulations require very small timesteps and many millions of elements, making them inherently expensive.
- Non-linear FEA: contact, plasticity and large-deformation analyses require iterative solutions that are computationally intensive per step.
- Optimisation: gradient-based or heuristic optimisers require many function evaluations, each of which may be a full simulation.
- Monte Carlo and reliability analysis: statistical sampling requires thousands or millions of evaluations, each of which is a simulation.
- Parameter studies: design of experiments, sensitivity analysis and multi-fidelity campaigns require many model evaluations across a parameter space.
Computing Architecture and Scaling Behaviour
The diagram below illustrates the conceptual difference between computing on a single workstation, a parallel cluster and a GPU-enabled platform, followed by a scaling curve that shows the characteristic diminishing returns of parallel speed-up. The scaling curve is conceptual — the actual scaling depends on the solver, the problem, the mesh, the decomposition and the interconnect — but the shape is typical: speed-up is near-linear at low processor counts, departs from linear as communication overhead grows, and may plateau or decrease at high processor counts where the per-processor workload is too small to amortise the communication cost. No specific vendor, architecture or scaling factor is promoted; the diagram illustrates the principle, not a measured result.
COMPUTING ARCHITECTURES — CONCEPTUAL
SINGLE WORKSTATION PARALLEL CLUSTER GPU-ENABLED
┌────────────────┐ ┌────────┐ ┌────────┐ ┌────────────────┐
│ Multi-core │ │ Node 1│ │ Node 2│ ... │ CPU + GPU │
│ CPU │ │ cores │ │ cores │ │ CPU manages │
│ │ │ memory│ │ memory│ │ GPU executes │
│ Shared memory │ └────────┘ └────────┘ │ massively │
│ Limited by │ Connected by high-speed │ parallel │
│ core count │ interconnect │ kernels │
│ and RAM │ └────────────────┘
└────────────────┘ Distributed memory Suitable for: dense
Suitable for: Suitable for: large linear algebra, FFT,
moderate models FE, CFD, explicit image processing
single-domain dynamics, optimisation Not all solvers
analyses loops, Monte Carlo support GPU
SCALING CURVE — PROCESSORS vs RUNTIME (CONCEPTUAL)
Runtime
(log)
│
│ ● Ideal (1/N)
│ ●
│ ●
│ ●
│ ● ← Near-linear region: Actual scaling
│ ● communication overhead ●
│ ● is small relative to ●
│ ● computation ●
│ ● ● ← Diminishing returns:
│ ● ● communication overhead
│ ● ● dominates; per-core
│ ● ● workload too small
│ ● ●
│ ● ● ● ● ● ● ● ● ● ● ● ● ← Plateau / decrease
│ adding cores does not
└─────────────────────────────────────────────────────── help (or hurts)
Processors (N)
The point of diminishing returns depends on:
• Problem size (larger problems scale further)
• Solver architecture (domain decomposition, communication pattern)
• Interconnect bandwidth and latency
• Memory per node
• I/O bandwidth
The scaling must be TESTED for the specific problem and solver.HPC Engineering Checklist
Before committing an analysis to HPC resources, the following questions should be answered. They address the two key risks of HPC: that the problem does not scale as hoped, and that the computational resources are consumed without delivering proportionate engineering value. HPC is not free — even when the compute is available, the engineer's time to set up, debug and manage the parallel run is a cost, and the results must justify that cost.
- Is the problem parallelisable? — Some problems are inherently serial or have large serial fractions. Understand the solver's parallel structure before assuming benefit.
- Is the memory sufficient? — A model that fits in RAM on a large shared-memory node may not fit when decomposed across distributed-memory nodes. Check per-node memory.
- Does the solver support the target architecture? — Not all solvers support GPU acceleration, and not all solvers scale efficiently on distributed clusters. Confirm the solver's capability.
- Is the I/O bottleneck understood? — Large transient analyses may be I/O-bound. Identify whether disk, network filesystem or serial output is the bottleneck.
- Has strong or weak scaling been tested? — Run the problem at increasing processor counts and measure the speed-up. Do not assume scaling — test it.
- Is the numerical result unchanged with processor count? — For some solvers, the decomposition can affect the numerical result (e.g. due to rounding in parallel reductions). Verify that the result is consistent across processor counts where this is expected.
- Is automated job failure detection present? — On a cluster, jobs may fail due to node faults, queue limits or resource exhaustion. The pipeline must detect and flag failed jobs.
- Is result reproducibility checked? — A job run twice on the same cluster should produce the same result. If it does not, investigate the cause before trusting either result.
- Is the cost of additional compute justified by additional engineering information? — More processors may produce a faster result, but if the engineering decision does not require the additional fidelity or scale, the compute is wasted.
HPC Considerations
The table below summarises the key HPC considerations, what each affects, how to assess it and the common misconception associated with it. The considerations are interdependent — memory affects decomposition, I/O affects scaling, solver architecture affects GPU suitability — and the engineer must consider them together, not in isolation.
| Consideration | What it affects | How to assess | Common misconception |
|---|---|---|---|
| Parallelism type (shared / distributed / GPU) | Which problems can be run and how efficiently | Check solver documentation; run benchmark on target architecture | Assuming all solvers support all parallelism types equally |
| Memory (per-node and total) | Maximum model size; decomposition strategy | Estimate model memory; check per-node RAM; test decomposition | Assuming total cluster memory equals usable model memory |
| I/O (disk, network filesystem, output format) | Wall-clock time for I/O-heavy analyses; restart capability | Measure I/O time as fraction of total; test parallel I/O if available | Ignoring I/O because the compute scales — I/O can dominate |
| Scaling character (strong / weak) | How many processors are worth using; how large a problem is feasible | Run scaling study at increasing processor counts; measure speed-up | Assuming linear scaling to any processor count |
| Solver architecture (domain decomposition, MPI, OpenMP, GPU) | Parallel efficiency; maximum practical problem size | Review solver parallel design; benchmark on representative problem | Assuming the solver scales the same on all problems |
| GPU suitability | Whether GPU acceleration provides benefit | Check if solver has GPU-enabled routines; benchmark GPU vs CPU | Assuming GPUs accelerate everything — they accelerate specific workloads |
| Batch scheduling | Job queue time, resource allocation, wall-clock limits | Review scheduler policies; estimate queue time; plan for limits | Assuming jobs start immediately — queue time can exceed run time |
| Cost (compute hours, storage, engineer time) | Whether HPC use is justified for the engineering question | Estimate compute hours; compare against engineering value of the result | Assuming more compute is always better — compute is a cost to be justified |
The Most Common HPC Mistake
The most common mistake in HPC use is assuming that running a model on more processors will automatically produce results faster, without testing the scaling behaviour. This assumption leads to two failures. The first is the waste of computational resources: a problem that has already reached its scaling limit is given more processors, and the additional compute produces no reduction in wall-clock time — the resources are consumed without benefit. The second is more subtle: a problem that runs on more processors but produces a different numerical result due to parallel rounding, decomposition effects or solver-specific parallel behaviour, and the engineer does not notice because the result "looks the same". The remedy is to perform a scaling study — run the problem at increasing processor counts, measure the speed-up, identify the point of diminishing returns and verify that the numerical result is consistent across processor counts — before committing to a large parallel campaign. The scaling study costs a modest amount of compute; the cost of running an entire campaign at a processor count that does not scale, or that produces numerically different results, is far greater.
ASSUMING THAT RUNNING A MODEL ON MORE PROCESSORS WILL AUTOMATICALLY PRODUCE RESULTS FASTER WITHOUT TESTING SCALING BEHAVIOUR CAN LEAD TO EXPENSIVE COMPUTATIONAL RESOURCES PRODUCING DIMINISHING ENGINEERING RETURNS. The scaling must be tested for the specific problem and solver. The point of diminishing returns must be identified. The numerical consistency of the result across processor counts must be verified. HPC is an opportunity, not a guarantee.
Key Takeaways
- HPC uses parallel resources — many cores, distributed memory, GPUs — to solve problems too large or expensive for a workstation
- Strong scaling: fixed problem, more processors → faster. Weak scaling: larger problem, more processors → same time
- Both scaling types must be tested; neither can be assumed
- Model decomposition, communication overhead, memory bandwidth and I/O all limit parallel efficiency
- Not all analyses scale efficiently — small problems may run faster on a single core
- HPC benefits CFD, explicit dynamics, non-linear FEA, optimisation, Monte Carlo and parameter studies
- A scaling study should be performed before committing to a large parallel campaign
- Numerical consistency across processor counts must be verified where expected
- The cost of additional compute must be justified by additional engineering information