Langford Analytic · Knowledge Base

Performance, Vectorisation, Parallelism & HPC

The fastest code is not necessarily the best engineering code if it becomes difficult to verify or maintain. From algorithm improvement to vectorisation, parallelism and HPC, performance should serve the engineering objective — not the other way round.

Article 18Engineering Software & Data12 min read
performancevectorisationparallel computingHPCoptimisation

The Engineering Problem

Some engineering computations are fast enough to run on a laptop in seconds. Others — large FEA models, CFD simulations, Monte Carlo campaigns, optimisation loops — require significant computational resources. Understanding how to improve performance, when to do so and what the trade-offs are is part of computational engineering.

Performance Hierarchy

The most effective performance improvement is almost always a better algorithm. Before optimising the hardware or the implementation, ensure the algorithm itself is efficient. Only then should implementation-level optimisations — vectorisation, compiled kernels, parallelism — be considered.

Algorithm improvement → Vectorisation → Compiled code → Parallel execution → HPC

Optimise the algorithm before optimising the hardware where practical.

Algorithm Before Hardware

A poor algorithm on the fastest hardware will often be slower than a good algorithm on modest hardware. Replacing an O(n²) algorithm with an O(n log n) algorithm can provide orders of magnitude improvement. No amount of parallelism or vectorisation compensates for a fundamentally inefficient algorithm.

  • Choose the right algorithm first — complexity class matters more than constant factors
  • Exploit problem structure — sparsity, symmetry, periodicity
  • Use established, optimised libraries — BLAS, LAPACK, FFTW
  • Profile before optimising — identify where time is actually spent

Vectorisation

Vectorisation replaces element-by-element operations with whole-array operations that use hardware-level SIMD (Single Instruction, Multiple Data) capabilities. In Python, this means using NumPy array operations instead of Python loops. In MATLAB, it means using matrix operations instead of element-wise loops. In Fortran and C++, it means array syntax or compiler auto-vectorisation.

# Python: element-by-element (slow)
for i in range(n):
    c[i] = a[i] + b[i]

# NumPy: vectorised (fast)
c = a + b   # operates on entire array at once

Compiled Kernels

For performance-critical sections that cannot be vectorised in a high-level language, compiled kernels — written in C, C++, Fortran or Cython — provide the performance of compiled execution. NumPy and SciPy already use compiled kernels internally. For custom computations, tools like Cython, Numba or direct C/Fortran extensions can provide compiled performance from Python code.

  • NumPy and SciPy already use compiled BLAS and LAPACK internally
  • Numba — JIT compilation of Python numerical code
  • Cython — Python-like syntax that compiles to C
  • Direct C/C++/Fortran extensions for maximum control

Parallel Computing

Parallel computing uses multiple processing units simultaneously. The appropriate form of parallelism depends on the problem and the hardware available.

Parallelism TypeHow It WorksEngineering Use
Multi-threadingShared memory, multiple threads in one processVectorised libraries, FEA solver internal parallelism
Multi-processingMultiple processes, separate memoryIndependent analyses — parametric studies, Monte Carlo
Distributed computingMultiple machines connected by networkLarge FEA/CFD models — MPI-based solvers
GPU computingMassively parallel on graphics hardwareDense linear algebra, CFD, ML training

Serial Fraction and Amdahl’s Law

The speedup from parallelism is limited by the serial fraction — the part of the computation that cannot be parallelised. Amdahl’s law states that if 10% of the computation is serial, the maximum speedup is 10× regardless of how many processors are used. Understanding the serial fraction helps set realistic expectations for parallel performance.

Amdahl's Law:

Speedup  =  1 / ( s + (1 − s) / p )

where:
s  =  serial fraction
p  =  number of processors

Communication and I/O

In parallel computing, communication between processes and I/O can dominate the runtime. A computation that is fast in serial may scale poorly in parallel if processes spend most of their time communicating rather than computing. Understanding the balance between computation and communication is essential for effective parallelism.

  • Communication overhead — data transfer between processes
  • Synchronisation — waiting for all processes to reach a point
  • Load balancing — ensuring all processors have similar workloads
  • I/O bottleneck — reading and writing large datasets

High-Performance Computing (HPC)

HPC refers to the use of large, parallel computing facilities — clusters, supercomputers — for demanding engineering simulations. HPC enables larger models, finer meshes and more extensive parametric studies than are possible on workstations. However, HPC does not change the engineering — it changes the scale of what can be computed. The same verification, traceability and engineering review requirements apply.

The fastest code is not necessarily the best engineering code if it becomes difficult to verify or maintain.

When Performance Matters

Performance optimisation is appropriate when the computation time is a practical constraint on the engineering work. If a calculation runs in seconds, optimisation is unnecessary. If a parametric study takes days and the engineer is waiting, optimisation is justified. If the same calculation is run thousands of times in a Monte Carlo campaign, performance improvements multiply in value.

  • Performance matters when computation time constrains the engineering work
  • Optimise the parts that are actually slow — profile first
  • Do not sacrifice readability for performance unless the gain is significant
  • Re-verify after optimisation — a faster but incorrect result is worthless

When Not to Optimise

Premature optimisation — making code faster before knowing whether speed matters — wastes engineering time and often produces code that is harder to verify, harder to maintain and harder for other engineers to understand. The simplest correct implementation should be the starting point. Optimisation should be targeted and based on measured performance, not speculation.

ENGINEERING CHECK: After optimisation, does the code still produce verified results? A faster but incorrect result is worthless. Re-run the full verification suite after any performance change.

Key Takeaways

  • Optimise the algorithm before optimising the hardware
  • Vectorisation replaces slow element-by-element loops with fast whole-array operations
  • Parallel speedup is limited by the serial fraction — Amdahl’s law
  • HPC changes the scale of computation, not the engineering requirements
  • Do not sacrifice verifiability or maintainability for performance unless justified