Verification, Testing & Debugging of Engineering Code
How analytical benchmarks, unit tests, limit cases and regression testing convert working software into credible engineering software. Software testing asks whether the code works as written. Engineering verification asks whether the code represents the intended mathematics and physics.
The Engineering Problem
The most dangerous engineering software error is often not a crash — it is a plausible-looking incorrect answer. A program that runs without errors, produces well-formatted output and gives a result that looks reasonable can still be wrong. Engineering code requires a different and more rigorous approach to verification than ordinary software, because the consequence of an undetected error can be an incorrect engineering decision.
Distinguishing Error Types
Engineering computational errors come in several distinct types. Confusing them leads to incorrect remediation. A software bug is not the same as a numerical error, which is not the same as a model error.
The most dangerous engineering software error is often not a crash — it is a plausible-looking incorrect answer.
| Error Type | What It Is | Example |
|---|---|---|
| Software bug | The code does not do what the programmer intended | Array index off by one; wrong variable used |
| Numerical error | The code does what was intended but the method introduces approximation | Truncation error in finite difference; round-off in summation |
| Model error | The mathematics is implemented correctly but does not represent the physics | Euler–Bernoulli beam used where shear deformation is significant |
| Input error | The code and model are correct but the inputs are wrong | Load in kN entered where N was expected |
| Interpretation error | The result is correct but the engineering conclusion is wrong | Von Mises stress compared against shear allowable |
The Verification Hierarchy
Engineering code verification is not a single activity. It is a hierarchy of checks, each addressing a different level of confidence. Moving up the hierarchy increases confidence that the code is not just running but producing credible engineering results.
Level 1: Syntax/Execution → Level 2: Software Test → Level 3: Numerical Benchmark → Level 4: Engineering Verification → Level 5: Validation
Level 1 — Syntax / Execution
The lowest level of confidence. Does the code run without crashing? This is necessary but nowhere near sufficient. A program that runs may still produce completely incorrect results.
"It runs" is the lowest level of confidence.
Level 2 — Software Test
Unit tests check that individual functions behave as intended. Each function should be tested with known inputs and expected outputs. Unit tests catch software bugs — cases where the code does not do what the programmer intended.
- Test each function with known inputs and expected outputs
- Test edge cases — zero input, empty array, single element
- Test error handling — what happens with invalid input?
- Use a test framework (pytest, MATLAB unit testing, Google Test)
- Run tests automatically on every code change
Level 3 — Numerical Benchmark
Benchmark tests compare code output against known analytical solutions. A function that computes beam deflection should be tested against the analytical Euler–Bernoulli solution. A linear solver should be tested against a small system with a known solution. Benchmark tests catch numerical errors — cases where the method introduces unacceptable approximation.
- Compare against closed-form analytical solutions where available
- Compare against published benchmark results from the literature
- Compare against results from a different, independent implementation
- Test at multiple discretisation levels and check convergence
Level 4 — Engineering Verification
Engineering verification asks whether the code implements the intended engineering model — the right equations, the right assumptions, the right failure modes. This goes beyond numerical correctness to physical correctness. A code can produce a numerically correct solution to the wrong equation.
- Does the code implement the intended engineering equation?
- Are the assumptions documented and consistent with the problem?
- Does the result agree with a simple analytical or order-of-magnitude calculation?
- Does the result behave physically — thicker structure → lower stress, stiffer support → less deflection?
- Are the units correct throughout?
Software testing asks whether the code works as written. Engineering verification asks whether the code represents the intended mathematics and physics.
Level 5 — Validation
Where physical data exists, validation compares the computational result against measured behaviour. Validation asks whether the engineering model — even when correctly implemented — agrees adequately with reality for its intended use. Validation is the highest level of confidence but is not always possible, because test data may not exist or may be too expensive to obtain.
Limit Cases
Limit cases test code at the boundaries of its intended use. They reveal whether the code behaves correctly in extreme or degenerate situations, which is often where hidden assumptions break down.
- Zero load — deflection should be zero; stress should be zero
- Very stiff limit — displacement should approach zero
- Symmetry — symmetric loading on symmetric structure should produce symmetric response
- Simple geometry — a uniform bar in tension should give uniform stress
- Known special case — simply supported beam with central point load has known analytical solution
Regression Tests
Regression tests ensure that future code changes do not alter validated behaviour unintentionally. When a function is verified against a benchmark, the benchmark result becomes a regression test. Any future change that alters the result is flagged, preventing subtle drift in numerical behaviour over time.
CODE CHECK: Does a code change produce the same results as the previous version? If not, the change has altered the numerical behaviour and must be investigated.
Dimensional Checks
Every engineering calculation should be checked for dimensional consistency. Force equals stress times area — the units must work out. A dimensional inconsistency is a definitive sign of an error, either in the equation or in the implementation. Automated dimensional checking, where practical, catches these errors before they reach the output.
ENGINEERING CHECK: Do the units of the result follow from the units of the inputs? If not, there is an error in the equation or the implementation.
Independent Calculation
The strongest single verification activity is an independent calculation — computing the same result using a different method, a different implementation or a hand calculation. If two independent approaches agree, confidence is high. If they disagree, the discrepancy must be understood before the result is used.
- Hand calculation for a simplified version of the problem
- Alternative implementation in a different language or library
- Commercial solver as a check on a custom implementation
- Published benchmark result as an independent reference
Debugging Engineering Code
Debugging engineering code is harder than debugging ordinary software because the symptom may not be an error message — it may be a result that is plausible but wrong. Standard debugging techniques apply — breakpoints, logging, step-by-step execution — but engineering debugging also requires checking the mathematics, the units and the physical behaviour of the output.
- Check the simplest case first — does the code work for a trivial problem?
- Print intermediate results and check them against hand calculations
- Compare against a known solution at each stage of the calculation
- Check units at every step, not just at the output
- Look for assumptions that may not hold for the case being tested
AI-Assisted Programming and Verification
AI-assisted code generation can produce code that looks correct, follows good style and runs without errors, but implements the wrong mathematics or uses incorrect units. AI-generated code must be subject to the full verification hierarchy — unit tests, benchmarks, limit cases, dimensional checks and engineering verification. AI can assist with generating test cases or suggesting debugging approaches, but it cannot replace the engineering judgement that determines whether a result is correct.
AI-generated code is not verified code. Code generation is not code validation. Every level of the verification hierarchy still applies.
Controlled Release
A verified engineering tool should be released with a defined version, documented purpose, known limitations, test suite results and benchmark comparison records. This makes the tool usable by other engineers with confidence and makes future changes traceable.
Equation → manual benchmark → code → test case → comparison → controlled release
Key Takeaways
- The most dangerous error is a plausible-looking incorrect answer, not a crash
- "It runs" is the lowest level of confidence — five levels of verification exist
- Software testing checks the code; engineering verification checks the mathematics and physics
- Limit cases reveal hidden assumptions that break down at the boundaries
- Independent calculation is the strongest single verification activity
- AI-generated code is not verified code — the full hierarchy still applies