Langford Analytic · Knowledge Base

Probability Density & Cumulative Distribution Functions

The PDF and CDF — what they represent, how they relate, and how engineers use them to evaluate exceedance probabilities, quantiles and characteristic values for structural assessment.

Article 04Probability for Engineers17 min read
PDFCDFprobability densitycumulative distributionexceedancequantilesengineering

Technical provenance

Applicable standards / specifications

  • ISO 2394 (2015 (confirmed 2026)) — General principles on reliability for structures

References

Engineering Context

Probability density functions and cumulative distribution functions describe the same continuous random variable from different viewpoints. The PDF shows how probability is distributed locally; the CDF gives the probability of being at or below a threshold. In engineering decisions the CDF, survival function and exceedance probability are often more directly useful than the visual peak of the PDF because design limits are threshold questions.

Core Mathematical Definition

For a continuous variable, the PDF f_X(x) is non-negative and integrates to one. Probability over an interval is the area under the density. The CDF F_X(x) increases monotonically from zero to one and, where differentiable, its derivative is the PDF. The survival function S=1-F gives the probability of exceeding x. For discrete variables, probability mass rather than continuous density is used, but the CDF definition remains valid.

F_X(x) = P(X ≤ x)

f_X(x) = dF_X(x)/dx

P(a < X ≤ b) = F_X(b) - F_X(a)

Survival: S_X(x) = 1 - F_X(x)

Evidence, Data & Population Definition

An empirical CDF is a powerful first look at measured data because it requires no assumed distribution. Each ordered observation increments cumulative probability. Plotting empirical and fitted CDFs can reveal tail mismatch that is less obvious in histograms. Histograms themselves depend strongly on bin width and should not be the sole basis for distribution selection.

Use in Structural Engineering

Structural allowables, characteristic strengths and environmental exceedance levels are naturally expressed through quantiles of the CDF. If a component fails when load exceeds L, the relevant quantity is a tail probability rather than the density at L. Likewise, if strength must exceed a required value, the lower tail of the strength CDF controls. Keeping the direction of exceedance clear prevents common sign and tail mistakes.

Integration with FEA & Simulation

When simulation produces a sample of response values, the output can be represented by an empirical CDF even if no parametric distribution is fitted. This is often the most defensible representation for Monte Carlo response. However, tail estimates from an empirical CDF are limited by sample size; a simulation with 1,000 runs cannot directly resolve a 10^-6 failure probability by crude counting.

Parameter Estimation & Data Quality

Kernel density estimates can smooth an observed PDF, but bandwidth selection affects apparent tail and modal structure. Parametric CDFs are easier to extrapolate but introduce model-form assumptions. For reliability work, fit quality should be checked where the limit state lies, not only around the distribution centre. A visually excellent central fit can still be poor in the tail that governs failure.

Sensitivity & Model Uncertainty

Threshold-based decisions can be extremely sensitive to small changes in tail shape. Comparing candidate CDFs, truncation assumptions and parameter confidence intervals at the actual engineering threshold is often more informative than comparing global goodness-of-fit statistics.

Engineering Interpretation & Decision-Making

Engineers should choose the representation that serves the decision. PDFs are useful for understanding concentration and modes; CDFs are useful for percentiles and non-exceedance; survival curves are useful for exceedance and life data. Switching between them is straightforward mathematically but can clarify the physical question dramatically.

Relationship to Deterministic Design & Standards

These concepts sit underneath deterministic design values rather than competing with them. Partial factors, allowables, characteristic values and qualification margins often contain implicit or calibrated reliability assumptions. A probabilistic study should therefore identify which conservatisms are already present before adding stochastic inputs; otherwise the same uncertainty can be counted twice. Conversely, using mean loads and mean strengths in a probabilistic model while comparing the result directly with a code-factored requirement can mix two design philosophies inconsistently. The correct relationship depends on the governing standard and purpose of the assessment. For internal design optimisation, the probabilistic model may operate on unfactored physical variables and a physical limit state. For certification, the probabilistic result may instead support sensitivity, equivalence or risk understanding while the formal compliance statement remains based on the prescribed deterministic framework.

Common Engineering Mistakes

  • Reading the height of a PDF as a probability
  • Forgetting that probability is PDF area over an interval
  • Using a histogram as though it were a unique estimate of the density
  • Extrapolating an empirical CDF far beyond the sample support
  • Checking fit only near the mean when the limit state is in a tail
  • Confusing lower-tail and upper-tail failure events

A Defensible Working Method

  1. Define the engineering quantity, population, units and reference condition before assigning any probability model.
  2. Identify the evidence source and separate measured variability from lack of knowledge or model-form uncertainty.
  3. Select candidate models using physical support and mechanism before applying statistical fit diagnostics.
  4. Represent dependence between variables where it arises from common manufacturing, loading or environmental causes.
  5. Propagate uncertainty through a verified engineering model using a method appropriate to nonlinearity and required tail probability.
  6. Check convergence and sensitivity specifically for the statistic or limit state used in the design decision.
  7. Document assumptions, data limitations, tail extrapolation and the effect of plausible alternative models on the conclusion.

Verification & Senior Review

The review should verify normalisation, support, tail direction and threshold interpretation. Empirical data and fitted distributions should be shown together where possible. If extrapolation beyond observed data is required, the physical and statistical justification for the chosen tail model should be explicit.

Engineering Review Checklist

  • The uncertain quantity and population are defined unambiguously.
  • Distribution support is compatible with the physics and any hard bounds.
  • Data provenance, sample size, censoring and measurement limitations are recorded.
  • Dependence between important inputs has been assessed rather than assumed away.
  • The statistical quantity used for acceptance matches the actual engineering limit state.
  • Tail behaviour and extrapolation are justified at the probability level used for the decision.
  • Sensitivity to uncertain parameters and plausible alternative models has been checked.
  • The probabilistic result is interpreted alongside consequence, deterministic requirements and model limitations.

Probability is useful only when the event, population, evidence and engineering consequence are defined as carefully as the mathematics. More sophisticated statistics cannot compensate for an ambiguous physical question.