
“No man ever steps in the same river twice, for it is not the same river and he is not the same man.” – Heraclitus
I felt that my last post needed an addendum. In the meantime, other events have overtaken my mind. Things I will write about in the fullness of time. Now they are too new and sudden to have gotten the necessary consideration. Now back to the thoughts about the philosophy of simulations of physical systems, engineered or otherwise.
“To consult the statistician after an experiment is finished is often merely to ask him to conduct a post-mortem examination. He can perhaps say what the experiment died of.” – Ronald A. Fisher
I’m constantly struck by the damaging role of wishful thinking in simulation and modeling. When we conduct very expensive experiments, we want to understand the results. Unfortunately, these expensive experiments are often done exactly once. Worse yet, these are often supposed to help characterize systems we want to make in large numbers. This can be especially bad when the experiment does not go well. We want to know why, and often we enlist simulations to understand it. This is where things start to go awry.
Most experiments and phenomena are not repeatable, especially as they become more complex and energetic. The complexity and energy often result in instabilities and turbulence. More generally, the physics travels through critical points where the evolution takes one route or another. More things are happening at the same time, which we call multi-physics. Together, these all mean the experiments will get different results every time, no matter how much we make things identical. This should be obvious. Yet most of the users of our codes (analysts and designers) want to treat every experiment like a one-off.
“Chance is only the measure of our ignorance.” – Henri Poincaré
Instead, the reality is that an experiment is a single draw from an unknown statistical distribution. Worse yet, we don’t know those statistics. We don’t know where on the distribution any given experiment is drawn unless we do many experiments. It could be in the tails or from the middle, or… The de facto assumption is that the experiment is simply the mean of the distribution. Even worse, we do almost nothing to figure it out. We don’t repeat experiments or structure our models to do the same. The damage of this attitude is hard to fathom. It is vast.
“The most important questions of life are, for the most part, really only problems of probability.” – Pierre-Simon Laplace

This is one of the biggest gaps in our current scientific and engineering practice. I saw it at Los Alamos and Sandia. It is a common attitude and stood in the way of learning more. It compounds our broad societal tendency to misunderstand statistics. While this is tolerable for the layman, it is unforgivable from our highest-level institutions. Yet, even there, the attitude rules actions and fuels stagnation.
Most simulations of large-scale experiments try to replicate the result of that specific experiment. The problem is that these experiments are not repeatable. In a broad sense, this is always true. If we tried to do the same experiment N times, we would get N answers. The results would form a statistical distribution. The validation problem is that we do very few repeat experiments; most are single shots. Modelers try to simulate that precise experiment. The real impact is modelers calibrating the models on the basis of this single experiment. If the experiment is not the mean or median, the calibration may be extremely harmful.
“There are never two beings in nature that are perfectly alike.” – Gottfried Wilhelm Leibniz
Part of the issue what I last wrote about. We homogenize the materials, treating them the same without regard to scale. To me, this is obviously an approximation. It becomes systematically worse with grid refinement, not better. For some systems, it will neutralize the advantage of high-performance computing. The gross large-scale features are studied. The problem is that real materials encode some randomness into all problems. The hope is that these do not matter. The problem is that we do not verify that this is true. We have scant knowledge about how these random details drive variation in reality. We just assume it away. It is a if we learned nothing after Newton and still live in an assumed deterministic universe. We do not.
If the small-scale details impact the results critically, this is ill-posed and hopeless. If the results depend upon unstable phenomena, it can push the model through critical points. At those points, the model will turn to one major branch or another. We sometimes get good results by calibrating the simulation. This fails the desire to be predictive. The fault is largely philosophical. We are trying to solve a well-posed initial value problem that is not a well-posed initial value problem. There is a fundamental bias among physicists to view each experiment this way. Even when it is obviously not true.
The consequence of this cognitive dissonance is a lack of knowledge. Generally speaking, we have little or no idea of the distribution of results for these engineered objects. I saw this happen time and time again in my career. It was the primary philosophy of simulation at all the Labs. All of them, and for virtually any experiment of consequence. I’ll roll out a few examples of the sort of situations where it was obviously the wrong thing to do.
Most of what the analysts and designers do is quite defensible. Many experiments have gross large-scale features that lead to the unusual results. Sometimes the construction or manufacturing is flawed, or standard precision is off. All of these are common things to pursue. These studies are generally limited by the ability to input things into the model. The painting of materials upon initialization is one such limitation. To change this, the basic nature of material initialization would have to be modified. The materials would need to be measured and characterized at the scale of these details.
“With four parameters I can fit an elephant, and with five I can make him wiggle his trunk.” – John von Neumann
A canonical case where single experiments are terrible representatives of the mean is car crashes. Here solid mechanics is the ruling physics. It is obvious that each crash is different. This is ruled by the complexity of the engineered item, the car (or truck), combined with the complexity of material response. It is a place where the structural detail of the material should be added stochastically. The homogenized nature loses a key part of the response. In cases where the vehicle is very expensive, a single experiment could be extremely misleading. Yet in practice I have observed the analysts calibrate the entire population of vehicles to this single point. We nullify some of the most important powers of simulation by failing to understand the statistics of these events.

“Since all models are wrong the scientist must be alert to what is importantly wrong. It is inappropriate to be concerned about mice when there are tigers abroad.” – George Box
High explosives are very heterogeneous materials. They are typically composed of three components: energetic material, binder, and void. Thus, at a small scale, the shock wave produced is uneven and corrugated. Their representation as an ideal shock in homogeneous material only works at a macroscopic level. If these corrugated shock waves collide or converge, we can expect the variations to be amplified. Is this an effect we account for? Clearly, it can be seen in images of explosions where large variations appear. Even more importantly, how do these variations impact technology and other materials combined with them? Again, this is a place where single experiments could be extremely misleading. The outcomes are likely to be highly statistical. Do we have any real understanding of this?
Fusion capsules are a third example. To achieve successful fusion. we need to compress the fuel to a massive degree. Anything non-homogeneous in the material or design is going to be amplified to the same degree. Thus, any material variation will also be amplified. How much of the ubiquitous mixing and turbulence arises from small heterogeneous bulk aspects of materials? Do we know? Again, this would be a place for proper simulations to shed light. Are we looking here, or is it a rock we refuse to turn over? My understanding is that the focus has been on things like surface finish and roughness. Materials are still painted homogeneously. It is also clear that the experiments are not repeatable and admit a statistical distribution of outcomes.
“Absence of evidence is not evidence of absence.” – Carl Sagan
There are many more examples out there. In a totally different vein, one can look at initial conditions for supernovae. The details of the star before it explodes are essential. Its structure and rotation, plus evolution, all imprint onto the details we can see. The cloud formed after the explosion and the radiation signature all indicate that these explosions are each unique to some degree. A companion star can also influence the outcome greatly. This would include the orbital mechanics at the precise moment of the explosion. The differences in the initial conditions matter greatly. It is a natural place to look at the statistics of what can be measured. We should look for explanations in this, and simulation could be helpful.
The important thing is that codes need to be structured to allow such modeling. My last post on heterogeneous scale-dependent materials is an example of such. It would simply also be a more realistic representation of the problem being simulated.
“It is better to be vaguely right than exactly wrong.” – Carveth Read
In all these cases, the issue is really more deeply philosophical. When we observe an event or an experiment, the tendency is to look at it as a single thing. It is not viewed as a part of a distribution. The statistical nature of the results is not taken into account. In the parlance of physics, it is seen as a single well-posed initial value problem. It is not. Most of the important cases are not. We should expect the results to be variable. There should be a different outcome if the identical event or experiment were conducted. Real understanding would unveil this bit of knowledge. Even in cases where specific features caused the result, the statistical nature does not fade. The same maxim holds there too.
“For a successful technology, reality must take precedence over public relations, for Nature cannot be fooled.” – Richard Feynman
hey Bill – while I agree with your central point that both experiments and simulations are too often treated as single representative events, I wouldn’t want your readers to assume that’s always the case – in at least two research areas I participated in while at ORNL, metal additive manufacturing and batteries, performing multiple experiments in order to study variability was/is common – here are a few examples if anyone’s interested (note that several of the AM ones are from your Sandia colleagues):
Metal Additive Manufacturing
High-throughput stochastic tensile performance of additively manufactured stainless steelhttps://doi.org/10.1016/j.jmatprotec.2016.10.023
Extreme-Value Statistics Reveal Rare Failure-Critical Defects in Additive Manufacturinghttps://doi.org/10.1002/adem.201700102
Automated high-throughput tensile testing reveals stochastic process parameter sensitivityhttps://doi.org/10.1016/j.msea.2019.138632
Batteries
Failure statistics for commercial lithium ion batteries: A study of 24 pouch cellshttps://doi.org/10.1016/j.jpowsour.2016.12.083
Estimation of Li-Ion Degradation Test Sample Sizes Required to Understand Cell-to-Cell Variabilityhttps://doi.org/10.1002/batt.202100148
Analysis of the number of replicates required for Li-ion battery degradation testinghttps://doi.org/10.1016/j.est.2024.114014
Comprehensive battery aging dataset: capacity and impedance fade measurements of a lithium-ion NMC/C-SiO cellhttps://doi.org/10.1038/s41597-024-03831-x
obviously this is much more difficult for the extremely difficult, expensive, and/or dangerous experiments we’re familiar with – just pointing out that beyond those particular applications it’s maybe not quite as bad as you think
-JT
John,
You’re quite right. There are places where lots of experiments are done. Material characterization is one area, and apparently batteries too. I’m sure ensembles are done in other areas. Calibration of materials are also done multiply. The question really comes down to simulations of the system and its aim. Do the simulations have the mechanisms leading to variable outcomes with in it. Most complex engineered systems (cars, planes, bombs, power plants etc, …) would have a varied response to operation or accidents. Multiple experiments become rarer as the systems become more complex and expensive. Simulations should be a way to explore the variability. The common practice I saw over and over at LANL and SNL was to try to simulation that one expensive experiment as a deterministic problem. The natural variability is simply ignored. The result is a really grim lack of knowledge for the systems that are the apex of the system.
Nevertheless I stand corrected and as usual there are places where practices are better. Thank god!