
When a measure becomes a target, it ceases to be a good measure. Marilyn Strathern, formulating Goodhart’s Law
Let me be precise about that, because the imprecise version is wrong. Badly done V&V is no trouble at all for most managers. Well-done V&V is a nuisance for management. The reason is that well-done V&V finds problems and changes plans. Badly done V&V just lets plans and programs proceed unchanged.
The reason is simple. V&V finds problems and shortcomings in programmatic work. Those problems are anathema to a plan. They require extra effort to reach the level of success that managers have already described to someone above them. Management has over-promised, and V&V is the instrument that measures the shortfall. It puts a number on the distance between what was promised and what was delivered. These days, the funding is attached to the over-promise. Thus, V&V is a threat to funding. That threat needs to be neutralized.
The Checkbox
The result is that V&V becomes a desired checkbox for quality while the managers who want the checkbox decline to pay for the quality. That is a corrosive position to occupy. Interest in doing the work well steadily declines, and what survives is a stamp that certifies poorly done work as acceptable. V&V simply becomes a vehicle to justify the management’s BS.
The conclusion I have come to is that management needs approval for a great many things it has already done. It has committed to aggressive schedules across many projects. It has promised that modeling and simulation would return something large. On top of all that, it has spent a fuck-ton of money on big, flashy, expensive computers. All of it has to pay off, and it has to pay off visibly. Real V&V would show that the return is smaller than promised. Fake V&V supports their chosen narrative.
V&V done well questions every piece of that. What it would find is that the aggressive schedules, the enormous budgets, and the effort have gone into a variety of the wrong places. Managers would need to adjust and fix things. They would need to admit that the errors and predictions are not as good as expected. Fault would need to be admitted. These days, those are all career-limiting actions.
With four parameters I can fit an elephant, and with five I can make him wiggle his trunk. John von Neumann, as recounted by Freeman Dyson
This began to dawn on me when I realized that in the programs at Sandia, virtually no validation was actually being done. They would say there was validation, but this was an illusion. There was preliminary validation. Pre-shot simulations were compared to post-shot results. Then the model was modified and calibrated until it matched those results better. Usually the tests had no defensible error bars, so what we held was a set of point values from a single experiment, often not repeatable, set against a calculation.
That is not validation. That is curve fitting with a schedule attached. Since they had calibrated with the only test data available, validation becomes impossible. What we are left with is modeling in isolation, where the model form error cannot be determined. We spent an enormous sum on simulation technology, then systematically shot ourselves in the foot.
Burying the Evidence
Unless the individual truly responsible can be identified when something goes wrong, no one has really been responsible. Admiral Hyman G. Rickover
Then there is the example I have spent a great deal of ink on. Straightforward verification was performed on a code used for important problems, against analytical results that are highly relevant to those problems. The results were negative. They indicated real trouble with the code. Management’s response was to bury them, misclassifying a record to keep it away from prying eyes. This is corrupt. It is mismanagement. Worse, it guarantees that a problem in an important code will never be addressed, or even acknowledged. Instead, everyone is told to look away.
The same thing happens with validation results. Those are often properly classified at a high level and legitimately kept from public view, but the practice of burial is common and pernicious regardless. In substance it is the larger problem, because it leaves us with a very poor idea of how faithful our physical models are to the systems we use them to analyze and certify.
The human understanding when it has once adopted an opinion draws all things else to support and agree with it. And though there be a greater number and weight of instances to be found on the other side, yet these it either neglects and despises, or else by some distinction sets aside and rejects, in order that by this great and pernicious predetermination the authority of its former conclusions may remain inviolate. – Francis Bacon
The real victim of all this is quality, and science.
I have written repeatedly that properly practiced V&V is simply the scientific method applied to computation. When you meet this kind of obstinacy, this inability to deal with a problem once it has been named is the crime. What suffers is the quality of the results and any warranted confidence in them. The opportunity to mount a focused scientific effort to improve the models, the codes, and the methods is swept away with the inconvenient finding.
The purpose of V&V is to provide evidence that the work is good enough, and that the money bought something real. By massaging and burying results as they see fit, management short-circuits the mechanism entirely. They install a low-quality result as the standard rather than pursue excellence.
What these trends portend is institutional decay. Rather than engines of innovation and progress, the labs have become engines of the status quo: engines of “good enough.” Instead, they are engines of mistakes buried under (improperly) applied classification. In my view, the root cause is the preeminence of money as the measure of all things. Money has become the stand-in for quality and success. It is in the place of any real measure of technical merit or scientific progress. V&V is merely the place where the trend becomes most visible, and the place that suffers first. All the follow-on work that V&V should spearhead is abandoned with it, killed in the cradle.
Management simply wants to broadcast that everything it does is high quality and good enough right now, and that it is succeeding. There is an absolute inability on its part to mount any rational response to a genuine problem with the work.

Exascale: The Hero Calculations and No V&V
There is nothing so useless as doing efficiently that which should not be done at all. – Peter Drucker
A good share of the blame belongs to the Exascale computing program. One of its most pathological features was a general lack of support for V&V. There was no V&V element in the program at all. A small piece survived inside the NNSA portion, and that was a holdover from ASC. The reason is not mysterious. V&V would have delivered bad news about Exascale, and that should have been obvious from the outset.
Computing power was never the long pole for improving accuracy. The problem is, and was, the physical models themselves. Their limited accuracy, together with the intrinsic statistical variability of many phenomena that experiments generally fail to capture, is the actual source of the largest errors. V&V would have confirmed this and shown the level of investment in hardware to be foolhardy.
Anyone can build a fast CPU. The trick is to build a fast system. – Seymour Cray
Some studies could get you punished. I remember using the computers at Oak Ridge National Laboratory, where you would be penalized if anyone discovered you were running UQ studies on their big machines. UQ via sampling was trivial to scale up to the entire machine.
The reason was that the whole program had staked its claimed impact on the power of those machines to perform massive single-purpose hero calculations. That posture ignored how inefficient added resolution is at improving accuracy when the order of convergence is low, typically first order or considerably worse, and worse still once you reach phenomenology like turbulence, where all the statistical variability lives. No single hero calculation is science. It is a stunt. You can do them, but they have no error bars. The meta-belief was that the single calculation was “perfect”. This is like the DNS myth. The hero calculation is as good as an experiment.
A well-done V&V study would have made this plain and cast substantial doubt on the efficacy of the investment. Instead, politicians and managers were unified in their desire to promote success commensurate with the money, and to supply reasons why the money had been well spent. Flashy computing and unvalidated results were what they wanted to show. Under no circumstances should real science look under the hood and assess the quality of any of it. That could only produce bad news and demonstrate that the effort had been misdirected.
The Real Problem Is Balance
You can see the computer age everywhere but in the productivity statistics. – Robert Solow
The real problem is the lack of balance.
Investment in high-performance computing hardware, and in the software to use it, is money well spent. That is not the complaint. The problem with Exascale was that it was utterly unbalanced. Alongside the absence of V&V, there was no serious focus on methods or algorithms, and none on physical models or experiments.
Modeling and simulation is an integration across the whole of science. For it to flourish, every part of the scientific enterprise feeding it has to be healthy and pushing the state of the art forward. The program ignored the fact that methods and algorithms have delivered as much improvement in effective computing capability as the hardware has, and arguably more. Instead the focus on hardware became pathological.
We are now watching the same pathology repeat with artificial intelligence: enormous investment in data centers and computing hardware, with the same inattention to everything else. It is inflating an economic bubble that will very likely end in pain. That pain is almost entirely a product of the imbalance we saw in scientific computing, inherited wholesale by the work on AI. It will end up an economic own goal, completely unnecessary, and it will do immense damage.
The Canary
Men, it has been well said, think in herds; it will be seen that they go mad in herds, while they only recover their senses slowly, and one by one. – Charles Mackay
V&V should be supplying detailed feedback and evidence to the scientific enterprise. AI too. It is the fulcrum of a balanced program that self-examines.
Instead it has become the canary in the coal mine.
Its neglect, and the generally poor quality of the practice that remains, is a harbinger of much larger problems: the absence of balance, and the inability to react flexibly to where the real bottlenecks in efficiency, efficacy, and quality actually sit in computational science.
It tells you something about how the AI effort has unfolded, too, and where the dangers in that activity are going to be found. We are probably too far in to avoid the collapse. One hopes we learn our lesson. Recent history seeds doubt.
“Science is a way of thinking much more than it is a body of knowledge.”– Carl Sagan




