“The most important figures that one needs for management are unknown or unknowable.” — W. Edwards Deming, Out of the Crisis (1986)

tl;dr

The miraculous nature of today’s world owes its existence to human progress. Some of that progress is serendipity. Some is targeted and guided. In science, measurement is the key both to inspiring progress and to recognizing it. For algorithms, methods, and computing, code verification is how we measure. The alternative to measurement is expert judgment. Both have value. But when you find a field that has stopped progressing, you will usually find too little measurement and too much expert judgment. If we want computational progress, measurement is our friend.

Progress Is Paramount

When I started graduate school, I worked on nuclear energy — specifically the more exotic aspects of it, in support of space exploration. As exciting as that might sound, I never developed any passion for it. It was a reaction to my professor, or perhaps just the topic.

When I started learning numerical methods for differential equations, though, I was full of excitement. The idea that you could simulate the real world on a computer fascinated me. I enjoyed the power and beauty of hyperbolic PDEs and their solution, especially the nonlinear methods that had recently come into vogue. I ate it up.

I learned numerical method after numerical method, gradually narrowing my focus to monotone shock-capturing schemes. FCT and TVD methods were in vogue then, and ENO methods were just being created. I was enchanted by the rapid progress of the previous decade and read paper after paper as the methods were developed and tested. I also grew fond of the associated mathematical literature, finding real beauty and elegance in how that mathematics supported both the methods and our understanding of them. As I devoured the literature, I began to study something else: how these methods were tested and demonstrated, and what the standard was for claiming one method was better than another.

What stands out in retrospect is that the prevailing view of “better” rested on a qualitative metric. The measure was the expert’s opinion about the vagaries and nuances of a visual plot. The judgment seemed defensible enough in one dimension. In two and three dimensions, acceptance of results was essentially an artistic verdict.

A few years later, after I had moved to Los Alamos, I noticed the working version of this standard: the gnarlier and swirlier the vorticity in a problem, the better the method was considered. And yet the progress made in these methods over a short period was astounding. As readers who know my general tendencies will guess, progress is something I care about deeply. Progress is what we should always be seeking.

“Working on the right things is what makes knowledge work effective. This is not capable of being measured by any of the yardsticks for manual work.” — Peter F. Drucker, The Effective Executive (1967)

Measurement Is Science

As this body of work became more familiar to me, the reliance on qualitative examination and expert opinion became more and more distasteful. I met a number of people who were genuine experts and who sat in judgment over methods and progress. Eventually I saw them for what they had become: gatekeepers. There was no objective measure of what was better. There was only expert opinion, and expert opinion favored those who already held power.

The problem with this status quo is that progress has screeched to a halt. Journal editors are thoroughly fed up with yet another paper developing and comparing limiters — and I am responsible for a couple of those myself. It became a cottage industry. Meanwhile, the standard of measurement for the problems that actually matter is nonexistent. This is the recipe for a stalled science. The key thing to recognize is the centrality of measurement to science. Unmeasured qualitative assessment is not science. The stagnation of progress is therefore no mystery — the two are correlated. Progress requires measurement.

This gets at the real tension over results and standards. The mathematics supporting this area is quite limited. Rigor comes with caveats and serious restrictions. Relatively few practical results carry any assurance from the theoretical foundations of the field. Into that vacuum steps qualitative gatekeeping, and the consequence is stagnation.

To overcome this, we need to measure results quantitatively. We also need a leap of faith: we must be willing to trust what we see empirically, in the places where theory cannot follow. We know qualitatively that modern methods were a quantum leap in capability. They let computations attempt problems that had been impossible, and they rapidly became standard. The problem is the threadbare theory. Consider the Lax equivalence theorem, which establishes that consistency plus stability is equivalent to convergence — but only for linear PDEs. All our experience says the theorem describes the utility of computing for nonlinear and far more complex problems. What we lack is the rigor to prove it. Stagnation has festered in exactly that gap.

“Absence of evidence is not evidence of absence.” — commonly associated with Carl Sagan

Methods, Algorithms, and Codes Need Metrics

As I matured as a scientist, I came to understand that measurement is feedback. Measurement is science, and method development needed it to flourish. This view generated a great deal of friction with the existing community — friction I found inside the Lab’s computational physics community as well as outside it. It came to a head in an exchange with an editor who demanded that I “just take that V&V shit out of the paper.” Gatekeeping at its finest.

That paper was the culmination of work at Los Alamos; it died on the vine at Sandia. In it I was looking for methods that could deliver quantitatively better solutions to discontinuous problems at an attractive cost — efficiency, in a word. Doing that requires unraveling which elements of method design actually contribute to better solutions, balanced against computational cost. The work grew out of an observation: WENO methods buy formal accuracy at high cost and fail to deliver practical accuracy. Jeff Greenough and I published that failure.

The next step was to turn that knowledge into a better method, which I believe we accomplished in the paper with Jeff and Jim Kamm. Code verification was central to demonstrating the progress, and that is precisely where we ran headlong into the gatekeeping. Results for problems with discontinuous initial conditions and shocks are evaluated qualitatively, full stop. Measurement and quantitative measures need not apply. We measure only for smooth problems, and only to confirm formal high-order convergence rates. That is the standard practice.

No wonder we are stagnating. No wonder there are thousands of papers tweaking limiters and WENO variants. Without measurement, progress is a random walk. I am still pushing this effort forward because I remain convinced it is necessary. Progress depends on aligning our approach with best practice; verification is part of that practice, and V&V must be quantitative. I got into V&V because it is how you do science correctly. The irony is hard to miss: engineers have embraced V&V while scientists and mathematicians treat it with genuine animosity. The problem is that engineers see V&V as a process and a check, not a fuel for progress.

“When you can measure what you are speaking about, and express it in numbers, you know something about it; but when you cannot measure it… your knowledge is of a meagre and unsatisfactory kind.” — Lord Kelvin, Popular Lectures and Addresses (1883)

Code Verification Is the Way

“Measure what is measurable, and make measurable what is not so.” — attributed to Galileo

Once I understood modern methods, I started to see the cracks in practice, and working alongside some of the leaders in the field amplified the view rather than softening it. In everything I did, I looked to measure the connection between theory and observed results. Not only comparisons against analytical solutions, but all results.

Analytical comparisons carry the great power of objective truth. Without them, you can estimate error, but you cannot know it. Given how soft the theoretical results are, comparison against analytic solutions deserves a spotlight. The catch is that the problems with analytic solutions are the ones furthest from application.

I had discovered that V&V was essential. V&V provides feedback on your code and your model, and that feedback is actionable. Qualitative comparison is actionable too, but the variations it suggests come from expert judgment — and experts have bias. That bias usually takes the form of projecting one’s own philosophy of algorithm design onto the results. I preferred something less prone to it. Code verification became my instrument for measuring progress, and the arc of my work through the late 1990s and 2000s reflects that. Along the way, I became more knowledgeable about code verification and more engaged with it broadly.

What is worth acknowledging is that my position runs counter to both communities. The algorithms and methods people carry animosity toward quantitative verification. The verification community that came out of engineering has its own bias: it adopts relatively simple practices and tends to ignore most of the mathematical expectations. Its members are dismissive of the equivalence theorem while adhering to its precepts far beyond the range where it rigorously applies. So I stood out as an outsider. This is a role I have grown comfortable with, and even proud of. The quantitative comparison of results gave me constant feedback and guided my work.

“Nullius in verba” (“Take nobody’s word for it”) — motto of the Royal Society, 1660

Why Is Code Verification Despised?

The logic of quantifying error and using it as feedback is clear enough. So why is the practice resisted so hard? Why the animosity?

The answer is the power bound up in gatekeeping. Experts who hold power in their own right have that power amplified by serving as the judges of progress. They ensure their own work remains the cornerstone. Their judgments involve a great deal of handwaving: the reading of solution structure and oscillations is judgment, and judgment can be made to serve its holder’s ends. Quantitative verification is a different animal, because numbers cannot be waved away — they still require expert interpretation, but they constrain it.

The animosity, then, is largely gatekeepers defending their hold on a field. They decide who publishes and who gets credit, and they reject those they dislike or who fail to reinforce their standing. This behavior is hardly confined to computational fluid dynamics. The turbulence community is equally stagnant, if not more so, and for the same reason: persistent gatekeeping throttling progress.

I saw gatekeeping at Los Alamos as well, where the nuclear weapons designers held the role. They were quantitative about validation results, but they achieved those results through extensive calibration. In particular, they would quite cavalierly calibrate away numerical error using physics submodels. This is a common practice as weather and climate modeling does the same basic thing.

The hostility toward quantitative measurement of code performance is deep and profound, and together with the CFD gatekeepers, it has blunted progress for decades. The result is a hegemony of older methods still in use long after they should have been replaced. It limits what modern computing can contribute to genuinely important problems. This remains a force that saps some of the value we should be enjoying from computing. Over the history of the computational age, methods and algorithms have provided efficiency equal to or greater than computing itself.

The case for being quantitative about practical problems is, in some ways, simple. Practical problems have extremely low convergence rates — first order is typically the best you can achieve, because the dynamics of numerical methods change considerably in the presence of discontinuities. Under those conditions you should care far more about producing small errors than about high convergence rates. Highly accurate, high-convergence-rate methods simply don’t pay off, particularly when they are expensive.

The real question is how to deliberately produce methods that achieve small error, remain robust, and selectively preserve the solution structures that matter. That is a fundamentally different design problem than chasing formal order of accuracy. Pursuing high-order methods formally can produce more error on practical problems. They only show their benefit when the problem becomes complex enough to contain a great deal of intricate structure. This points toward greater adaptivity to tie the two uses together cleverly. Quantitative assessment is key to success.

“It is wrong to suppose that if you can’t measure it, you can’t manage it — a costly myth.” — W. Edwards Deming, The New Economics (1993)

References

Sod, Gary A. “A survey of several finite difference methods for systems of nonlinear hyperbolic conservation laws.” Journal of computational physics 27, no. 1 (1978): 1-31.

Boris, Jay P., and David L. Book. “Flux-corrected transport. I. SHASTA, a fluid transport algorithm that works.” Journal of computational physics 11, no. 1 (1973): 38-69.

Harten, Ami. “High resolution schemes for hyperbolic conservation laws.” Journal of computational physics 49, no. 3 (1983): 357-393.

Jiang, Guang-Shan, and Chi-Wang Shu. “Efficient implementation of weighted ENO schemes.” Journal of computational physics 126, no. 1 (1996): 202-228.

Rider, William J., and Douglas B. Kothe. “Reconstructing volume tracking.” Journal of computational physics 141, no. 2 (1998): 112-152.

Rider, William J. “Revisiting wall heating.” Journal of Computational Physics 162, no. 2 (2000): 395-410.

Rider, William J., and Len G. Margolin. “Simple modifications of monotonicity-preserving limiter.” Journal of Computational Physics 174, no. 1 (2001): 473-488.

Greenough, J. A., and W. J. Rider. “A quantitative comparison of numerical methods for the compressible Euler equations: fifth-order WENO and piecewise-linear Godunov.” Journal of Computational Physics 196, no. 1 (2004): 259-281.

Rider, William J., Jeffrey A. Greenough, and James R. Kamm. “Accurate monotonicity-and extrema-preserving methods through adaptive nonlinear hybridizations.” Journal of Computational Physics 225, no. 2 (2007): 1827-1848.

Banks, Jeffrey W., T. Aslam, and William J. Rider. “On sub-linear convergence for linearly degenerate waves in capturing schemes.” Journal of Computational Physics 227, no. 14 (2008): 6985-7002.