This article hasn’t been deeply studied or thought through yet, and may still be substantially revised in the future.
Anomalies in Complex Systems
Whenever a major product failure happens, people’s first reaction is astonishingly consistent: find the person who made the mistake, or the part that failed. We eagerly look for a clear “root cause,” because it gives us an illusory sense of control—as if fixing this one point will put everything back on track.
In some situations, this simple way of thinking isn’t wrong—simple means fast. When you’re in a hurry to resolve a customer complaint and to get the on-site staff out of trouble, quickly pinpointing a “root cause” and claiming to have solved it is a workable solution. We often feel satisfied, even smug, when we have “solved a problem,” but we need to clearly realize that this kind of “solution” solves the people problem, not the problem of the whole system.
Whether it is an energy grid, a company, a piece of equipment, or a software system, these things are complex by design or in practice. No one can easily understand exactly how they currently work. They are in fact full of all kinds of minor faults; it’s just that enough redundant design lets them keep functioning normally.
At some point, a few minor faults suddenly join forces, a planned task cannot be completed, and an incident erupts. We have to deal with the incident and, based on our shallow understanding, slap on a patch that doesn’t look too bad, persuade the person who found the problem, and claim it is solved. I once accompanied a classmate to deal with an administrative dean at our school. Although my classmate and I cursed her for being so fussy at the time, she said something I think is quite philosophical—“Every swallow leaves a trace; everything you do has consequences.” In the same way, a patch we rush to apply brings more hard-to-detect minor faults into the whole system.
In engineering management, people tend to cling to a superficial “root cause” and ignore the real root cause. That’s because the people who actually do the work are giving “an account” to those above them, and an account is easier to get approved when blame is pinned on a particular person or thing. But in the end, this only conceals the systemic root cause.
This way of thinking belongs to the early part of the second stage of quality control, the “statistical quality control stage,”[^three-stages-of-quality-management] and is a legacy of the Ford assembly-line era, with its excessive focus on decomposition and standardization and its strong sense of causality. Compared with the first stage, the “quality inspection stage,” this is of course a huge methodological improvement. But in today’s world, the complexity of systems—products, engineering, society, organizations—has risen sharply, forming complex systems that are multivariable, nonlinear, changing in real time, and made up of variables that influence one another. It is already very hard for people to grasp the causal relationships among all the links, yet our way of thinking remains stuck in the old model of primitively and instinctively handling simple, linear relationships, creating a huge cognitive gap.
Focus on the System
If a coffee shop’s quality is good one day and bad the next, and one day a customer complaint appears, the manager’s first reaction is always “firefighting.” Call an emergency meeting, quickly find the barista on duty, accuse them of not calibrating the coffee machine properly, fine them, and compensate the customer.
Is that enough? The complaint was handled very promptly, but why do the same things keep happening? In the coffee shop as a whole system, the essential reliance on the barista’s craft to produce the coffee has not changed.
There are almost no coffee shops like this left in our time. Think about it: for any chain coffee shop, isn’t the product taste almost the same in every store? Of course we know this taste is not excellent; perhaps it’s not as good as what that occasionally fallible barista made. But this is the positioning of chain coffee shops. We simply produce products at this quality, and naturally we only serve customers who are satisfied with this quality.
Deming[^deming] divided all quality problems into two types.
The first type is called a “controllable failure.” It’s like your computer suddenly crashing with a blue screen. This is an abnormal, sudden disturbance with a clear cause—maybe an operating mistake, a broken piece of hardware, or a crashed driver. For this kind of problem, you must act immediately, find it, fix it, and make sure it never happens again. This is firefighting, and it must be done immediately.
But the second type is more common and more troublesome. Deming called it an “occasional failure.” This is more like your computer’s overall speed being fast at some times and slow at others. It is not caused by a single, clear fault; it is inherent in the system. Maybe the operating system is a bit bloated, too many programs are running in the background, the hard disk is low on space... countless tiny, random factors work together to create that overall, indescribable “laggy feeling.” This is the system’s “background noise,” and it is always there.
Obviously, because baristas are human, the quality problems caused by baristas are “occasional failures.” Chain coffee shop managers have cleverly downgraded their target customers, built a stable and sound coffee bean supply system, and reduced the complexity of barista operations to a minimum, using these as means to optimize the system.
Deming’s suggested path is to first extinguish all those “controllable failures” that suddenly catch fire. Through a set of standards (to be explained later), scientifically judge which signals are truly abnormal. When all “controllable failures” have been eliminated, the system enters a “stable state.” There are still problems and fluctuations at this point, but these are all normal noise.
At this point, the truly important improvement is just beginning. From now on, the root cause of every problem is no longer a particular person or thing, but the system itself. Managers need to improve the system more intelligently and carefully, and keep repeating the process of thinking and improvement.
How to Tell Whether a System Has Entered a “Stable State”
There are some mathematical methods and metrics, but I haven’t understood them yet.
PDCA and PDSA
First, the Deming cycle is the concept of repeatedly going through several phases to achieve system optimization.
PDCA is the “Deming cycle” widely accepted today. It stands for Plan-Do-Check-Act: plan, execute, evaluate, improve. It may be a misattribution, since Deming himself explicitly said he never proposed it.
PDSA is the purist “Deming cycle.” It stands for Plan-Do-Study-Act: plan, execute, learn, improve.
Repeating the four phases brings about a stepwise improvement of the system.
Modern “methodologies” sometimes mention the concept of “large cycles within small cycles,” and some phases are pretentiously expanded—for example, expanding C into 4C: Check, Communicate, Clean, Control. But my view is that a methodology should not be over-refined; refine it to the extreme and you end up with no methodology at all.
I think overemphasizing the cycle as the driving force kills a system’s innovativeness, which is fatal in certain stages of a system. At the same time, this approach lowers the system’s ceiling. The brilliant breakthroughs and innovations of a system naturally bring more minor faults. Therefore, I believe this quality optimization framework is only suitable for a “stable state,” and you should execute it by treating yourself as a tool.
In The New Economics for Industry, Government, Education (2nd Edition) | yono’s files, Deming also expresses a view similar to “don’t be dogmatic about methods; adapt to circumstances instead.” This book was Deming’s last work, and I strongly recommend downloading and reading it. The book also contains ideas such as the uselessness of performance rankings, everyone working to optimize the system, and rewarding employees with respect rather than material incentives—all quite idealistic ideas. As with PDCA/PDSA, just understand the master’s thinking and that’s enough.

image-20250624173735952
There is also the famous Deming’s Fourteen Points, which you can search for and study on your own. It is not very relevant to small fry like us.
Reflections
My biggest takeaway is to stop blaming a single point. When a problem occurs, it is actually the system’s design that has a flaw. So there is no need to be anxious or blame yourself—these are all problems for the people further up.
- three-stages-of-quality-management: 1. Quality inspection stage: before the 18th century, products usually came from workshops, and quality assurance depended on the skill and experience of the manual worker, with an experienced hand doing the final check. This inertia lasted into the early 20th century, but it was actually only “after-the-fact inspection” that picked defective products out of the finished goods. 2. Statistical quality control stage: it mainly used statistical methods and the process control charts proposed by Shewhart to detect defects in a process in time and improve them. 3. Total quality management stage: proposed in a 1956 TQC paper, it argued that quality problems arising in the manufacturing process account for only 20%, and put forward the idea of total quality management that fully considers market research, design, production, and service.Returnthree-stages-of-quality-management
- deming: An American quality management master who laid a solid foundation for quality management in Japanese business.Returndeming