Beyond Moore’s Law: How Algorithms, Data Movement, and Co-Design Drive Performance
Key Highlights
- Moore's Law slowing down is primarily due to wiring resistance and power consumption, not transistor shrinking, leading to a shift from speed to parallelism and efficiency.
- Smarter algorithms, like adaptive mesh refinement, can significantly reduce computational workload and energy use by focusing resources where needed most.
- Data movement inside chips consumes more energy than computation itself; reducing data precision and optimizing data flow are key strategies for efficiency.
On October 1, 2026 in Computing, Dev Tools, EIT 2026, Power, Wide Bandgap by Mouser Technical Content Staff
Based on an interview with John Shalf, Senior Scientist and Department Head for Computer Science, Lawrence Berkeley National Laboratory
For most of computing history, software engineers could rely on a predictable pattern. Write code today, and next year’s hardware would run it faster, with no extra effort. As time went on, clock speeds climbed, transistors shrank, and performance arrived almost automatically. That pattern is now weakening, and the engineers who understand why will be prepared for the design decisions ahead.
John Shalf, Senior Scientist and Department Head for Computer Science at Lawrence Berkeley National Laboratory, has witnessed this shift firsthand. His own career reflects how dramatically the field has changed. When he interviewed with companies in the early 1990s, a common refrain was, “Why would you do parallel computing? Why would you waste your career like that?”
For Shalf, that perspective changed at a 1992 conference, where he watched astrophysicists collide simulated black holes inside of the CAVE, an early virtual reality environment. Inspired by what he saw, Shalf changed course and dedicated himself to parallel code. Today, that once-common question sounds almost antiquated. Parallelism, the practice of running many calculations simultaneously rather than sequentially, has become a challenge for the entire computing industry. As computing evolves in a post Moore’s Law era, understanding how the industry arrived at this point is key to knowing what comes next.
Shalf has spent decades studying the limits and opportunities of high-performance computing (HPC). Mouser sat down with Shalf to examine the reasons why processor speed gains have slowed and why design engineers are adopting smarter algorithms, reducing energy-intensive data movement, and leveraging architectural specialization through continuous hardware-software co-design to continue improving computing efficiency and performance as the automatic progress of transistor scaling fades.
John Shalf is a Senior Scientist and Department Head for Computer Science at Lawrence Berkeley National Laboratory, where he has led research in high-performance computing (HPC) for more than 25 years. Shalf previously served as Deputy Director of Hardware Technology for the US Department of Energy’s (DOE’s) $1.6 billion Exascale Computing Project and Chief Technology Officer (CTO) for the National Energy Research Scientific Computing Center (NERSC), which supplies computing power to the Office of Science for DOE.
An expert in parallel computing, computer architecture, and semiconductor technologies, Shalf focuses on advancing next-generation computing systems through innovations in algorithms, chip design, and hardware-software co-design. In addition to co-authoring “The Landscape of Parallel Computing Research: A View from Berkeley,”[1] an influential whitepaper on the future of parallel computing, he is a Distinguished Lecturer for the IEEE Electronics Packaging Society, a Senior Member of the IEEE, and a recipient of the Seymour Cray Award for his work and leadership in HPC.
What Is Actually Slowing Down
The phrase “Moore’s Law is ending” gets used loosely, and Shalf is careful to distinguish what is actually slowing from what is not. Packing more transistors onto a chip has not stopped. “The ability to shrink transistors is continuing, and will likely continue with 3D stacked gates and other topological approaches,” he says. The real change came from a related trend that ended around 2005.
For years, each new process generation allowed designers to lower transistor operating voltages and reduce transistor size. Because a chip’s power consumption rises with both its clock frequency and the square of its voltage, even a modest voltage reduction produced a large efficiency gain. When voltage scaling ended, the easy speedups stopped with it. This effect showed up in clock frequency, with Shalf explaining, “clock frequency stalled in 2005, and we’ve made up the difference with massive parallelism.” Instead of faster chips, the industry shipped chips with more processing cores, so many calculations could run side by side. However, that is not the end of the challenges to scaling.
Lately, the problem has moved into the wiring. Transistors still improve in performance as they shrink, but the wires connecting them do not. “Wires, when you make them smaller, the resistance goes up,” Shalf says. That rising resistance now dominates a chip’s energy budget, and the consequences appear in power bills and cooling systems. The graphics processing units (GPUs) that handle much of today’s heavy computation clearly illustrate this trend. “Each generation of GPU is dissipating double the amount of watts of the previous generation,” he notes. Efficiency is still improving, “but not at historical rates.” That widening gap between rising density and rising power is what most people mean when they say Moore’s Law is ending.
Why Processing Limits Reach Beyond Supercomputers
It would be easy to treat this shift as an HPC concern alone, but Shalf argues that it touches nearly every design. “Parallelism has become pervasive across all scales,” he contends. A phone or laptop now carries multiple central processing units (CPUs) and GPUs. “It’s not just an esoteric problem anymore. It’s everybody’s problem. And so, we have to have parallel programming systems for the masses.”
The deeper issue is that the historical workaround, cutting a problem into ever smaller pieces, is reaching its own limit. The common approach today exploits “data parallelism” by dividing a large problem into many small chunks that can be processed at the same time. “We are really close to the limits for how much you can chop up a problem before it doesn’t scale up anymore,” Shalf warns. For simulations such as weather modeling or computational fluid dynamics, which models how air and liquids flow, finer spatial detail also forces smaller time steps to keep the simulation stable, so the total work grows faster than the resolution. With clock rates frozen, more time steps mean more waiting. “There comes a point where it’ll take longer to simulate the phenomenon than it actually takes to just do it. And that’s a failure mode.”
The fix, in his view, involves more than hardware. “It’s not just about the hardware. It’s not just about the software, but it’s about the actual underlying mathematics and algorithms.”
Smarter Algorithms Shrink the Problem
An algorithm is the step-by-step method a program uses to reach an answer, and the same answer can often be reached by very different methods—some far cheaper than others. When hardware improvements were automatic, an inefficient algorithm mattered less, because next year’s chip would make up the difference. That automatic help is gone, so the choice of algorithm now directly affects how much computing a problem requires. A better-designed method can produce the same quality of result while performing far fewer operations.
Shalf offers a concrete example from engineering simulation. To model how air flows over an aircraft, a program divides the surrounding space into a grid and calculates the physics at each point. A finer grid gives a more accurate result, but it also means more points to compute. The traditional approach makes the grid uniformly fine everywhere, including the calm regions far from the aircraft where little is changing. Adaptive mesh refinement is a smarter method. It computes finer detail only where the most important physics occurs, such as close to the airframe and in turbulent regions, while leaving the grid coarse elsewhere, often achieving comparable accuracy with far fewer points. “You can use a tiny fraction of the amount of compute and memory to solve the same problem,” Shalf says.
This lesson extends well beyond aircraft. In computing, an “Op” is a single memory or arithmetic operation, and as Shalf puts it, “the most energy efficient Op is the one that you don’t perform.” The cheapest calculation is the one you find a way to avoid.
Energy Bottlenecks: The Real Cost Is Moving Data
If one idea deserves to reshape how engineers think about performance, it is the energy cost of moving data. Inside a chip, a small block of circuitry does the actual arithmetic, and the operands (the data) it works on must travel to it from memory across the chip’s wiring. The same resistance problem driving up power has made moving those operands, rather than computing with them, the part that burns the most energy. “The wires consume more than 10 times more energy than the transistors that they’re connected to,” Shalf says. By his account, the energy spent moving data to the arithmetic “now exceeds by a factor of 10” the energy the arithmetic itself consumes. The computational floating-point unit (FPU) is nearly free, but delivering the numbers to the FPU is what consumes the energy.
One way to cut that energy is to make the numbers themselves smaller. Every value in a calculation is stored as a string of binary digits, or bits, and more bits mean a more precise number. These bits also mean more data to move across the chip. Reduced-precision arithmetic uses fewer bits, which narrows the data path and lowers the energy each operation costs. The tradeoff is accuracy. A 64-bit number is so precise that “you can measure the distance from earth to the sun in human hairs,” Shalf says, which is far more precision than many calculations need. On the other hand, cutting to 16-bit or 8-bit numbers can make some calculations “fall apart,” and “predicting when it’s going to fail is a challenge.”
Specialization, Chiplets, and Their Limits
The architect’s usual response is specialization, designing a chip to do one job or a range of more specific jobs extremely well. “Architectural specialization is definitely the architect’s go-to,” Shalf says. The GPU is the clearest example; a processor built for graphics whose math happened to suit scientific work. Much of the benefit comes from arranging the chip so data only travels where it needs to go. Specialization is now ubiquitous. Analysis of modern mobile SoCs shows that much of the die area is now devoted to specialized accelerators, and companies such as Amazon, Google, and others have built custom silicon for their own workloads.[2]
Chiplets push the idea further. Rather than build one large expensive chip, designers can fabricate smaller reusable blocks and package them together, which can improve yield, enable modular reuse, and reduce some system-level communication distances, depending on the package and interconnect architecture. But specialization has limits. “Not everything can be successfully specialized,” Shalf cautions. A problem limited by how fast data can be fetched rather than how fast it can be processed will see only so much benefit.
Behind all these choices is co-design. Traditionally, hardware designers, software developers, and algorithm designers work in sequence: the hardware arrives, and everyone else adapts to it. Co-design replaces that handoff with a continuous conversation and adjustments during the design process, so the chip, software, and algorithm are shaped around each other instead of one at a time. As the automatic gains from smaller transistors fade, that coordination is where much of the remaining performance is found.
Shalf illustrates it with a negotiation. An algorithm designer might ask for more memory bandwidth, which is the rate at which data can be delivered to the processor. The architect, knowing bandwidth is expensive, asks whether repeating some arithmetic in a way that avoids data movement could cut the amount of data that has to move instead. Working together, they reach a design neither would have found alone. He has also seen the cost of skipping that conversation. One CPU vendor concluded that a small high-speed memory called the cache should always be 16kB, not realizing that software teams had simply optimized their code to fit the previous generation’s 16kB cache. After revisiting the question and having the software experts re-tune the algorithm to match the size of larger caches, they found the optimal hardware-software design improved performance roughly tenfold. “Hardware, algorithm, software, co-design, and chiplet modularity is the future of the electronics industry,” Shalf says.
Conclusion
The era of automatic performance gains from faster hardware is fading, forcing engineers to rethink how computing systems are designed. As transistor scaling delivers diminishing returns and data movement becomes a dominant source of energy consumption, future advances will depend on a more holistic approach to system optimization. Smarter algorithms can reduce computational workloads, architectural specialization and chiplet-based designs can improve efficiency, and hardware-software co-design can uncover performance gains that isolated development efforts often miss.
For design engineers, the practical shift is to ask different questions earlier. Is this the right workload, and is the current software the best way to express it, or merely the way it was inherited? Another shift is to measure progress by useful output rather than by the chip’s raw specifications. Transistor counts and raw arithmetic rates describe the hardware, Shalf argues, not the work it accomplishes. What matters is how much useful output a system produces for the energy it consumes. In the post-Moore’s Law era, choosing the right measure of progress is itself part of the engineering.
Sources
[1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2006/EECS-2006-183.html
[2] https://arxiv.org/pdf/1907.02064; https://people.eecs.berkeley.edu/~ysshao/assets/papers/shao2016-dissertation.pdf

