The Missing Layer in AI Systems Design is Predictable, Scalable Data Movement

As AI systems scale, moving data predictably through the fabric is becoming just as important as compute and memory.

What you’ll learn:

  • How predictable data movement helps maintain performance as AI systems scale.
  • Why AI workloads need more than best-effort traffic handling from the interconnect.
  • Why sustained utilization may matter more than peak performance in next-generation AI hardware.
  • How larger architectures need tighter control over data movement across dies and domains.

Over the years, AI chip designers have focused on scaling systems, connecting compute, memory, and accelerators. It’s led to an evolution in chiplet architectures. Significant progress has been made in standardizing die-to-die interfaces and advancing packaging, making it easier to assemble more capable systems.

These technological advances have received most of the spotlight in this arena. Far less attention has been paid to what happens after everything is connected, specifically how data moves, how it’s prioritized and scheduled, and how reliably it’s delivered under real, multi-tenant AI workloads.

It’s not simply a question of bandwidth — it’s a question of deterministic data and control within the fabric. Data-movement predictability ultimately drives performance. When several workloads compete for bandwidth, memory, or interconnect, systems may appear efficient in ideal conditions. However, they quickly lose performance as real load increases. Utilization can become unpredictable, not because of insufficient computing, but because data isn’t delivered when and where it’s needed.

When Best Effort Breaks Down

The system tries to move data as quickly as possible, but under contention, it often lacks guarantees on latency, bandwidth, or ordering. That model breaks down in AI systems where multiple engines share memory and IO, and where training, inference, and control traffic all have different requirements.

AI workloads expose the limits of fabrics because the traffic they generate is fundamentally uneven and competing. Model weights move in large bursts and can tolerate latency, while other AI traffic has very different bandwidth, latency, and prioritization requirements.

That’s why the network-on-chip (NoC) interconnect must manage these flows predictably under load. When they share the same fabric without isolation or enforced quality of service, the system begins to behave erratically. Control traffic can be delayed behind bulk transfers, tail latency increases, and compute resources sit idle waiting for data.

This failure isn’t gradual. As contention rises, arbitration becomes less predictable, and tail latency increases. Adding more bandwidth doesn’t fully solve the problem, because the issue isn’t capacity, but rather how traffic is scheduled. In AI systems, performance is increasingly determined by scheduling behavior rather than peak bandwidth.

Architecting Control

Predictability must be treated as a core design requirement, not as a secondary optimization layered on later. In practical terms, this means defining explicit traffic classes, enforcing isolation between competing workloads, and establishing quality-of-service mechanisms that hold under load. Arbitration must behave deterministically even as contention rises, coherency must scale without introducing instability, and the state of the fabric must be both observable and controllable.

These issues can’t be addressed at the physical layer, solved through packaging, or corrected with firmware after the fact. They require a system-level approach to how data moves through the fabric. This effectively defines a control plane for data movement within the chip — a layer essential to system behavior but still lacking a clearly defined owner in most architectures.

As systems scale, the control plane becomes even more critical because data movement is no longer confined to a single block or even a single die. When traffic starts backing up in one part of the fabric, the effects can spread through the rest of the system, and without coordinated control, a small local problem can quickly turn into a system-wide performance issue.

That’s where a system-level NoC interconnect approach becomes essential, one that can manage traffic hierarchically, maintain coherency across domains, and provide visibility into how data flows under real conditions. Rather than treating interconnect as a passive pathway, it must actively manage prioritization, enforce isolation, and sustain performance as workloads interact. This shifts the role of the fabric from simple connectivity to a central mechanism for maintaining stability and efficiency across the entire system.

Sustained Utilization

Multi-tenant AI is the forcing function that makes such a shift unavoidable. Next-generation systems will run multiple models concurrently, combining real-time and throughput-driven workloads on shared silicon across users and services. That introduces fundamentally different requirements: Some workloads demand strict latency bounds while others prioritize throughput, and both must coexist without interfering with each other.

In these environments, the absence of predictable data movement leads to contention where one workload can consume disproportionate bandwidth, another misses timing requirements, and overall utilization drops. The typical response is to over-provision resources, adding more silicon and power to compensate for inefficiency.

With predictability built into the fabric, systems can enforce isolation, maintain performance under mixed load, and achieve higher sustained utilization. It not only improves efficiency, but also becomes a point of differentiation, where real performance under load matters more than peak specifications. In this context, fabrics that lack predictable data movement become inefficient to scale and increasingly costly over time.

By 2028, AI systems are likely to compete less on peak TFLOPS and more on sustained utilization under real workloads. What matters next isn’t simply building larger AI systems, but building systems whose behavior remains stable in the face of increasing concurrency, sharing, and workload diversity. Such a shift will separate designs that scale economically from those relying on overprovisioning. The company that defines predictable data movement will define the next layer of system architecture.

Between physical connectivity and software scheduling lies a critical layer where traffic is shaped, prioritized, and controlled. Arteris provides system IP that manages on-chip data movement via coherent and non-coherent interconnects, cache, and automation technologies, designed to improve performance, power efficiency, and scalability. This effectively forms the fabric's control plane, where quality of service, arbitration, and visibility are enforced as part of the architecture.

As AI systems continue to scale, defining this layer becomes increasingly important. If it’s not addressed within the fabric itself, it will emerge elsewhere as an additional abstraction above it, shifting control away from the NoC interconnect.

About the Author

Andy Nightingale

Andy Nightingale

VP of Product Management and Marketing, Arteris

Andy Nightingale, VP of Product Management and Marketing at Arteris, has over 40 years of experience in the high-tech industry. He's a seasoned global business leader with a diverse background in engineering and product marketing. Andy is a Chartered Member of the British Computer Society and the Chartered Institute of Marketing, and has over 35 years of experience in the high-tech industry. 

Throughout his career, Andy has held a range of roles, including engineering and product management positions at Arm, where he spent 23 years leading a product marketing team specializing in system IP products that involved network interconnects, memory and interrupt controllers, and system MMUs. In his current role at Arteris, Andy oversees the Magillem system-on-chip deployment tooling and FlexNoC and Ncore network-on-chip products. 

Sign up for our eNewsletters
Get the latest news and updates

Comment About the Article

To join the conversation, and become an exclusive member of Electronic Design, create an account today!