Beyond GPUs: What drives AI performance at scale

July 30 2026, by Macquarie Data Centres | Category: Data Centres
Beyond GPUs: What drives AI performance at scale

For the past three years, AI infrastructure investment has become a board-level priority for many organisations. IDC’s 2026 outlook puts global spending at US$497 billion this year, representing 56% year-on-year growth. By 2029, it expects the market to surpass US$1 trillion.

Yet businesses often assume that scaling AI architecture is simply a matter of adding more GPUs. It isn’t. An organisation can invest heavily in enterprise H200s, but if those units can’t communicate or keep AI workloads connected at scale, even the most powerful setup will be bottlenecked by the architecture holding it together.

AI performance depends on every layer working together: the GPU, server, storage layer, network fabric and core network. This guide explains how AI connectivity architecture works, why it determines real-world performance and what organisations should look for when running AI at scale. 

The buying criteria for an AI-ready data centre.

Before looking at how AI has changed data centre architecture, it helps to understand what a modern AI-ready data centre needs to deliver.

The following criteria support the density, connectivity and freedom of movement that AI workloads require:

  • High-density rack support: AI infrastructure can place significant power and cooling demands on each rack, particularly as GPU deployments scale.
  • Liquid cooling readiness: High-density AI hardware often requires more than traditional air cooling. Liquid cooling readiness should therefore be a priority.
  • Dense fibre pathways: AI workloads require fast data movement between servers, racks, storage and networks, especially as leaf-spine architecture becomes more common.
  • Carrier-neutral connectivity: Access to multiple carriers gives organisations the freedom to choose the network paths and partners that best suit their operations.
  • Low-latency interconnection: AI systems need high-performance connections to cloud platforms, data sources, users and, where required, other data centres.
  • Scalable power availability: AI demand can grow quickly, so the facility needs sufficient power capacity to support future expansion.
  • Secure or sovereign hosting: Sensitive AI use cases often need to meet local compliance, security or data sovereignty requirements.
  • Fungible infrastructure: AI requirements are changing quickly. A fungible data centre should be able to adapt as rack density, cooling, power and workload requirements evolve.
  • Proven support for hyperscale or AI workloads: Look for evidence that the facility can support high-density, highly connected environments in practice, not just in theory.

Ultimately, a modern data centre needs to be more than a facility for storing servers. It needs to be an AI and cloud-first environment built from the ground up for scalability. To understand why, we first need to look at how AI has changed the way data moves and what that means for data centre design. 

AI workloads change the data centre traffic model.

For decades, data centre traffic was relatively predictable. A user sent a request, the application retrieved the required information and a response was sent back. This created a neat north-south traffic pattern.

AI breaks that model.

A single inference prompt, such as “Which customers are likely to churn this quarter, and what should we do about it?”, can trigger work across hundreds of GPUs at once. The system might pull information from CRM and analytics platforms, retrieve support and billing histories, and run multiple model components in sequence before producing a single answer.

This turns traffic sideways. Instead of simply moving in and out of the data centre, data travels east-west: from server to server and GPU to GPU, often in large, bursty volumes.

The design challenge is no longer just about serving one request quickly. It is about keeping thousands of processors synchronised so that none of them are left idle while waiting for the network.

An east-west architecture only performs when every layer can keep pace. Even the most powerful GPUs will hit a wall if the data centre relies on interconnects that can’t support the rest of the workflow. And stalled GPUs are expensive. One idle week on a 512-GPU H100 cluster can represent an estimated US$80,000–$120,000 in stranded compute.

AI connectivity has to hold everything together.

Industrial AI workloads can’t be expected to flow through infrastructure designed for far simpler applications.

Supporting AI workloads at scale requires a connected architecture that begins inside the server and extends through the data hall fabric and into the broader network:

GPU layer: Each GPU server needs a fast path to CPUs, memory, local storage and network interface cards, preventing GPUs from sitting idle while waiting for information.

Cluster layer: Multiple GPU servers within the same rack need high-speed interconnects so they can communicate and share workloads efficiently.

Data hall layer: Racks throughout the data centre need dense network fabric and sufficient switching capacity to connect GPU clusters, applications and storage.

Core network layer: AI environments need reliable links to cloud platforms, enterprise systems, data sources and users.

The strength of this chain depends on whether each layer can pass work to the next without slowing down the entire environment. Here is how those layers work together.

GPU clusters scale AI beyond one server.

As AI workloads grow beyond what a single machine can handle, the challenge becomes keeping multiple systems connected, synchronised and aligned.

That connectivity begins at the rack level with GPU clusters. These environments connect multiple GPU servers through high-speed server-to-server interconnects, giving distributed computing systems far more capacity than any single GPU could provide.

Different GPU clusters put different demands on the network.

Cluster architecture begins with workload behaviour. The task being performed will determine how many GPUs need to be connected, how closely they need to synchronise and what the network must prioritise.

There are three main variations of an AI cluster network when classified by workload focus. Each places different demands on compute cluster connectivity and requires different network capabilities.

The three main types of GPU cluster architectures at a glance.

GPU clusterWhat does it mean?What does it do? What does the network need? 
Training clusterBuilding or fine-tuning AI models at scale using large datasetsWorks through large datasets and shares updates in real time as the model learnsUltra low latency + high bandwidth networking + tight synchronisation
Inference clusterRunning trained models (such as chatbots) for live users or applicationsHandles live requests and return outputs quickly and reliablyHigh throughput + fast response times
Hybrid cluster Supporting multiple workload types, like training and inference, in one environment Shifts between different AI workloads without one slowing down the others Flexibility + ability to isolate workloads + scalability

This is why attempting to stitch a GPU cluster onto an existing network is rarely enough. Just as organisations choose GPU infrastructure to support specific AI requirements, the data centre fabric needs to be designed around the workloads it will support.

At the data hall layer, network design matters.

As organisations scale their AI efforts to the cluster level, network fabric becomes the layer that can make or break performance.

It determines whether GPU servers can remain synchronised, move data quickly and keep workloads running without persistent bottlenecks.

More network options mean more architecture decisions.

InfiniBand is the standard network architecture for high-performance computing (HPC) clusters, but the arrival of Ethernet alternatives like RoCE v2 means businesses now have more options to consider: 

  • InfiniBand is purpose-built for ultra-low latency networks and high-bandwidth clusters. It’s the gold standard for complex AI training because of its immense scalability. 
  • RoCE v2 is built on a standard Ethernet network fabric. It’s fast and cost-effective, and it integrates with other vendors and hardware. However, it needs careful congestion control mechanisms to minimise bottlenecks and avoid packet loss. 

The available options are continuing to evolve. The Ultra Ethernet Consortium recently released its open-standard Ultra Ethernet specification as part of a broader effort to make Ethernet more suitable for AI at scale. Vendor-specific technologies, such as DriveNets Fabric-Scheduled Ethernet, are working towards a similar goal.

The point is not that one interconnect will be the right choice in every situation. It is that organisations can no longer treat the network as the “as long as it works” layer. Network architecture is now a strategic design decision that needs to be considered as carefully as any other part of the AI environment.

Pre-integrated AI infrastructure reduces the assembly problem.

The complexity of these design decisions is one reason pre-integrated AI architectures such as NVIDIA DGX SuperPOD are gaining attention.

These solutions bring GPU infrastructure, high-performance networking, storage and software together within a coordinated environment. This reduces some of the integration risk involved in building an AI architecture that needs to perform reliably in production.

Prebuilt architectures can simplify the assembly process, but they do not complete the entire AI environment.

Once workloads move into production, the architecture needs to extend beyond the data hall and connect with the wider business. That is the role of the core network.

The core network layer and the AI factory.

A GPU cluster can process a workload, and network fabric can keep data moving through the data hall. But AI only becomes useful when that internal environment can connect to the rest of the organisation.

At this point, the data centre begins to operate less like a hosting facility and more like an AI factory.

Raw data enters from the business. AI models process that data and turn it into something useful. The completed output is then delivered to the teams, applications and workflows that need it. In that sense, the environment acts like a manufacturing plant for intelligence.

The core network is the AI factory’s central nervous system.

The core network connects GPU clusters and the network fabric architecture to the systems and people that AI workloads depend on. This turns data centre AI infrastructure into a more complete production environment where AI workflows can link to: 

  • Storage systems and data platforms
  • Cloud environments 
  • Enterprise applications like CRMs and ERPs
  • Users, partners, carriers, and internet services
  • Other data centres through data centre interconnects (DCIs)

While the network fabric keeps workloads moving internally, the core network determines whether the workload can reach external systems. If it fails, inputs can’t get in, and outputs can’t get out. Then your AI factory turns into a remote island.  

So, how do you ensure this kind of connectivity is possible? Much like GPUs and network fabrics, organisations need to treat the core network as a critical component. It needs virtually no latency and strong interconnect throughput. Lossless network and deterministic performance under load are also essential. 

What this means for AI data centre design and selection.

Organisations investing in AI architecture can no longer rely on data centres that provide server hosting and call it a day. Modern AI requires a custom environment built for dense GPU infrastructure, complete connectivity, and data movement in real time.

Many AI data centres can support the GPU and cluster layer, but production AI often depends on what happens beyond the data hall itself. As such, data centre location, carrier access, cloud connectivity, and interconnection are now a vital part of the buying conversation.

A deeply interconnected, carrier-neutral data centre gives organisations more freedom to choose the network paths and connectivity partners that fit their workloads. It can also adapt as workloads and connectivity needs change over time. 

Build AI infrastructure that keeps performance moving.

AI performance isn’t determined by the GPUs you deploy. More often than not, it depends on the architecture that connects them. Think the network fabric inside the data hall, the core network beyond it, and the data centre environment that keeps workloads connected. 

The more transformative AI becomes, the more important it will be to build an infrastructure that can support it. The organisations that scale AI successfully will be the ones that treat compute, data centre fabric, and the core network as one interconnected environment. 

If you’d like to see what that integrated environment can look like, start with Macquarie Data Centres. Our Tier III certified data centres help enterprises and businesses build the deep infrastructure that powerful AI workflows depend on.

Contact us today, learn more about our team, or book a tour at our Sydney Data Centre to see how Macquarie Data Centres can support your AI infrastructure. 

FAQs.

What is AI data centre connectivity architecture?

Think of AI data centre connectivity architecture as the full chain of connections that make AI workloads work at scale. It refers to every component, from GPUs and racks to server interconnect architecture and networks to data halls, that sync together to handle the enormous scalability that AI demands. It combines a network fabric and core networking with server GPU cluster interconnect architecture and data centre design. 

What is the difference between a training cluster and an inference cluster?

A training cluster is used to build or fine-tune AI models using large datasets. An inference cluster runs trained models for live users, applications, or agents. Training usually needs tight synchronisation, while inference depends more on speed and packet delivery efficiency.

How do I know if my data centre can support AI workloads?

Look for high-density rack support, liquid cooling readiness, dense fibre pathways, carrier-neutral connectivity, and scalable power. The data centre should also support bandwidth capacity planning, data infrastructure scaling, and hybrid or cloud infrastructure design as AI workloads grow.

What is scale-out networking for AI clusters?

Scale-out networking is a way of connecting more servers, GPUs, and racks so they can work together as one cohesive system. Instead of relying on one big machine, organisations add more connected infrastructure and use high-speed network fabric to move workloads between them. 


About the author.

Macquarie Data Centres is Australia’s most trusted data centre provider. They house and protect the data for the world’s biggest hyperscalers, Global Fortune 500 companies and 42% of the Australian Federal Government. Part of the ASX-listed Macquarie Technology Group, they have been successfully building and operating data centres in Australia for over 20 years. Macquarie Data Centres currently owns and operates three data centres campuses, two in Sydney and one in Canberra, all of which are Certified Strategic by the Australian Government. Offering the confidence of a 100% uptime guarantee, their Tier III data centres provide the highest levels of security, sovereignty, service and compliance for their customers.

See all articles by this author

Get in touch.

1800 004 943

Enquiry Sent.

Thank you for contacting us. Our specialists will get in touch with you shortly.