AI performance isn’t decided by your GPUs
For three AI infrastructure investment has become a board-level priority for many organisations. IDC’s 2026 outlook puts global spending at US$497 billion this year, a 56% YoY growth. By 2029, they expect the market to surpass US$1 trillion.
Businesses often believe that scaling AI architecture is a challenge best solved by adding more GPUs. That’s the wrong impression. Businesses can invest every dollar they have into enterprise H200s, but if those units can’t communicate and keep AI workloads connected at scale, even the most powerful setup will be bottlenecked by the architecture holding it together.
AI performance depends on how every layer works together: GPU, server, storage layer, network fabric, and core network. This guide explains how AI connectivity architecture works, why it decides real-world performance, and what to look for to run AI at scale.
The buying criteria for an AI-ready data centre
Before we examine how AI has changed data centre architecture, it helps to have a grasp on what the modern AI-ready data centre actually needs to deliver.
These are the criteria for a system that can support the density, connectivity, and freedom of movement that AI workloads require:
- High-density rack support: AI infrastructure can place heavy power and cooling demands on each rack, especially as GPU deployments scale.
- Liquid cooling readiness: High-density AI hardware often needs more than traditional air cooling. Go for liquid cooling as a priority.
- Dense fibre pathways: AI workloads need fast movement between servers, racks, storage, and networks, especially as leaf-spine architecture becomes the norm.
- Carrier-neutral connectivity: Access to multiple carriers gives organisations more freedom to choose the network paths and partners that fit their operation.
- Low-latency interconnection: AI systems need high-performance routing to cloud platforms, data sources, and users, as well as other data centres as needed.
- Scalable power availability: AI demand can grow quickly, so the facility needs enough power capacity to support expansion.
- Secure or sovereign hosting: Sensitive AI use cases often need to meet local requirements for compliance or sovereignty.
- Fungible infrastructure: AI requirements are changing quickly. The data centre should be able to adapt as rack density, cooling, power, and workload needs evolve.
- Proven support for hyperscale or AI workloads: Look for evidence that the facility supports high-density, high-connectivity environments in practice.
Ultimately, data centres today need to be less a server storage facility and more an AI factory built from the ground up for scalability. To see why that matters, we need to go back a step and look at how AI has changed the way data moves and how that transforms data centre demands.
AI workloads change the data centre traffic model
For decades, data centre traffic was predictable. A user sent a request, the application pulled the right information, and the response went back: a neat and tidy north-south pattern.
AI breaks that model. A single inference prompt like “which customers are likely to churn this quarter, and what do we do about it?” can trigger work across hundreds of GPUs at once, pulling from CRMs and analytics platforms, retrieving support and billing history, and running multiple model components in sequence before a single answer is produced.
That turns traffic sideways. Instead of flowing in and out, data now moves east-west. Server to server, GPU to GPU, in huge, bursty volumes. The design challenge is no longer serving a request quickly; it’s keeping thousands of processors synchronised so none of them sit idle waiting on the network.
An east-west architecture only delivers when every layer keeps pace. Even the most powerful GPUs will hit a wall and sit idle when the data centre is limited by interconnects that can’t keep up with the rest of the workflow. And a stalled GPU isn’t cheap. One idle week on a 512-GPU H100 cluster can cost an estimated US$80,000–$120,000 in stranded compute.
AI connectivity has to hold everything together
You can’t expect industrial AI workloads to flow like water through home plumbing. Supporting AI workloads at scale means building a connected architecture that starts inside the server and extends through the data hall fabric and broader network:
- GPU layer: Each GPU server needs a fast path to CPUs, memory, local storage, and network interface cards so they aren’t left idle and waiting for information.
- Cluster layer: Multiple GPU servers in the same rack need high-speed interconnects so they can communicate and share workloads.
- Data hall layer: Racks inside the data centre need dense network fabric and switching capacity to connect GPU clusters and applications.
- Core network layer: The AI environments need reliable links to cloud platforms and enterprise systems in addition to data sources and users.
How strong data centres make this chain depends on whether each layer can pass work to the next without slowing everything down. Let’s take a look at how it works:
GPU clusters scale AI beyond one server
As AI workloads scale beyond what a single machine can handle, the new challenge is ensuring systems can communicate and stay aligned.
That kind of connectivity starts at rack level with GPU clusters. These environments connect multiple GPU servers via high-speed server-to-server interconnects, giving distributed computing more capacity than any single GPU could provide. .
Different GPU clusters put different demands on the network
Cluster architecture starts with workload behaviour. The job that the AI is being asked to do will determine how many GPUs need to be connected and how closely they synchronise.
There are three main variations of an AI cluster network when classified by workload focus. Each comes with its own level of compute cluster connectivity and network priorities to support a core AI workflow.
The three main types of GPU cluster architectures at a glance
| GPU cluster | What does it mean? | What does it do? | What does the network need? |
| Training cluster | Building or fine-tuning AI models at scale using large datasets | Works through large datasets and shares updates in real time as the model learns | Ultra low latency + high bandwidth networking + tight synchronisation |
| Inference cluster | Running trained models (such as chatbots) for live users or applications | Handles live requests and return outputs quickly and reliably | High throughput + fast response times |
| Hybrid cluster | Supporting multiple workload types, like training and inference, in one environment | Shifts between different AI workloads without one slowing down the others | Flexibility + ability to isolate workloads + scalability |
This is why trying to stitch together a GPU cluster with an already-built network won’t suffice. Just as organisations choose GPU infrastructure to support AI, the data centre fabric now needs to be carefully designed around the workload it’s going to support.
At the data hall layer, network design matters
As businesses scale their AI efforts to cluster scale, network fabric becomes the layer that makes or breaks performance. It determines whether GPU servers can stay synchronised, shift data quickly, and keep workloads moving without constant bottlenecks.
More network options mean more architecture decisions
InfiniBand is the standard network architecture for high-performance computing (HPC) clusters, but the arrival of Ethernet alternatives like RoCE v2 means businesses now have more options to consider:
- InfiniBand is purpose-built for ultra-low latency networks and high-bandwidth clusters. It’s the gold standard for complex AI training because of its immense scalability.
- RoCE v2 is built on a standard Ethernet network fabric. It’s fast and cost-effective, and it integrates with other vendors and hardware. However, it needs careful congestion control mechanisms to minimise bottlenecks and avoid packet loss.
The possibilities are evolving all the time, too. The UEC recently released its open-standard Ultra Ethernet as part of a wider movement to make Ethernet more suitable for AI at scale. Vendor-specific versions, like DriveNets Fabric-Scheduled Ethernet, also have a similar goal.
The message here isn’t that one interconnect wins every time, but rather that organisations can no longer treat networks as the “as long as it works” layer. It’s now a design decision that needs to be as carefully considered as any other AI component.
Pre-integrated AI infrastructure reduces the assembly problem
The complexity of making those design decisions is one reason pre-integrated AI architectures like NVIDIA DGX SuperPod are gaining attention. These solutions bring GPU infrastructure and high-performance networking, storage, and software into one coordinated environment. That reduces the risk of building something that crumbles in production.
Prebuilt architectures simplify assembly. Even so, they don’t complete the entire AI environment. Once workloads move into production, the architecture needs to move beyond the data hall and into the wider business. This is the responsibility of the core network.
The core network layer and the AI factory
A GPU cluster can process the workload, and the network fabric can keep the data flowing through the data hall, but even so, AI is only useful when that internal environment can connect to the rest of the business.
At that point, the data centre starts to behave less like a host and more like an AI factory: Raw data comes in from the business, AI models process it and turn it into something useful, and the finished output gets sent back to teams and workflows that need it. Think of it like a manufacturing plant for intelligence.
The core network is the AI factory’s central nervous system
The core network connects GPU clusters and the network fabric architecture to the systems and people that AI workloads depend on. This turns data centre AI infrastructure into a more complete production environment where AI workflows can link to:
- Storage systems and data platforms
- Cloud environments
- Enterprise applications like CRMs and ERPs
- Users, partners, carriers, and internet services
- Other data centres through data centre interconnects (DCIs)
While the network fabric keeps workloads moving internally, the core network determines whether the workload can reach external systems. If it fails, inputs can’t get in, and outputs can’t get out. Then your AI factory turns into a remote island.
So, how do you ensure this kind of connectivity is possible? Much like GPUs and network fabrics, organisations need to treat the core network as a critical component. It needs virtually no latency and strong interconnect throughput. Lossless network and deterministic performance under load are also essential.
What this means for AI data centre design and selection
Organisations investing in AI architecture can no longer rely on data centres that provide server hosting and call it a day. Modern AI requires a custom environment built for dense GPU infrastructure, complete connectivity, and data movement in real time.
Many AI data centres can support the GPU and cluster layer, but production AI often depends on what happens beyond the data hall itself. As such, data centre location, carrier access, cloud connectivity, and interconnection are now a vital part of the buying conversation.
A deeply interconnected, carrier-neutral data centre gives organisations more freedom to choose the network paths and connectivity partners that fit their workloads. It can also adapt as workloads and connectivity needs change over time.
Build AI infrastructure that keeps performance moving
AI performance isn’t determined by the GPUs you deploy. More often than not, it depends on the architecture that connects them. Think the network fabric inside the data hall, the core network beyond it, and the data centre environment that keeps workloads connected.
The more transformative AI becomes, the more important it will be to build an infrastructure that can support it. The organisations that scale AI successfully will be the ones that treat compute, data centre fabric, and the core network as one interconnected environment.
If you’d like to see what that integrated environment can look like, start with Macquarie Data Centres. Our Tier III certified data centres help enterprises and businesses build the deep infrastructure that powerful AI workflows depend on.
Contact us today, learn more about our team, or book a tour at our Sydney Data Centre to see how Macquarie Data Centres can support your AI infrastructure.
FAQs
What is AI data centre connectivity architecture?
Think of AI data centre connectivity architecture as the full chain of connections that make AI workloads work at scale. It refers to every component, from GPUs and racks to server interconnect architecture and networks to data halls, that sync together to handle the enormous scalability that AI demands. It combines a network fabric and core networking with server GPU cluster interconnect architecture and data centre design.
What is the difference between a training cluster and an inference cluster?
A training cluster is used to build or fine-tune AI models using large datasets. An inference cluster runs trained models for live users, applications, or agents. Training usually needs tight synchronisation, while inference depends more on speed and packet delivery efficiency.
How do I know if my data centre can support AI workloads?
Look for high-density rack support, liquid cooling readiness, dense fibre pathways, carrier-neutral connectivity, and scalable power. The data centre should also support bandwidth capacity planning, data infrastructure scaling, and hybrid or cloud infrastructure design as AI workloads grow.
What is scale-out networking for AI clusters?
Scale-out networking is a way of connecting more servers, GPUs, and racks so they can work together as one cohesive system. Instead of relying on one big machine, organisations add more connected infrastructure and use high-speed network fabric to move workloads between them.