Scan Logo
Partner Logo

NVIDIA Rack Scale Clusters

NVIDIA Rack Scale Clusters

AI factory infrastructures for the most demanding agentic and physical AI challenges.

Turnkey AI Factories

Turnkey AI Factories

NVIDIA rack scale infrastructures are the de facto blueprint for the AI factory. Each rack unit contains 72 NVIDIA GPUs, coupled with 36 CPUs and unified memory combined in a Superchip architecture. These communicate via an advanced chip-to chip (C2C) link within each server node and via NVLink, NVSwitch and ConnectX networking across the rack, providing hugely scalable and groundbreaking compute performance.

Comprising compute nodes, an AI-ready software stack, management and orchestration tools and enterprise-level support, a Superchip rack scale solution is truly unprecedented.

As an NVIDIA Elite partner and the UK's only DGX Managed Service Provider, Scan can advise on the optimum configuration and install the complete infrastructure.

Enterprise AI Made Simple Diagram

AI Factories Made Simple

NVIDIA AI Enterprise, included with rack scale clusters, is an end-to-end software platform for building and deploying AI applications.

This collection of libraries, frameworks and NVIDIA Inference Microservices (NIMs) is specifically designed to reduce the complexity of building AI applications from scratch. Explore the range of optimised frameworks.

NVIDIA Rack Scale Generations

Find out more about the features of the different generations of Superchips in these systems.

Vera Rubin Architecture Features

Systems with Vera Rubin Superchips were announced in 2026 and are available to pre-order.

Driving Down Inferencing Costs

Driving Down Inferencing Costs

The Vera Rubin Superchip features a specialised multi-agent engine, and delivers one-tenth the cost per million tokens compared to Grace Blackwell for highly interactive, deep reasoning agentic AI.

Revolutionising Throughput

Revolutionising Throughput

The Vera Rubin Superchip delivers up to 10x more tokens per megawatt than Grace Blackwell, scaling intelligence within the same power footprint.

NVIDIA NVLink

Next-gen Connectivity

Sixth-generation NVLink and NVSwitch provide 28.8TB/s of bandwidth between GPUs, whilst ConnectX-9 SuperNICs and DPUs deliver networking at up to 800Gb/s.

Grace Blackwell Architecture Features

Systems with Grace Blackwell Superchips were announced in 2025 and are available now.

Revolutionising Throughput

Maximising Throughput

Grace Blackwell delivers a 10x boost in user responsiveness and a 5x improvement in throughput. Together, these translate into a remarkable 50x leap in overall AI factory output over the previous Hopper generation.

Dense NVFP4

Dense NVFP4

There are two types of Grace Blackwell superchips. The Grace Blackwell Ultra chip in the GB300 NVL72 are specially optimised for low precision calculations, supporting several proprietary NVIDIA formats such as dense NVFP4, which can boost performance by up to 50% more than the standard Grace Blackwell chips in the GB200 NVL72.

Dense NVFP4

Rapid Connectivity

Fifth-generation NVLink and NVSwitch provide 14.4TB/s of bandwidth between GPUs, whilst ConnectX-8 SuperNICs and DPUs deliver networking at up to 800Gb/s.

NVIDIA Rack Scale Cluster Platforms

Compare the specifications of the different Superchip systems.

Comparison of NVIDIA Vera Rubin NVL72, GB300 NVL72 and GB200 NVL72 specifications
Specification Vera Rubin NVL72 GB300 NVL72 GB200 NVL72
GPUs 72x Rubin 72x Blackwell Ultra 72x Blackwell
Cooling Liquid Liquid Liquid
GPU memory 20.7TB HBM4 20TB HBM3e 13.4TB HBM3e
GPU memory bandwidth 1,580TB/s 576TB/s 576TB/s
NVLink 6th gen 5th gen 5th gen
NVSwitch 6th gen 5th gen 5th gen
NVLink bandwidth 260TB/s 130TB/s 130TB/s
FP4 performance 3,600 PFLOPS 1,440 PFLOPS 1,440 PFLOPS
CPUs 36x Vera 36x Grace 36x Grace
System memory 54TB LPDDR5X 17TB LPDDR5X 17TB LPDDR5X
Networking NVIDIA ConnectX-9 VPI – up to 800Gb/s NVIDIA InfiniBand and Ethernet

NVIDIA BlueField-4 DPUs – up to 800Gb/s NVIDIA InfiniBand and Ethernet
NVIDIA ConnectX-8 VPI – up to 800Gb/s NVIDIA InfiniBand and Ethernet

NVIDIA BlueField-3 DPUs – up to 400Gb/s NVIDIA InfiniBand and Ethernet
NVIDIA ConnectX-7 VPI – up to 400Gb/s NVIDIA InfiniBand and Ethernet

NVIDIA BlueField-3 DPUs – up to 400Gb/s NVIDIA InfiniBand and Ethernet
Storage EDSFF E3.S SSDs EDSFF E1.S SSDs EDSFF E1.S SSDs
Software NVIDIA Mission Control
NVIDIA Base Command
NVIDIA UFM
NVIDIA NMX
NVIDIA Run:ai
NVIDIA Slurm
NVIDIA Mission Control
NVIDIA Base Command
NVIDIA UFM
NVIDIA NMX
NVIDIA Run:ai
NVIDIA Slurm
NVIDIA Mission Control
NVIDIA Base Command
NVIDIA UFM
NVIDIA NMX
NVIDIA Run:ai
NVIDIA Slurm
Power usage 190-230kW 132-155kW 120-132kW
Form factor 42U 42U 42U

Discover how the different DGX appliances and other platforms compare in our AI training and inferencing hardware buyers guide.

Data Centre

Deploying your Rack Scale Cluster

Scan is proud to be the UK's only NVIDIA DGX-Ready Managed Services Provider, offering in-country and global customer support. Our managed service portfolio offers a one-stop shop for AI factory customers deploying, managing, and maintaining AI supercomputing infrastructure.

Clusters have much more demanding power and cooling requirements than conventional servers. Our technical specialists will perform a site survey to see whether your facilities are suitable or can be adapted. This includes evaluating liquid cooling options, alongside differing power delivery methods, such as integrated PSUs versus busbars. Alternatively, we can arrange installation at one of our NVIDIA-approved hosting partners.

Speak to an Expert

Discuss your DGX project requirements with Scan's specialist team.

Frequently Asked Questions

Find answers to common questions about NVIDIA Superchip servers, rack scale clusters and AI factories.

A Superchip server is a SoC (system on chip) architecture where the CPU, GPU and memory are combined on a single chip, rather than separate components.

A cluster is a collection of connected server nodes within a rack environment. It is designed from the ground up as a single architecture, with each server node connected to the next, rather than separate servers simply placed in the same rack.

Whereas several connected servers can deliver enhanced performance by acting as a single coherent unit, a cluster is designed from the ground up to remove as many bottlenecks and latencies as possible, so the resulting unified architecture will outperform the equivalent number of connected servers.

AI factories are the next step for the datacentre. A normal datacentre stores data. An AI factory turns that data into real-time insight; it's intelligence, not data, that comes out the other end.

It does this by managing every stage of an AI project: data preparation, training, fine-tuning, and inference. Fine-tuning and inference matter most, because that's where the intelligence gets made.

Think of an AI factory as a manufacturing plant for intelligence. Raw data goes in. Insights, trained models and real-time predictions come out. Unlike a general-purpose datacentre, it's built to continuously train and deploy AI models, turning data into useful intelligence at scale.

Any organisation that wants to scale AI use efficiently and responsibly can benefit from an AI factory, including:

  • Enterprises managing large datasets and multiple AI projects
  • Healthcare providers using AI for diagnostics or patient management
  • Financial institutions deploying AI for fraud detection or risk modelling
  • Retailers and e-commerce companies focused on personalisation and demand forecasting
  • Manufacturers aiming to automate processes or predict maintenance
  • Government and research bodies working with sensitive or high-volume data

AI factories offer several benefits that help businesses mature their AI use:

  • Faster results – automated data prep, training and deployment cut lead times
  • Scalability – reuse pipelines and models across teams and use cases
  • Consistency – standardised workflows reduce errors and technical debt
  • Better ROI – lower operational costs, stronger model performance
  • Governance and compliance – centralised control over data usage, lineage and versioning
  • Adaptability – works with cloud, hybrid or on-premise environments

Yes. Small businesses can use cloud platforms with pre-built templates and low-code tools, making it easier and more affordable to deploy AI workflows without building complex infrastructure from scratch. Contact our Scan Cloud team for a readiness check.

An AI factory helps your business scale AI by systematising the model development pipeline, allowing teams to reuse components, reduce development time, ensure consistency, and deploy models faster. This scalability is essential for companies looking to integrate AI into multiple departments or products.

Yes, AI factories can integrate across environments. Hybrid and multi-cloud setups allow companies to process data where it lives – whether in the cloud, at the edge (e.g., IoT devices), or in secure on-premises systems – always ensuring flexibility and compliance. Contact our AI team for your readiness check today.

Yes. AI factories can support real-time AI inference and batch processing simultaneously, using intelligent workload management to balance low-latency applications with large-scale AI workloads while maximising GPU utilisation and performance.

Yes, as a leading NVIDIA Elite partner and the UK's only certified NVIDIA DGX Managed Services Provider (MSP), Scan sells all the individual elements to create your own AI factory. This includes NVIDIA DGX and NVIDIA Run:ai software licences. We also offer a range of professional services for ease of deployment and management.

Migrating to an AI factory model is a strategic move that requires planning, the right tools, and experienced support. Scan acts as your specialist partner to streamline this transformation, bringing technical expertise, cutting-edge infrastructure, and industry-proven workflows.

Here's how the migration journey typically unfolds with Scan by your side:

  • Discovery & assessment – starting with a full audit across the whole business that identifies bottlenecks in data prep, model development, and deployment
  • AI factory blueprint design – our teams work with you to design a scalable and secure AI factory architecture tailored to your goals, including standardised data ingestion pipelines, reusable model training and deployment workflows, governance and monitoring frameworks, and on-premise, cloud, or hybrid infrastructure integration
  • Infrastructure & platform setup – choose infrastructure and tools that support scalability, are built for performance, security and long-term growth
  • Automation and MLOps integration – we help operationalise your AI workflows with best-in-class MLOps practices, introducing version control, CI/CD pipelines, model monitoring, and retraining strategies to enable continuous delivery of AI models across teams and business units
  • Team enablement – Scan not only builds the factory, we empower your team to run it. Alongside NVIDIA, we can provide documentation, training, and ongoing support to ensure your data scientists, engineers, and decision-makers can maintain, scale, and improve the system confidently
  • Ongoing optimisation & support – your AI factory isn't a static solution. Scan offers long-term support and optimisation services, helping you manage compute resources efficiently, track carbon usage, refine models, and adopt emerging AI tools