Infrastructure & architecture practice

Sovereign AI infrastructure sizing calculator

Work out what it actually takes to serve an open-weight model from Canadian infrastructure — GPU count, memory headroom, latency, topology and cost. Every figure below shows the arithmetic that produced it, so you can check it rather than trust it.

Workload

Presets fill the fields below and are approximate. Verify against the model card; every value stays editable.

Differs only for MoE. Memory holds all; compute reads active.

GQA models have far fewer than attention heads. This drives KV cache size.


Batch size. Raising it lifts throughput and worsens per-stream latency.

Time to first token — prefill, compute-bound.

Decode step time — memory-bandwidth-bound.


Verify against a current quote — list prices move and vary by commitment.

Efficiency assumptions

0.6–0.8 is realistic for a tuned vLLM / TensorRT-LLM deployment.

Recommended configuration

—

—

—
Tensor parallel degree
—
Replicas
—
8-GPU nodes

Memory

Weights—
KV cache at full concurrency—
Activations & runtime overhead—
Total per replica—
GPUs required to fit—
Concurrent sequences this memory allows—

—

Latency & throughput

Time to first token—
Inter-token latency—
Per-stream output rate—
Decode throughput per replica—
Concurrency implied by arrival rate—
Aggregate output tokens/sec—

—

Topology

—

Network topology

—

—

Cost & Canadian availability

—

—

What this cannot know

These figures are a defensible starting point, not a bill of materials. The variables that move a real deployment by ±40% are all things a form cannot ask: your actual token-length distribution rather than an average, prefix-cache hit rate, how bursty arrivals are against your tail-latency target, scheduler and chunked-prefill configuration, speculative decoding, tolerance for quality loss under quantisation, and the failure domain you need to survive.

We size clusters against real traces rather than averages, and design the fabric, segmentation and residency controls underneath them.

Method

How these numbers are derived

Memory

Weights are parameters × bytes per parameter. KV cache per token is 2 × layers × KV heads × head dim × bytes — the factor of two being K and V. Total cache scales with concurrency and sequence length, which is why it, not the weights, is usually what actually runs you out of memory at long context.

Decode

Token generation is bound by memory bandwidth, not arithmetic. Each forward pass streams the weights once plus the KV cache for every sequence in the batch, so step time is bytes read ÷ (bandwidth × utilisation). Batching amortises the weight read across concurrent requests, which is why throughput climbs steeply with batch size until the point where it goes compute-bound.

Prefill

Time to first token is compute-bound: roughly 2 × active parameters × input tokens FLOPs against the GPU's dense throughput, divided across the tensor-parallel group. Long prompts hurt TTFT; they do not affect inter-token latency.

Topology

Tensor parallelism inside a node runs over NVLink. Replicas are independent and exchange nothing. So for most inference the inter-node fabric carries no traffic on the critical path and its choice is irrelevant — a claim worth making precisely because so much marketing implies otherwise. It changes only when the model is too large to shard within one node.

GPU reference figures

GPU Memory Bandwidth BF16 dense Board power Intra-node link

Vendor-published figures, shown so the arithmetic above can be audited. Dense BF16 throughput is quoted — not the doubled sparsity number that marketing material usually leads with.