Choosing hardware

How to write a GPU server specification

The workload, operating requirements and acceptance evidence needed to compare proposals before placing an order.

5 min readUpdated 6 October 2026
Editorial illustration

A useful technical specification connects hardware to a measurable business task. It lets you compare proposals on equal terms and decide in advance what successful acceptance means. A short brief covering the five areas below is enough for an initial discussion. Mark unknown parameters as questions to validate instead of filling them with assumptions.

1. Describe the task and identify the model

Start with the outcome: answering staff questions from a knowledge base, processing documents, recognising images or another specific operation. Separate inference, which runs an already trained model, from training or fine-tuning. They have different workload profiles; a configuration that serves answers does not establish that it can train the model.

Provide the model link, exact weight revision, format and licence restrictions. Include sample inputs and criteria for useful output. If the model is still being selected, list acceptable candidates and define a validation stage. “A server for AI” specifies neither the memory requirement nor how to judge the result.

2. Define demand and acceptable waiting time

Describe normal and peak demand separately: simultaneous requests, input and output lengths, and whether queuing is acceptable. Text-model context is measured in tokens, the pieces of text a model processes. An average user count is not enough to size a system without these details.

Set measurable targets: time to the first response fragment, request completion time, error rate and test duration. Keep single-conversation speed separate from aggregate throughput. Choose a realistic growth scenario and state which limits can change: context, concurrency, batch size or model quality.

3. Specify the whole server

Alongside the GPUs, list the runtime and driver, library and container versions. Check that the selected weights and computational operations are supported on the actual platform. If a model spans accelerators, agree the interconnect and partitioning method: equal total memory does not make multi-GPU systems behave identically.

ResourceWhat to specifyWhat to verify
CPUData preparation and supporting servicesWhether processing keeps the GPUs supplied
RAMModel loading, caches and permitted GPU offloadPeak usage during startup and operation
StorageWeights, data, logs, working copies and backupsCapacity, loading speed and recovery
NetworkAPI, storage and links between nodesThroughput and access arrangement

4. Agree operations and responsibilities

State the location, available power and cooling, rack requirements and network connection. Liquid cooling requires a separate site compatibility check. Identify who installs and updates the software, monitors the system and receives incident reports.

An SLA, or service-level agreement, should define specific commitments, responsibilities and measurement rules. Separate response time from restoration time. Specify maintenance windows, redundancy, data return and administrator access. Having several GPUs does not by itself make a service resilient to failure.

5. Turn requirements into an acceptance test

Prepare anonymised requests that resemble production: short and long inputs, difficult cases and the agreed peak. Record software versions, model settings and the measurement method. The report should cover output quality, latency, errors, peak memory, restart and recovery. Demonstrating one answer is not a substitute for this test.

Map every requirement to a proposed configuration, a test result or an open question. Put unverified conditions into a separate stage with an owner and a completion criterion. Before ordering, agree the supplied components, acceptance procedure, warranty and work needed at the site.

The deliverable is a testable chain: task → demand → configuration → acceptance. It connects the hardware choice to your workload and keeps an accelerator name from replacing evidence.

Checklist

Include these in your initial enquiry

Sources and documentation

These guides help you prepare requirements. The exact configuration and terms are set out in the quotation.