AI infrastructure / Kazakhstan

GPU servers.
Without limits
for your ideas.

GPU servers for businesses in Kazakhstan.
NVIDIA HGX / NVL · AMD Instinct · Huawei Ascend
Purchase, rental and deployment planning.

OMNAR LLP / KAZAKHSTAN / B2B
OMNAR / COMPUTE
01 — 08
Total GPU memory2,160 GB
Accelerators08 GPU
ECOSYSTEMS & WORKLOADSNVIDIA HGXAMD InstinctHuawei AscendLLM / INFERENCETRAINING / HPC
01 / Hardware

Power for your scale.

Compare platforms and find the right configuration for your workload.

Specifications verified: 06.10.2026
13 / 13
NVIDIABlackwell Ultra
Supermicro SYS-822GS-NB3RT · system example

HGX B300 · 8 GPU

Eight Blackwell Ultra GPUs with NVLink 5 Switch. Configuration and cooling are agreed with the server manufacturer.

Accelerator memory
2,160 GB
Architecture and software
SXM6 · NVLink 5 SwitchCUDA
from $634,046
NVIDIABlackwell
Supermicro SYS-822GS-NBRT · system example

HGX B200 · 8 GPU

Eight Blackwell GPUs in one NVLink domain. The server and cooling are selected for the infrastructure.

Accelerator memory
1,440 GB
Architecture and software
SXM6 · NVLink 5 SwitchCUDA
from $480,285
NVIDIAFor large models
Supermicro SYS-821GE-TNHR · system example

HGX H200 · 8 GPU

Training and inference across eight accelerators with high-speed interconnects.

Accelerator memory
1,128 GB
Architecture and software
SXM5 · NVSwitchCUDA
from $404,504
NVIDIAFlexible platform
Supermicro SYS-422GA-NRT · system example

H200 NVL · 8 GPU

Eight PCIe GPUs in two NVLink domains of four GPUs each. Server and bridge compatibility is confirmed in the quotation.

Accelerator memory
1,128 GB
Architecture and software
PCIe · 2 × NVLinkCUDA
from $469,168
AMD2.3 TB HBM3e
Supermicro AS-A126GS-TNMR · system example

Instinct MI355X · 8 GPU

Eight MI355X GPUs. Cooling, power limits and site requirements depend on the server and are agreed in the quotation.

Accelerator memory
2,304 GB
Architecture and software
OAM · Infinity FabricROCm
from $402,824
AMD2 TB HBM3e
Supermicro AS-8126GS-TNMR · system example

Instinct MI325X · 8 GPU

Eight accelerators for memory-intensive workloads.

Accelerator memory
2,048 GB
Architecture and software
OAM · Infinity FabricROCm
Request a quote
Showing: 6 / 13

Images show examples of complete systems for each platform. The exact chassis, installed accelerator count and configuration are confirmed in the quotation.

Prices are budget estimates, not binding offers or confirmation of stock. VAT, exact specifications, delivery, warranty and lead time are set out in the quotation. Accelerator memory is a total; usable capacity depends on software and interconnect topology.

Platform finder

Choose a model. Compare memory and cost.

Inference only

Platforms 13

Filters
Platform type
GLM-5.3 · FP8 / BF16 / F32 · Reserve: 128 GB · Memory fit: smallest sufficient first
PlatformMemoryMemory fitDiscuss this setup
H200 NVL · 8 GPUNVIDIA · CUDAfrom $469,1681,128 GBCapacity fitsremaining 244.4 GB
HGX H200 · 8 GPUNVIDIA · CUDAfrom $404,5041,128 GBCapacity fitsremaining 244.4 GB
HGX B200 · 8 GPUNVIDIA · CUDAfrom $480,2851,440 GBCapacity fitsremaining 556.4 GB
Instinct MI300X · 8 GPUAMD · ROCmfrom $356,1681,536 GBCapacity fitsremaining 652.4 GB
Instinct MI325X · 8 GPUAMD · ROCmRequest a quote2,048 GBCapacity fitsremaining 1,164.4 GB
HGX B300 · 8 GPUNVIDIA · CUDAfrom $634,0462,160 GBCapacity fitsremaining 1,276.4 GB
Showing 6/13

Nominal memory indicates capacity. Configuration, software, topology and availability require separate validation.

Source check:
How to estimate LLM memory
Calculation & sources

GLM-5.3

753 B — checkpoint parameters. 39 B — active per token.

Published checkpoint: FP8 / BF16 / F32. Weight payload from the published checkpoint or official deployment guide. This is stored payload, not measured GPU allocation; runtime conversion can increase it. Operating reserve is added separately.

Weight format: Start with the published checkpoint. Extra quantization may affect quality; converting quantized weights to BF16 does not restore lost precision.

Operating reserve: Memory beyond weights for KV cache, activations, runtime buffers and headroom across the selected configuration, not per GPU. Set it from context length and concurrent requests, then measure peak use. The presets are examples, not universal recommendations.

Model card ↗Published weight payload ↗Published deployment guide ↗

This is the limit declared in the model configuration, not a context verified on the selected hardware.

DSA uses kv_lora_rank=512 and full/shared indexer types. Cache sizing requires an architecture-specific calculation.

Model config ↗

Context config checked: .

Runtime validation required. Published deployment guide: HGX H200 · 8 GPU, Instinct MI355X · 8 GPU, Instinct MI325X · 8 GPU, Instinct MI300X · 8 GPU.

How to read this estimate

Weights = total parameters × bytes per parameter: BF16 2, FP8 1, INT4 0.5. Decimal GB. MoE models count every expert, not only active parameters. This is a lower bound: quantization scales, unquantized layers and metadata are not separately modelled; they must fit within the reserve.

A green result is a capacity estimate, not a tested deployment. GPU memory is distributed across devices. Tensor and expert parallelism, interconnects, model kernels and framework support must all be checked. Long context and parallel users can need much more reserve. Fine-tuning and training require a separate calculation.

KV cache depends on context length, concurrent requests, cache precision and the model’s attention architecture. Sliding-window, linear-attention and compressed-cache models need architecture-specific estimates; validate the peak in the selected runtime. This field changes the capacity estimate, not server settings.

Hugging Face: KV cache strategies ↗vLLM: Hybrid cache allocation ↗
  • Capacity fits. The estimated memory budget fits. Validate software, sharding and performance.
  • Limited headroom. Under 10% free at this reserve. Validate memory use before choosing.
  • Exceeds capacity. This scenario exceeds total GPU memory. Change precision or platform.

BF16 / FP8 / INT4 are size scenarios. The format names do not assert native execution support or equal model quality.

NVL topology and sharding require separate validation; an 8-GPU PCIe server is not the same fabric as HGX/NVSwitch. Validate the exact checkpoint and kernels in ROCm before deployment.

Huawei Ascend uses CANN. Capacity results do not confirm that the selected checkpoint, operators or precision run on these stacks. A model port and runtime validation are required.

The calculation uses 270 GB per GPU for HGX B300 and 279 GB per GPU for GB300, following the NVIDIA datasheet. The 288 GB figure is not used. CPU RAM is not added to GPU memory. NVIDIA Blackwell Ultra (PDF) ↗

02 / Infrastructure rental

Focus on results.
You don’t have to
own the server.

Let’s define a dedicated configuration for your workload. Rental term, hosting, access and support are agreed before launch.

Get a rental quote↗
01
LLM & AI inferenceModels, RAG and enterprise AI services
02
Training & fine-tuningMemory, interconnect and storage planning
03
Computer vision & HPCCompute matched to your workload
03 / How we work

From workload
to configuration.

From the first specification to equipment handover. We work with companies in Kazakhstan and select infrastructure for their workloads.

01 /

Understand the workload

Model, precision, memory, user count and performance targets.

02 /

Compare the options

NVIDIA, AMD or Huawei. We compare memory, accelerator interconnects, software compatibility and total cost of ownership.

03 /

Agree the delivery

Specifications, hardware checks, contract, invoice payment and warranty terms.

Supply across Kazakhstan

Global technology.
A local partner.

OMNAR is a Kazakhstan-based GPU server supplier. We select equipment for companies in Kazakhstan and prepare purchase and dedicated rental proposals.

01 / CONFIGUREProjectWorkload and configuration
02 / OPERATEDeploymentDelivery and acceptance
Specification → contract → delivery → acceptance
01

Contract with Omnar LLP

A Kazakhstan legal entity, quotations in tenge and payment against an invoice.

02

Transparent pricing

Equipment, delivery, taxes and additional services are itemised in the quotation.

03

Your infrastructure model

A server at your facility or a dedicated rental with agreed hosting and access.

Start with your workload.

Three starting points for planning your infrastructure.

01 /

An AI service for your team

Define the model, context length, concurrent users and acceptable response time.

What to prepare
02 /

Training and experiments

Describe the data, fine-tuning methods and run duration. Compare buying and renting for the same workload.

What to prepare
03 /

Your own infrastructure

Check the site, power, cooling and network before committing to servers or a full rack.

What to prepare

Answers that help you decide.

From the first memory estimate to server acceptance, the detail lives in dedicated guides.

All guides →
Let’s get started

What are you
building next?

Tell us about your model and expected workload. Send a server purchase or rental enquiry with the details we need to prepare a quotation.

AI infrastructure specialist at a laptop
Salessales@omnar.kzEmail is being prepared for launch. You can send an enquiry using the form.

All fields are required.

Consent does not include marketing messages. Sending a request does not place an order or commit you to a purchase.

Selecting “Send enquiry” sends your details and consent to OMNAR. Do not include identity numbers, documents, passwords or information unrelated to the enquiry.

Good to know

Before your first deployment.

How do I choose between NVIDIA, AMD and Huawei?

Start with your model and software stack: CUDA for NVIDIA, ROCm for AMD and CANN for Huawei. Then verify weight formats, memory placement and performance on your workload. Equal memory capacity does not make platforms interchangeable.

What do the catalog prices mean?

Pricing depends on the configuration and delivery terms. Request a quotation with the final price in KZT, VAT details, delivery lead time and warranty. Availability is confirmed in the quotation.

How does 8 × H200 NVL differ from HGX?

HGX H200 uses an SXM platform with NVSwitch. A PCIe build with H200 NVL needs compatibility checks and may consist of two groups of four GPUs. This is not equivalent to the eight-GPU HGX interconnect fabric.

How is rental arranged?

Configuration, term, data center, connectivity, access and support are agreed individually. Capacity and start dates must be confirmed; instant provisioning is not promised.