Reserved GPUs.
For the stack you run.
NVIDIA RTX PRO 6000 Blackwell with 96 GB per GPU, reserved for you in Germany and billed per GPU per month. You bring the models, training jobs and containers. We run the hardware and the network.
Available on request. We confirm capacity and start date for each inquiry.
Yours alone.
Your GPUs are reserved for you for the whole contract and exclusive to you: no other customer runs workloads on them, and a training run never waits for free capacity.
A fixed monthly line.
Billed per GPU per month. The bill does not depend on how many hours or tokens you use.
As a VM or a Kubernetes node.
Take the GPUs as VMs with root access, or as GPU nodes in your enum Kubernetes cluster next to your object storage. Both on request.
One GPU type.
Sized for open-weight models.
We offer one GPU: the NVIDIA RTX PRO 6000 Blackwell. With 96 GB per card, gpt-oss-120b, Mistral Small and Qwen3 32B each run on a single GPU.
| Size | GPU memory | Price |
|---|---|---|
| 1 GPU | 96 GB | On request |
| 2 GPUs | 192 GB | On request |
| 4 GPUs | 384 GB | On request |
| 8 GPUs | 768 GB | On request |
| More than 8 | On request | On request |
Indicative sizes, each as a VM or as GPU nodes in your enum Kubernetes cluster, all with NVIDIA RTX PRO 6000 Blackwell (96 GB GDDR7 with ECC). Term, capacity and start date are confirmed per inquiry. Pricing on request.
Your stack.
On a VM or on Kubernetes.
Take the GPUs as VMs with root access and install what you like, or as a GPU node pool in your enum Kubernetes cluster. There, workloads request them through the standard nvidia.com/gpu resource, so vLLM, Ollama, PyTorch jobs and your Helm charts run without changes. Both are available on request, and we set them up with your team.
# vllm.yaml (excerpt)
containers:
- name: vllm
image: vllm/vllm-openai
args: ["--model", "openai/gpt-oss-120b"]
resources:
limits:
nvidia.com/gpu: 1What teams run on them.
Their own models, their own way.
Inference you operate.
Serve open-weight or in-house models with vLLM, TGI or Triton, with your own scaling and monitoring.
Fine-tuning.
Adapt open-weight models to your data with LoRA, or fine-tune smaller models fully. Training data stays in your object storage in Germany.
Batch jobs.
Embeddings for large document sets, transcription, classification: work that runs for hours on a GPU that stays yours.
The hardware.
In numbers.
Contract and compliance.
Where we stand.
Send an inquiry
We confirm GPU count, capacity and start date after a short scoping call. Pricing on request.
Questions.
About reserved GPUs.
Which GPUs does enum offer?
The NVIDIA RTX PRO 6000 Blackwell with 96 GB of GDDR7 memory per GPU. You reserve them from one GPU upward.
How soon can we start?
Reserved GPUs are available on request. Tell us how many GPUs you need and from when, and we confirm capacity and a start date for your inquiry.
What does it cost?
You pay per GPU per month, in euros. We quote the price together with capacity and start date, so pricing is on request.
Is this dedicated hardware?
Your GPUs are reserved for you for the term of your contract and are exclusive to you: no other customer runs workloads on them. The network, storage and Kubernetes control planes around them run on enum's shared public cloud platform.
How do we use the GPUs?
As GPU VMs or as GPU nodes in your enum Kubernetes cluster, both on request. On a VM you have root access and install your own stack. On Kubernetes, pods request the GPUs through nvidia.com/gpu, so you deploy your inference server or training job like any other workload.
Which models fit on one GPU?
With 96 GB, gpt-oss-120b (MXFP4), gpt-oss-20b, Mistral Small 24B and Qwen3 32B each fit on one GPU. Llama 3.3 70B fits on one GPU in FP8 with limited context; for longer context or more users, plan two. These figures are indicative, because context length and concurrency change the sizing.
Where are the GPUs?
In Frankfurt. The data center is certified to ISO 27001 and EN 50600; those certificates belong to the facility operator. enum GmbH is a German company with no US parent, so the US CLOUD Act does not reach the data on them.
Is there a minimum term?
We agree the term per inquiry. Tell us how long you need the GPUs, and the quote covers that term.
Can enum run the models for us instead?
Yes. With Models as a Service, enum runs the serving stack on the same reserved GPUs, with a private OpenAI-compatible API and the same per-GPU billing.
Need GPUs in Germany?
Tell us what you run.
Send us the GPU count, your timeline and what you plan to run. We confirm capacity, a start date and a price.