---
canonical: https://enum.co/models-as-a-service
locale: en
---

# Private model endpoints. On GPUs reserved for you.

enum Models as a Service is billed per GPU per month, not per token. Customers reserve NVIDIA RTX PRO 6000 Blackwell GPUs (96 GB) in Germany that are exclusive to them, and enum runs open-weight models such as gpt-oss, Llama and Mistral, or the customer's own weights, on them behind a private OpenAI-compatible API. Tokens are not billed: the limit is what the reserved GPUs can process. Prompts and outputs are not stored by default and never used for training. enum GmbH is a German company with no US parent. Available on request: capacity and start date are confirmed per inquiry, and pricing is on request.

Models as a Service at enum is billed per GPU per month, not per token. We run gpt-oss, Llama, Mistral or your own weights on GPUs reserved for you in Germany, and you call them through a private OpenAI-compatible API. Available on request. We confirm capacity and start date for each inquiry.

- **No per-token bill.** You pay per GPU per month. Send as many tokens as your GPUs can process; none of them shows up on the bill.
- **Only your traffic.** The reserved GPUs are exclusive to you and answer your requests and nobody else's. Prompts and outputs are never used for training.
- **Change the base URL.** The API is OpenAI-compatible. Existing clients, SDKs and frameworks work once you swap the base URL and key.

## How it works. Three steps to your endpoint.

1. **Pick models and GPUs.** Choose open-weight models from the catalog, or bring your own weights. We size the GPUs with you.
2. **We deploy and run them.** enum installs the inference stack on your reserved GPUs in Germany, runs it and keeps it up to date.
3. **Call your endpoint.** Use any OpenAI-compatible client with your private base URL and API key.

## What enum takes care of. So your team can build the application.

- Inference server, NVIDIA drivers and CUDA on the reserved GPUs
- Model deployment, and version changes when you ask for them
- API keys for your applications
- Monitoring of the inference stack

## The models. Open weights you know.

We start with the models German companies ask for most. Each one runs only on enum's GPUs in Germany; no request goes to the model's publisher.

| Model | Publisher | Licence | GPUs (RTX PRO 6000, 96 GB) |
| --- | --- | --- | --- |
| gpt-oss-120b | OpenAI | Apache 2.0 | 1 (MXFP4) |
| gpt-oss-20b | OpenAI | Apache 2.0 | 1 |
| Llama 3.3 70B | Meta | Llama 3.3 Community License | 1 in FP8 with limited context, 2 for longer context |
| Mistral Small 24B | Mistral AI | Apache 2.0 | 1 |
| Qwen3 32B (Open weights. Runs only on enum servers in Germany, with no connection to the publisher.) | Qwen (Alibaba) | Apache 2.0 | 1 |
| DeepSeek (Larger DeepSeek models need several GPUs. We size them per inquiry.) | DeepSeek | Varies by model | On request |
| Your own weights | You | Yours | Sized with you |

GPU counts are indicative. Context length, quantisation and the number of concurrent users change the sizing, so we confirm it per inquiry.

## Per token or per GPU? When each one fits.

Reserved GPUs pay off when the load is steady. For occasional or spiky use, a shared per-token API usually costs less, and we tell you that in the first call.

|  | Shared per-token API | Reserved GPUs at enum |
| --- | --- | --- |
| Billing | Per million tokens | Per GPU per month |
| Monthly cost | Grows with usage | Fixed |
| Who uses the GPUs | Many customers | Only your requests |
| Rate limits | Set by the provider | Set by your GPU capacity |
| Own or fine-tuned weights | Rarely possible | Yes |
| Best fit | Prototypes, low or spiky volume | Steady internal assistants, document pipelines, agents |

## Your data. Under German law.

- GPUs and endpoints in Frankfurt, provided by enum GmbH, a German company with no US parent
- Prompts and outputs are not stored after the response by default, and never used for training
- Operational logs keep metadata such as timestamps and token counts, not content
- A GDPR data processing agreement in every contract
- Model cards and licences passed on to you, for your AI Act documentation

- **Data center ISO 27001 and EN 50600** (In place): Certificates of the Frankfurt facility operator.
- **GDPR data processing agreement** (Every contract): Under German law, with a German company and no US parent.
- **ISO 27001 for enum** (Q4 2026): In progress, target Q4 2026.
- **BSI C5** (2027): On the roadmap for 2027.

## Where teams start. Indicative sizing.

- **Pilot · 1 GPU**: One team, one model, for example gpt-oss-120b or Mistral Small.
- **Department · 2 to 4 GPUs**: Several teams, a 70B-class model with longer context, or a chat model next to an embedding model.
- **Company-wide · 8 GPUs**: Several models, many parallel users and room for failover. Larger setups on request.

Indicative only. We size every setup per inquiry.

## Questions. About private model endpoints.

### Is this pay per token?

No. You pay per GPU per month. Tokens are not billed, so you can send as many as your reserved GPUs can process. If you need more throughput, you add GPUs.

### What does "Models as a Service" mean at enum?

enum runs open-weight models for you on GPUs reserved for you, and you pay per GPU per month. Some providers use the same term for a shared API billed per token; ours is the reserved, per-GPU version. If you want to run the models yourself, reserve GPUs and bring your own stack.

### Is this the same as dedicated inference endpoints?

Close. Other providers call this dedicated endpoints or dedicated inference: a model served on GPUs that only your requests use. At enum the GPUs are reserved for your project and billed per month rather than per GPU hour.

### Who can see our prompts?

Your requests are processed on your reserved GPUs in Germany. By default we do not store prompts or outputs after the response, and we never use them for training. Operational logs keep metadata such as timestamps and token counts, not content.

### Can we use our own fine-tuned model?

Yes. Send us the weights or point us to a bucket in your enum object storage, and we serve them like any catalog model. The weights stay yours, and you take them with you if you leave.

### Do Qwen or DeepSeek send data to China?

No. Open weights are files: we load them onto enum's GPUs in Germany and serve them from there. The model has no connection to its publisher, and your data does not leave enum. Qwen is in the catalog; DeepSeek is available on request.

### Where do the models run?

In Frankfurt. The data center is certified to ISO 27001 and EN 50600; those certificates belong to the facility operator.

### How soon can we start, and what does it cost?

Models as a Service is available on request. For each inquiry we confirm capacity, a start date and the price per GPU per month. Pricing is on request.

### Is this GDPR compliant?

Processing happens in Germany under a GDPR data processing agreement with enum GmbH, a German company with no US parent. Whether a specific use is lawful stays your decision as the controller; we give your data protection officer the documents they need.
