GPU rental · on request
NVIDIA's Blackwell workstation-class card with 96 GB of GDDR7, joining the AxForge fleet as a dedicated EU-hosted machine. It is on the way — reserve it now and an engineer confirms timing with you.
Specifications
| Card | NVIDIA RTX 6000 Pro |
|---|---|
| Memory | 96 GB GDDR7 |
| Class | Blackwell workstation class |
| Availability | On request — reservable now, not rentable today |
| Tenancy | Dedicated machine, your traffic only |
| Pricing | €1.73 / GPU-hr · billed monthly (launch pricing) |
Fit
| Use case | Why it fits |
|---|---|
| Large-context inference | 96 GB leaves room for long KV caches next to the model weights. |
| Fine-tuning small models | Enough memory to train small models on a single dedicated card. |
| Image / video generation | Workstation-class card with headroom for generation pipelines. |
Community data
On the one benchmark that runs identically across NVIDIA cards, the RTX 6000 Pro decodes at 274.2 t/s with 14,855 t/s prompt processing — about 3.6× an RTX 3060. An RTX 5090 is actually a touch faster single-stream on a small model (290 t/s). That is not the point of this card.
Same-workload numbers from the llama.cpp CUDA scoreboard (Llama 2 7B Q4_0, tg128, full GPU offload) — community measurements, not ours. We have not measured this card on our own nodes yet; when it joins the fleet, we publish our own numbers. Full ladder on the GPU catalogue.
Community data
GamersNexus ran the same LM Studio workloads on an RTX 4090, RTX 5090 and RTX 6000 Pro. Small models: the cards are close. Models that crowd 24–32 GB: the consumer cards start spilling out of VRAM and collapse, while 96 GB keeps the whole model on the GPU.
| Model | RTX 4090 · 24 GB | RTX 5090 · 32 GB | RTX 6000 Pro · 96 GB |
|---|---|---|---|
| DeepSeek Llama 8B Distil | ~62 t/s | ~81 t/s | ~81 t/s |
| Phi-4 8-bit | ~43 t/s | ~51 t/s | 62 t/s |
| Qwen 2.5 | ~32 t/s | ~35 t/s | 44 t/s |
| InternLM | ~12 t/s | ~40 t/s | 50 t/s |
| Mistral Small (~26 GB) | 6.4 t/s | 17 t/s | 42.4 t/s |
| Gemma 3 27B | ~4 t/s | ~5 t/s | 29 t/s |
| Llama 3.3 70B Q4 | very low | very low | vastly faster |
The RTX 6000 Pro is not 6× faster silicon — it simply never has to offload. That makes it a different class of card: large models entirely on one GPU, with workstation-class bandwidth behind them. Approximate figures (~) read from the source's charts.
vs RTX 5090
The 5090 is the fastest-per-euro choice for models that fit in 32 GB. This card is for the models that don't.
vs DGX Spark
DGX Spark holds even more (128 GB unified) for less money, but at ~273 GB/s bandwidth — capacity first, decode speed second. This card gives you both, at a card price.
Data & privacy
Like every AxForge machine: your model, your traffic, our hardware — prompts never persisted. Only request metadata (token counts, timestamps, status) is kept for billing and operations. Full policy at axforge.ai/privacy.
FAQ
Not yet — the RTX 6000 Pro is on the way to the fleet. You can reserve it or register interest now, and an engineer confirms timing with you before you commit.
€1.73 per GPU-hour, billed monthly for the dedicated machine (launch pricing). We watch the market and price under it: that is 90% of the lowest comparable on-demand listing we found on 2026-08-26.
With 96 GB of GDDR7 in a Blackwell workstation-class card: large-context inference, fine-tuning small models, and image or video generation.
96 GB of GDDR7 — the largest memory in AxForge's RTX tiers.
In an AxForge EU region. Stockholm and Málaga are live today, with more EU regions in deployment; placement is confirmed with your reservation.
Dedicated. Your model, your traffic, our hardware — prompts never persisted. See the privacy policy.