Machines we will build you.
Three starting points, from an entry fine-tuning node to a dual-GPU training rig. None of them is a fixed product — we adjust the spec to whatever you actually plan to run.
They are named for the three metals in a catalytic converter: palladium, platinum and rhodium, in ascending order of how rare and expensive they are. It seemed more useful than a number.
Two reasons,
and only one is money.
The obvious one. If you would keep a node busy most of the month, owning beats renting inside a year, and after that the marginal hour costs you electricity. We will do that arithmetic with you honestly, including the case where it says rent.
The one that actually decides it. Some work cannot go anywhere else. Export-controlled drawings, patient records, fab process data under NDA, privileged documents. For that, a machine you own, sitting on a network you control, is not a cost optimisation — it is the only configuration that is allowed. We build those, document them for an audit, and can hand one over having never connected it to the internet at all.
What sovereign deployment involves →Specced, sourced, built, burned in.
A documented machine
The full parts list with part numbers, thermal numbers under sustained load, and a build guide detailed enough that someone else could put it together.
A working software stack
OS, NVIDIA drivers, CUDA, Docker with the container toolkit, and whichever frameworks you actually use — installed, pinned and checked, not left for you.
A burn-in report
We hammer it before it ships and record thermals and clock behaviour, so a flaky part shows up on our bench instead of three weeks into your training run.
Ongoing support
Optional, and worth having. Hardware fails eventually; when it does you will be talking to whoever put the machine together.
Three starting points.
Best for QLoRA fine-tuning of 7B–14B models and bf16 LoRA up to about 8B, small-batch inference, RAG indexing, and Jupyter-based research. Fits under a desk in a mid-tower, idles around 85 W, and takes about three hours from boxes to first boot.
Bill of materials
| GPU | 1× RTX 5080 — 16 GB GDDR7, Blackwell SM, PCIe 5.0 ×16, 256-bit, 960 GB/s, 360 W TDP |
|---|---|
| CPU | AMD Ryzen 7 7800X3D — 8C/16T, 96 MB L3 V-Cache, 4.2 GHz base / 5.0 GHz boost, Zen 4 |
| Memory | 64 GB DDR5-6000 CL30 (2×32 GB), EXPO, dual-channel |
| Storage | 1× 2 TB PCIe 4.0 NVMe (WD Black SN850X) + 1× 4 TB SATA SSD for datasets |
| Power | Corsair RM850x — 850 W 80+ Gold, fully modular, ATX 3.1 |
| Software | Ubuntu 24.04 LTS · CUDA 12.8 · PyTorch 2.7 · vLLM |
Best for QLoRA fine-tuning past 100B at 4-bit, full-weight fine-tuning to roughly 8B on one GPU, production inference with vLLM continuous batching, molecular dynamics, diffusion training, and multi-user JupyterHub. 96 GB also holds a 70B model at 4-bit with a long context. ECC memory earns its keep once a run lasts days. This is TROY-1 a generation on — we run the earlier revision ourselves.
Bill of materials
| GPU | 1× RTX PRO 6000 Blackwell Workstation Edition — 96 GB GDDR7 ECC, 1,792 GB/s, 512-bit, PCIe 5.0 ×16, 600 W maximum power |
|---|---|
| CPU | AMD Ryzen 9 9950X — 16C/32T, Zen 5, TSMC 4 nm, 80 MB cache, 4.3 GHz base / 5.7 GHz boost, 170 W TDP |
| Memory | 128 GB DDR5-6000 CL36 (2×64 GB), EXPO — two DIMMs rather than four, because four dual-rank sticks will not hold 6000 on AM5 |
| Storage | 2 TB PCIe 5.0 NVMe (Crucial T700) on the CPU-attached M.2 + 8 TB PCIe 4.0 NVMe for datasets. AM5 does not have the lanes for a second Gen5 drive plus U.2 alongside a ×16 GPU — that configuration needs TRX50. |
| Power | Seasonic Vertex PX-1200 — 1,200 W 80+ Platinum, ATX 3.1, native 12V-2×6. ~940 W under load, so about 78% of rating. |
| Motherboard | ASUS ProArt X870E-Creator WiFi, or an equivalent AM5 board with PCIe 5.0 ×16 plus a Gen5 M.2 |
| Cooling & chassis | 360 mm AIO on the CPU; full-tower with clearance for a 304 mm card |
| Software | Ubuntu 24.04 LTS · CUDA 12.8 · Docker 27.x · k3s · MLflow |
Best for tensor-parallel serving of 70B–123B models, full-parameter fine-tuning to roughly 10–13B with DeepSpeed ZeRO-3 across two GPUs, LoRA and QLoRA well past 70B, large-scale embedding and GNN training, and multi-tenant inference partitioned with MIG or MPS. The 192 GB is two 96 GB address spaces sharded in software, not a hardware-pooled pool — there is no NVLink on these cards.
Bill of materials
| GPU | 2× RTX PRO 6000 Blackwell Max-Q — 192 GB GDDR7 ECC total, 300 W each, PCIe 5.0 ×16 each with ~64 GB/s peer-to-peer. RTX PRO Blackwell does not support NVLink; the blower-style Max-Q is also the SKU meant for stacking two cards. |
|---|---|
| CPU | AMD Threadripper 7960X — 24C/48T, Zen 4, 152 MB total cache, 4.2 / 5.3 GHz, 350 W TDP, sTR5, 48 PCIe 5.0 lanes |
| Memory | 256 GB DDR5-5600 RDIMM ECC (4×64 GB), registered, quad-channel — 5600 is an EXPO profile; AMD's official rating for this CPU is 5200 |
| Storage | 2× 2 TB PCIe 5.0 NVMe for scratch, plus 2× 7.68 TB U.2 NVMe (Kioxia CM7-V) in RAID-0 for datasets — ~26 GB/s aggregate read across the U.2 pair |
| Power | Seasonic Prime TX-1600 — 1,600 W 80+ Titanium, single +12V rail, 3× EPS12V and 2× native 12V-2×6. ~1,150 W under load with Max-Q cards. |
| Networking | Onboard 10 GbE, and a spare PCIe 5.0 ×16 slot for a future 25/100 GbE NIC. IPMI is available on TRX50 only as an optional expansion card, which costs a slot. |
| Motherboard | ASUS Pro WS TRX50-SAGE WIFI — sTR5, quad-channel RDIMM, onboard 10 GbE |
| Cooling & chassis | sTR5 air or AIO cooler; full-tower or 4U rackmount with spacing for two dual-slot cards |
| Electrical | Draws ~10 A at 120 V under load, so it fits a standard 15 A circuit. Specced with 600 W Workstation Edition cards instead, it needs a dedicated 20 A circuit or 208/240 V. |
| Software | Ubuntu 24.04 LTS · CUDA 12.8 · NCCL 2.25 · DeepSpeed 0.16 · Slurm 24.05 · Enroot + Pyxis |
Get a build quoted.
Tell us what you plan to run and what you have to spend, and we will come back with a spec and an itemised quote.