Two nodes, in a rack, in Troy.
A custom Blackwell workstation and an NVIDIA DGX Spark, in a 20U open-frame rack on a UPS. We specced them, bought them and built them, which is why we can tell you exactly what is inside.
Good for. QLoRA on a 70B base takes about 38 GB of the 96, so there is headroom to go past 100B at 4-bit — this is the node for adapter fine-tuning at scale. Full-weight, with every parameter trainable, is a much heavier budget: roughly 8B with an 8-bit optimiser and gradient checkpointing. Also large-batch inference through vLLM or TensorRT-LLM, diffusion training, and molecular dynamics.
Full specification
| GPU | 1× NVIDIA RTX PRO 6000 Blackwell Workstation Edition — 96 GB GDDR7 with ECC, 128 MB L2 cache, PCIe 5.0 ×16, 4× DisplayPort 2.1, 600 W maximum power. Full Blackwell SM architecture: 5th-gen Tensor Cores with FP4/FP6/FP8/FP16/BF16/FP32, 4th-gen RT Cores, 9th-gen NVENC and 6th-gen NVDEC. 24,064 CUDA cores, 188 RT cores, 752 Tensor cores. FP64 runs on the CUDA cores at 1:64 rate, not on the Tensor Cores. |
|---|---|
| Memory bandwidth | 1,792 GB/s (GDDR7, 512-bit bus) |
| CPU | AMD Ryzen 9 7900X — 12C/24T, 4.7 GHz base / 5.6 GHz boost, Zen 4, TSMC 5 nm, 76 MB total cache (12 MB L2 + 64 MB L3), 170 W TDP, socket AM5 |
| System memory | 64 GB DDR5-6000, dual-channel, CL30-38-38-96, 2×32 GB DIMMs, EXPO enabled |
| Storage | 2× 2 TB PCIe 5.0 NVMe M.2 (TLC NAND, DRAM-cached, Phison E26) in RAID-0 — ~14 GB/s sequential read, ~12 GB/s write, ~1.5M IOPS random read per drive. Aggregate is bounded by the AM5 lane budget: with the GPU at ×16, only one M.2 is CPU-attached at PCIe 5.0 ×4 and the second sits behind the chipset link. |
| Networking | Realtek 2.5 GbE (onboard RTL8125BG). Remote management is over SSH; there is no BMC on this board. |
| Power | Seasonic Vertex PX-1200 — 1,200 W 80+ Platinum, fully modular, Cybenetics Platinum (ETA-A), ATX 3.1 / PCIe 5.1 compliant with a native 12V-2×6 connector. Measured system draw under sustained load is roughly 950–1,000 W at the wall. |
| Cooling | Noctua NH-D15S chromax.black — dual-tower, 6× heatpipe, 140 mm NF-A15 PWM. Chassis: 3× Noctua NF-A12x25 PWM intake, 1× exhaust. GPU: NVIDIA's dual-fan double-flow-through cooler. |
| Chassis & power protection | Fractal Design Meshify 2 XL (E-ATX), rackmount conversion in a 20U open-frame StarTech rack. CyberPower CP1500PFCLCD 1,500 VA / 1,000 W line-interactive UPS, pure sine wave output, AVR boost/buck. At full GPU load the system sits close to that 1,000 W ceiling, so the UPS is sized for ride-through and orderly shutdown rather than extended runtime. |
| Software stack | Ubuntu 24.04 LTS (kernel 6.8 HWE), CUDA 12.8.1, cuDNN 9.7, NCCL 2.25, NVIDIA driver 570.x, Docker CE 27.x with nvidia-container-toolkit, JupyterHub 5.x multi-user, Prometheus 3.x + Grafana 11.x (node_exporter and dcgm-exporter for GPU telemetry) |
| Typical workloads | LLM LoRA/QLoRA fine-tuning past 100B at 4-bit, full-weight fine-tuning to ~8B params (8-bit optimiser, gradient checkpointing), large-batch inference serving (vLLM, TensorRT-LLM), molecular dynamics (GROMACS, OpenMM), diffusion model training (SD3, FLUX.1), GNN training on large graphs |
Good for. The interesting bit is the 128 GB unified memory pool: CPU and GPU address one flat space over a coherent ~600 GB/s link, so there are no PCIe copies and no pinned staging buffers. NVIDIA rates it for fine-tuning models up to 70B, which matches what we see: a 70B base at 4-bit is about 38 GB, leaving plenty of room for adapters, optimiser state and a long context. It also suits agent workflows, RAG pipelines, and anything whose working set is awkwardly bigger than a normal card’s VRAM.
Full specification
| Superchip | NVIDIA GB10 Grace-Blackwell — 20-core ARMv9.2 CPU (10× Arm Cortex-X925 + 10× Cortex-A725) plus a Blackwell-class GPU on NVLink-C2C. Multi-die package on TSMC 3 nm, CPU die co-designed with MediaTek. Shared 128 GB unified memory pool addressable as a flat space by both CPU and GPU. |
|---|---|
| Unified memory | 128 GB LPDDR5x, 256-bit bus, 273 GB/s bandwidth. All 128 GB directly addressable by the GPU — kernels operate on CPU-allocated tensors with no transfer. |
| GPU compute | Blackwell-class GPU IP, 5th-gen Tensor Cores, FP4/FP6/FP8/FP16/BF16/FP32, dedicated Transformer Engine with FP8/FP4 block quantisation. Up to 1 petaFLOP FP4 with sparsity (1,000 NVFP4 TOPS). 9th-gen NVENC (AV1/HEVC), NVDEC, OFA. |
| NVLink-C2C | ~600 GB/s coherent fabric between the CPU and GPU dies, roughly 5× PCIe Gen 5 ×16. Cache-coherent across both domains; no PCIe bus contention. |
| Networking | NVIDIA ConnectX-7 SmartNIC — 200 Gb/s RDMA over 2× QSFP56-DD; separate 10 GbE RJ-45; Wi-Fi 7 (802.11be, 4×4 MIMO, 320 MHz channels), Bluetooth 5.4 |
| Storage | 1 TB PCIe Gen4 NVMe M.2 (TLC, DRAM-cached) — ~7 GB/s sequential read, ~5.2 GB/s write |
| Operating system | DGX OS 7 — NVIDIA-optimised Ubuntu 24.04 LTS derivative, NVIDIA DGX software stack, NCCL, CUDA, PyTorch containers, NVIDIA AI Workbench, NeMo Framework and NeMo Curator preinstalled |
| Software stack | CUDA 12.8.1, PyTorch 2.7+, TensorRT-LLM 0.16+, vLLM 0.8+, NVIDIA AI Workbench, NeMo Curator 0.8, NeMo Framework 24.12, Docker 27.x with nvidia-container-toolkit, optional Triton Inference Server |
| Typical workloads | LoRA/QLoRA fine-tuning up to ~70B (NVIDIA's rated figure for this machine) and full-weight fine-tuning to ~10B, agentic AI with long context windows, RAG pipelines (ColBERT, hybrid search, multi-vector indexing), sovereign local LLM hosting, ONNX/TRT-LLM optimised serving, GPU-accelerated Jupyter development |
Work out whether your job fits in 96 GB.
The question everyone actually has. This is an estimate, not a promise — but it shows its working, so you can see which term is the one hurting you and go argue with it.
—
Assumptions. Batch size 1, gradient checkpointing on, AdamW. KV cache is calibrated against a 70B grouped-query model at fp16 and scales linearly with context. Real numbers move with architecture, attention implementation and how much the allocator fragments — treat this as the right order of magnitude and the right answer about whether to bother trying. SPARK-1’s 128 GB unified pool is more forgiving again.
Parts,
not partnerships.
▓███████████████████████▒
▓███████████████████████▒
░▒▓▓█░ ░▒▓████████████████▒ ░░░░░░░░░░░ ░░░░░ ░░░░░ ░░░░ ░░░░░░░░░ ░░░░ ░░░░░░
░▒▓██▓▒░░▓█▓▓▒░ ░▓█████████████▒ ░████████████▓▒░ ▒████▒ ░████▒ ░████▒ ▓██████████▓▓░ ░████▒ ▒██████▓
▒▓██▒░ ░▒▒▒▒▓▓██▓░ ▒███████████▒ ░███████████████▒ ▓████░ ▓███▓ ░████▒ ▓█████████████▓░ ░████▒ ░████████░
░▒███▒ ░▒██▓▓▒░ ▒▓█▓▒ ▒█████████▒ ░████▒ ░▒█████░ ░████▒ ░████░ ░████▒ ▓████░ ░▒████▓ ░████▒ ▓███░▒████
▓███░ ▓██▒░ ▓█▓░ ▒███░ ░▓████████▒ ░████▒ ▒████▒ ▓████ ▓███▓ ░████▒ ▓████ ▒████░ ░████▒ ▒███▓ ▓███▓
░▓██▓ ▒██░ ▓███░▒███▒ ▒██████████▒ ░████▒ ░████▒ ░████▒ ░████░ ░████▒ ▓████ ░████▒ ░████▒ ░████░ ░████░
░▓██▒ ▒██▒ ▓██████▒░ ▒███▓░░▒▓████▒ ░████▒ ░████▒ ▓████ ▓███▓ ░████▒ ▓████ ░████▒ ░████▒ ▓███▒ ▒████
▓██▓ ░▓██▒▒▓▓▓▒░ ░▓███▓░ ░▓██▒ ░████▒ ░████▒ ░████▒░████░ ░████▒ ▓████ ▓████░ ░████▒ ▒████▓▓▓▓▓▓████▓
▒▓██▒ ░▒▓▒░░░░▒▓███▓▒░ ░▓████▒ ░████▒ ░████▒ ▒████▓███▒ ░████▒ ▒████▒▒▒▒▓█████▓ ░████▒ ░████████████████░
▒▓██▓░░ ▓█████▓▒░ ░▒▓███████▒ ░████▒ ░████▒ ████████ ░████▒ ▓█████████████▒ ░████▒ ████▓ ▓████░
░▒▓███▒ ░░▒▓▓███████████▒ ░▓▓▓▓░ ▓▓▓▓░ ▒▓▓▓▓▓▓▒ ░▓▓▓▓░ ▒▓▓▓▓▓▓▓▓▓▒▒░ ░▓▓▓▓░ ░▓▓▓▓░ ░▓▓▓▓▒
░▒▓▓▓▓▓██████████████████▒
▓███████████████████████▒
The manufacturers whose parts are in our machines. We list them because knowing what a server is made of matters — not because we have any relationship with them.
Want time on one of them?
Tell us what you are running and we will tell you which node fits and how long it ought to take.