Catalytic AITroy, New York
INFRASTRUCTURE2 NODES · 20UTROY · NY

Two nodes, in a rack, in Troy.

A custom Blackwell workstation and an NVIDIA DGX Spark, in a 20U open-frame rack on a UPS. We specced them, bought them and built them, which is why we can tell you exactly what is inside.

Node 01
TROY-1 Custom Blackwell GPU server Blackwell
GPU
RTX PRO 6000
VRAM
96 GB
Bandwidth
1,792 GB/s
CUDA cores
24,064
CPU
Ryzen 9 7900X
System RAM
64 GB

Good for. QLoRA on a 70B base takes about 38 GB of the 96, so there is headroom to go past 100B at 4-bit — this is the node for adapter fine-tuning at scale. Full-weight, with every parameter trainable, is a much heavier budget: roughly 8B with an 8-bit optimiser and gradient checkpointing. Also large-batch inference through vLLM or TensorRT-LLM, diffusion training, and molecular dynamics.

Full specification
GPU1× NVIDIA RTX PRO 6000 Blackwell Workstation Edition — 96 GB GDDR7 with ECC, 128 MB L2 cache, PCIe 5.0 ×16, 4× DisplayPort 2.1, 600 W maximum power. Full Blackwell SM architecture: 5th-gen Tensor Cores with FP4/FP6/FP8/FP16/BF16/FP32, 4th-gen RT Cores, 9th-gen NVENC and 6th-gen NVDEC. 24,064 CUDA cores, 188 RT cores, 752 Tensor cores. FP64 runs on the CUDA cores at 1:64 rate, not on the Tensor Cores.
Memory bandwidth1,792 GB/s (GDDR7, 512-bit bus)
CPUAMD Ryzen 9 7900X — 12C/24T, 4.7 GHz base / 5.6 GHz boost, Zen 4, TSMC 5 nm, 76 MB total cache (12 MB L2 + 64 MB L3), 170 W TDP, socket AM5
System memory64 GB DDR5-6000, dual-channel, CL30-38-38-96, 2×32 GB DIMMs, EXPO enabled
Storage2× 2 TB PCIe 5.0 NVMe M.2 (TLC NAND, DRAM-cached, Phison E26) in RAID-0 — ~14 GB/s sequential read, ~12 GB/s write, ~1.5M IOPS random read per drive. Aggregate is bounded by the AM5 lane budget: with the GPU at ×16, only one M.2 is CPU-attached at PCIe 5.0 ×4 and the second sits behind the chipset link.
NetworkingRealtek 2.5 GbE (onboard RTL8125BG). Remote management is over SSH; there is no BMC on this board.
PowerSeasonic Vertex PX-1200 — 1,200 W 80+ Platinum, fully modular, Cybenetics Platinum (ETA-A), ATX 3.1 / PCIe 5.1 compliant with a native 12V-2×6 connector. Measured system draw under sustained load is roughly 950–1,000 W at the wall.
CoolingNoctua NH-D15S chromax.black — dual-tower, 6× heatpipe, 140 mm NF-A15 PWM. Chassis: 3× Noctua NF-A12x25 PWM intake, 1× exhaust. GPU: NVIDIA's dual-fan double-flow-through cooler.
Chassis & power protectionFractal Design Meshify 2 XL (E-ATX), rackmount conversion in a 20U open-frame StarTech rack. CyberPower CP1500PFCLCD 1,500 VA / 1,000 W line-interactive UPS, pure sine wave output, AVR boost/buck. At full GPU load the system sits close to that 1,000 W ceiling, so the UPS is sized for ride-through and orderly shutdown rather than extended runtime.
Software stackUbuntu 24.04 LTS (kernel 6.8 HWE), CUDA 12.8.1, cuDNN 9.7, NCCL 2.25, NVIDIA driver 570.x, Docker CE 27.x with nvidia-container-toolkit, JupyterHub 5.x multi-user, Prometheus 3.x + Grafana 11.x (node_exporter and dcgm-exporter for GPU telemetry)
Typical workloadsLLM LoRA/QLoRA fine-tuning past 100B at 4-bit, full-weight fine-tuning to ~8B params (8-bit optimiser, gradient checkpointing), large-batch inference serving (vLLM, TensorRT-LLM), molecular dynamics (GROMACS, OpenMM), diffusion model training (SD3, FLUX.1), GNN training on large graphs
Node 02
SPARK-1 NVIDIA DGX Spark GB10
Superchip
GB10
Unified memory
128 GB
NVLink-C2C
~600 GB/s
Memory BW
273 GB/s
CPU
20-core ARMv9
OS
DGX OS 7

Good for. The interesting bit is the 128 GB unified memory pool: CPU and GPU address one flat space over a coherent ~600 GB/s link, so there are no PCIe copies and no pinned staging buffers. NVIDIA rates it for fine-tuning models up to 70B, which matches what we see: a 70B base at 4-bit is about 38 GB, leaving plenty of room for adapters, optimiser state and a long context. It also suits agent workflows, RAG pipelines, and anything whose working set is awkwardly bigger than a normal card’s VRAM.

Unified memory topology: Grace CPU and Blackwell GPU sharing one 128 GB address space over NVLink-C2C GRACE CPU BLACKWELL GPU 20-CORE ARMv9 5TH-GEN TENSOR ~600 GB/s NVLINK-C2C · COHERENT 128 GB UNIFIED MEMORY LPDDR5x · 273 GB/s · SINGLE FLAT ADDRESS SPACE
Fig. 02 — SPARK-1 unified memory topology
Full specification
SuperchipNVIDIA GB10 Grace-Blackwell — 20-core ARMv9.2 CPU (10× Arm Cortex-X925 + 10× Cortex-A725) plus a Blackwell-class GPU on NVLink-C2C. Multi-die package on TSMC 3 nm, CPU die co-designed with MediaTek. Shared 128 GB unified memory pool addressable as a flat space by both CPU and GPU.
Unified memory128 GB LPDDR5x, 256-bit bus, 273 GB/s bandwidth. All 128 GB directly addressable by the GPU — kernels operate on CPU-allocated tensors with no transfer.
GPU computeBlackwell-class GPU IP, 5th-gen Tensor Cores, FP4/FP6/FP8/FP16/BF16/FP32, dedicated Transformer Engine with FP8/FP4 block quantisation. Up to 1 petaFLOP FP4 with sparsity (1,000 NVFP4 TOPS). 9th-gen NVENC (AV1/HEVC), NVDEC, OFA.
NVLink-C2C~600 GB/s coherent fabric between the CPU and GPU dies, roughly 5× PCIe Gen 5 ×16. Cache-coherent across both domains; no PCIe bus contention.
NetworkingNVIDIA ConnectX-7 SmartNIC — 200 Gb/s RDMA over 2× QSFP56-DD; separate 10 GbE RJ-45; Wi-Fi 7 (802.11be, 4×4 MIMO, 320 MHz channels), Bluetooth 5.4
Storage1 TB PCIe Gen4 NVMe M.2 (TLC, DRAM-cached) — ~7 GB/s sequential read, ~5.2 GB/s write
Operating systemDGX OS 7 — NVIDIA-optimised Ubuntu 24.04 LTS derivative, NVIDIA DGX software stack, NCCL, CUDA, PyTorch containers, NVIDIA AI Workbench, NeMo Framework and NeMo Curator preinstalled
Software stackCUDA 12.8.1, PyTorch 2.7+, TensorRT-LLM 0.16+, vLLM 0.8+, NVIDIA AI Workbench, NeMo Curator 0.8, NeMo Framework 24.12, Docker 27.x with nvidia-container-toolkit, optional Triton Inference Server
Typical workloadsLoRA/QLoRA fine-tuning up to ~70B (NVIDIA's rated figure for this machine) and full-weight fine-tuning to ~10B, agentic AI with long context windows, RAG pipelines (ColBERT, hybrid search, multi-vector indexing), sovereign local LLM hosting, ONNX/TRT-LLM optimised serving, GPU-accelerated Jupyter development
Will it fit?

Work out whether your job fits in 96 GB.

The question everyone actually has. This is an estimate, not a promise — but it shows its working, so you can see which term is the one hurting you and go argue with it.

Estimated peak VRAM · TROY-1

Assumptions. Batch size 1, gradient checkpointing on, AdamW. KV cache is calibrated against a 70B grouped-query model at fp16 and scales linearly with context. Real numbers move with architecture, attention implementation and how much the allocator fragments — treat this as the right order of magnitude and the right answer about whether to bother trying. SPARK-1’s 128 GB unified pool is more forgiving again.

Built with

Parts,
not partnerships.

GPU · Superchip
CPU
DGX Spark
Memory
NVMe storage
Power supply
Cooling
Chassis

The manufacturers whose parts are in our machines. We list them because knowing what a server is made of matters — not because we have any relationship with them.

Want time on one of them?

Tell us what you are running and we will tell you which node fits and how long it ought to take.