Equipment

Hardware, purpose-built for AI.

From training clusters to edge terminals, every system ships pre-validated, fully observable, and ready to integrate. Designed, configured, and supported end-to-end by Nexus Forge engineers.

Starting prices are indicative for a base configuration. Final pricing depends on specification, region, and current supply — contact us for a tailored quote.

Training Clusters

Dense GPU systems for foundation-model pre-training and fine-tuning.

Forge GX-8 Cluster

Request pricing
8× NVIDIA H200 · 1.4TB HBM3e · NVLink 4

Flagship dense training node engineered for foundation-model fine-tuning, RLHF, and long-context multi-modal pipelines. Ships pre-validated against MLPerf reference runs.

  • 3.2 Tbps NVLink 4 fabric · 900 GB/s per-GPU
  • Direct-to-chip liquid cooling · 40 kW thermal envelope
  • Dual 56-core Intel Xeon Platinum · 2TB DDR5 ECC
  • 60TB NVMe Gen5 scratch · GPUDirect Storage
  • 8× 400Gb NDR InfiniBand · BlueField-3 DPUs
  • PyTorch, NeMo, Megatron-LM, DeepSpeed pre-tuned
Ideal for
70B+ parameter fine-tunes, RLHF, multi-modal training

Forge GX-4 Cluster

Request pricing
4× NVIDIA H100 · 640GB HBM3 · NVLink

Mid-scale training node for domain adaptation, continued pre-training, and embedding model development — a balanced footprint for teams growing past single-GPU experimentation.

  • 900 GB/s NVLink switch · NVSwitch backplane
  • Air-cooled 4U chassis · standard rack-ready
  • Dual 48-core AMD EPYC Genoa · 1TB DDR5 ECC
  • 30TB NVMe Gen5 · dual 200GbE
  • BMC + redundant Titanium PSUs
  • Ubuntu 24.04 LTS · CUDA 12.6 · NGC stack
Ideal for
7B–34B fine-tunes, embedding model training, LoRA at scale

Forge GX-Pod 32

Request pricing
32× H200 · 400GbE InfiniBand · 5.6TB HBM

Pre-validated 4-node training pod with non-blocking NDR interconnect, shared parallel storage, and rack-level liquid cooling — delivered as a turnkey unit with scheduler and observability online.

  • Non-blocking NDR InfiniBand · SHARP in-network reduction
  • Slurm + Kubernetes + Run:AI scheduling ready
  • Rear-door heat exchanger · 100 kW per rack
  • 1PB WekaFS parallel tier · GPUDirect Storage
  • Prometheus + Grafana + DCGM telemetry stack
  • Validated MLPerf Training v4.0 reference runs
Ideal for
Foundation-model pre-training, large-scale RLHF, multi-tenant research

Inference Servers

Low-latency serving nodes for agents, RAG, and real-time generation.

Forge IS-2 Inference Node

Request pricing
2× NVIDIA L40S · 192GB RAM · 100GbE

Cost-efficient 2U serving node for agentic workflows, RAG pipelines, and mid-sized model APIs — tuned out of the box for sub-100ms time-to-first-token and high concurrency.

  • 2× L40S (48GB GDDR6 ECC, FP8 native)
  • Dual 32-core AMD EPYC · 192GB DDR5 ECC
  • vLLM, TensorRT-LLM, and SGLang pre-tuned
  • PCIe Gen5 · 8TB NVMe Gen5 · hot-swap bays
  • Dual 100GbE + 1GbE BMC · redundant Titanium PSUs
  • Continuous batching, paged-attention, speculative decoding
Ideal for
13B–34B model serving, embeddings, RAG, agent backends

Forge IS-8 Inference Node

Request pricing
8× NVIDIA H100 NVL · 188GB HBM3

High-throughput 4U serving node engineered for long-context generation, multi-tenant model hosting, and production-grade chat platforms with strict latency SLOs.

  • 8× H100 NVL · NVLink-bridged pairs · 1.5TB HBM3 aggregate
  • Tensor + pipeline parallel ready · FP8 + INT4 quantization
  • Dual 56-core Xeon Platinum · 2TB DDR5 ECC
  • Triton Inference Server + KServe + Ray Serve
  • Dual 200GbE + BlueField-3 DPU · 30TB NVMe Gen5
  • Speculative decoding & lookahead decoding stack
Ideal for
70B+ chat serving, multi-tenant agent platforms, long-context APIs

Forge IS-Lite

Request pricing
1× L4 · 64GB RAM · 25GbE · 1U

Compact 1U serving node for small models, classifiers, vision pipelines, and edge gateways — silent enough for office deployment, dense enough to rack.

  • NVIDIA L4 (24GB GDDR6, 72W) · passively cooled
  • 16-core AMD EPYC Embedded · 64GB DDR5 ECC
  • 2TB NVMe Gen4 · dual 25GbE + 1GbE BMC
  • ONNX Runtime + TensorRT + Triton optimized
  • Whisper, Stable Diffusion, YOLO, SBERT preset stacks
  • <35 dBA acoustic · 350W max draw
Ideal for
Vision, ASR, small LLM serving, edge inference gateways

Workstation Terminal - Forge W7

Request pricing
RTX PRO 6000 Blackwell · 96GB GDDR7 ECC · FP4 native

Deskside AI workstation for model developers, researchers, and ML engineers — pre-configured and ready on day one.

  • AMD Threadripper PRO 7975WX (32C / 64T)
  • 256GB DDR5 ECC · 4TB NVMe Gen5 OS + 8TB NVMe data
  • 10 GbE · dual 4K display output
  • Ubuntu 24.04 LTS · CUDA 12.6 · PyTorch 2.5 · vLLM · Ollama
Ideal for
Local fine-tuning, prototyping, on-prem inference dev

Inference Server — Forge R4000 (4U)

Request pricing
8× NVIDIA L40S · 384GB GDDR6 ECC · FP8 native

Dense 4U inference server for high-throughput serving, batched embeddings, and multi-tenant model hosting.

  • 2× AMD EPYC 9354 (32C / 64T) · 512GB DDR5 ECC
  • 15TB NVMe array (RAID-10) · dual 10 GbE + BMC
  • Redundant 1600W Titanium PSU
  • Ubuntu Server 24.04 LTS · vLLM + Triton ready
Ideal for
Production inference, batched embeddings, multi-tenant serving

Edge & Embedded

Ruggedized systems for on-device perception, robotics, and field deployment.

Forge Edge Terminal

Request pricing
Jetson AGX Orin 64GB · 2TB NVMe · IP65

Rugged edge unit for on-device perception, sensor fusion, and autonomous decisioning — sealed for outdoor, vehicle, and industrial deployment without an enclosure.

  • 275 TOPS AI compute · 12-core Arm Cortex-A78AE
  • -40°C to 70°C operating · IP65 sealed aluminum chassis
  • GMSL2 (8-camera) + CAN-FD + isolated GPIO + RS-485
  • 5G sub-6 + Wi-Fi 6E + dual GbE + GNSS RTK
  • 9–36V DC input · ignition sense · 60W fanless
  • JetPack 6 · ROS 2 Humble · DeepStream · Isaac ready
Ideal for
Robotics, autonomous vehicles, mining, smart facilities

Forge Edge Micro

Request pricing
Jetson Orin Nano · 512GB SSD · DIN-rail

Compact DIN-rail embedded controller for industrial sensors, computer vision, and PLC-adjacent inference — drop-in for existing automation cabinets.

  • 40 TOPS AI compute · 6-core Arm Cortex-A78AE
  • 8W idle / 25W peak · fanless · DIN-rail mount
  • Dual GbE with TSN · USB 3.2 · 4× isolated DI/DO
  • OPC-UA, MQTT, Modbus TCP, EtherNet/IP
  • PoE+ powered · 9–28V DC backup
  • Ubuntu 22.04 + Docker + Triton edge runtime
Ideal for
Quality inspection, predictive maintenance, machine vision

Forge Vision Gateway

Request pricing
8× PoE cameras · L4 GPU · 16TB storage

Plug-in 2U gateway for multi-camera vision pipelines with on-prem retention, hardware-accelerated decode, and role-based access — no cloud round-trip required.

  • NVIDIA L4 GPU · 64GB DDR5 · 16TB enterprise NVMe
  • 8× PoE++ camera ports · ONVIF + RTSP ingest
  • Frigate, DeepStream, and custom Triton pipelines
  • Hardware H.264/H.265/AV1 decode · 200+ streams
  • RBAC + SSO + tamper-evident audit log
  • Encryption at rest · 30-day rolling retention default
Ideal for
Retail analytics, perimeter security, workplace safety monitoring

Networking Fabric

Lossless interconnect engineered for GPU-to-GPU collective traffic.

Forge IB-NDR Switch

Request pricing
64-port NDR InfiniBand · 400Gb/s

Non-blocking 1U spine for multi-node training pods with adaptive routing and in-network collectives — the backbone of MLPerf-class training fabrics.

  • 64× 400Gb NDR ports · 51.2 Tb/s aggregate · <130ns latency
  • SHARPv3 in-network reduction · adaptive routing
  • PFC + ECN lossless · congestion control offloads
  • UFM Enterprise telemetry & fabric management
  • Redundant hot-swap PSUs + N+1 fans
  • Liquid-cooling ready · front-to-back airflow
Ideal for
Training pod spine, GPUDirect RDMA, parallel storage fabric

Forge EN-400 Switch

Request pricing
32× 400GbE · RoCEv2 · cut-through

Open-network Ethernet leaf for inference clusters, storage backplanes, and AI-Ethernet fabrics — runs SONiC or vendor NOS with zero-touch provisioning.

  • 32× 400GbE QSFP-DD · 12.8 Tb/s · cut-through forwarding
  • RoCEv2 + DCQCN + PFC watchdog tuned for AI
  • SONiC, Cumulus, and SAI-compatible NOS options
  • BGP EVPN + VXLAN · MC-LAG · streaming telemetry
  • Redundant 1+1 PSUs · 5+1 fans · hot-swap
  • gNMI / gNOI · Ansible & Terraform providers
Ideal for
Inference fabric, storage network, AI-Ethernet leaf-spine

Storage Systems

Parallel and object storage tuned for training datasets and model artifacts.

Forge DS-Parallel 200

Request pricing
200TB NVMe · 80 GB/s read · RDMA

All-flash parallel filesystem for training dataloaders, checkpoint streams, and shared scratch — engineered to keep GPUs saturated under heavy shuffle.

  • 200TB usable NVMe Gen5 · 80 GB/s read · 45 GB/s write
  • WekaFS or Lustre · POSIX + NFS + S3 multi-protocol
  • GPUDirect Storage · NVMe-oF over RDMA
  • Snapshots, clones, and policy-based tiering to object
  • Dual active-active controllers · HA failover
  • Inline compression + end-to-end checksums
Ideal for
Training dataset hot tier, fast checkpointing, shared scratch

Forge DS-Object 1P

Request pricing
1PB raw · S3 API · erasure coded

On-prem object store for datasets, model registry, and long-term retention — S3 + Iceberg compatible so your data lake and AI stack share one source of truth.

  • 1PB raw · 12+4 erasure coding · multi-site replication
  • S3 + Iceberg + Parquet optimized · presigned URLs
  • Versioning, object lock, and WORM compliance modes
  • AES-256 encryption at rest · KMS / Vault integration
  • Lifecycle policies + cross-region replication targets
  • IAM-style RBAC · audit log streaming to SIEM
Ideal for
Cold data, model registry, dataset lake, audit archive

Orchestration & Racks

Pre-integrated racks with scheduling, observability, and security baked in.

Forge Orchestrator Rack

Request pricing
42U · Kubernetes + Slurm · 25kW

Pre-validated 42U control rack delivered with scheduler, observability, secrets management, and identity wired in — your platform team starts on day one, not month three.

  • Kubernetes + Slurm + Run:AI / Kueue scheduling
  • Prometheus + Grafana + Loki + DCGM exporters
  • OIDC SSO · HashiCorp Vault · sealed-secrets
  • Harbor private registry + Argo CD GitOps
  • Out-of-band BMC fabric · serial console aggregation
  • Power, thermal, and per-port network telemetry
Ideal for
Turnkey on-prem AI platform, multi-team GPU sharing

Forge Liquid Rack 60kW

Request pricing
Rear-door heat exchanger · 60kW · 48U

High-density liquid-cooled 48U rack for H200, B200, and GB200-class systems — drops into existing data halls without raised-floor chilled water.

  • Rear-door heat exchanger · 60kW sustained · 80kW peak
  • Integrated CDU · facility water or self-contained loop
  • Dual 3-phase 415V feeds · per-outlet metering
  • Leak detection · auto-shutoff · pressure monitoring
  • Hot-aisle containment kit · blanking & brush strips
  • Seismic Zone 4 anchoring · ZeroU PDU pair included
Ideal for
Dense training pods, HPC workloads, B200 / GB200 deployments

Every system ships turnkey.

Hardware is just the start. We handle the full lifecycle so your team can focus on the models, not the rack.

Site survey & power planning
Thermal, power, and network readiness assessment before delivery.
White-glove installation
Rack, cable, burn-in, and validation onsite by our engineers.
Software stack bring-up
Drivers, schedulers, observability, and your first model deployed.
24/7 support & spares
Next-business-day hardware replacement and on-call SRE coverage.

Need a custom configuration?

Tell us your workload, footprint, and timeline. We'll come back with a validated bill of materials and a deployment plan.

Request a quote