Jay Miracola

AI

Read time: 5 mins

Read time: 5 mins

Racked, Stacked, Inference Ready

Racked, Stacked, Inference Ready

Share

Share

Racked, Stacked, Inference Ready

Racked, Stacked, Inference Ready

The Problem - I want MetalPlane

I wanted a way to drive my hardware from a simple API to make inference ready for Modelplane. I set out with the goal of having my Dell R730 not only be provisioned from a single API, but be complete end to end. Turn on, provision, bootstrap Kubernetes, and be inference ready. Then make it true of the reverse: remove the API which cascades a flow to tear down the machine and turn it off. This way my server can spin up and be inference ready in ~21 minutes, while saving me some power costs after teardown when unneeded.

The Tech

First I standardized on the control plane, which was Crossplane. But to be able to drive the machine towards Kubernetes I needed to add some extra flavor. First I went with Metal3 which if you know OpenStack is just a Kubernetes flavored Ironic to drive the iDRAC. Then layered in CAPI (Cluster API) to help drive the Kubernetes provisioning. I went with these because my use case was enterprise hardware. Ironic mounts the boot ISO over Redfish rather than PXE, so my control plane can live anywhere that can reach the BMC instead of having to share an L2 segment with the machine. There are a lot of great projects out there that could replace pieces of this stack like Talos, Tinkerbell, and more to accommodate PXE boot for other hardware.

Once the stack was tested, I layered in my Crossplane configuration:
https://github.com/jaymiracola/configuration-metal-k8s

Now I could drive the machine by picking it out of a pool of available hosts to provision like so:

apiVersion: metal.kube.farm/v1alpha1
kind: MetalCluster
metadata:
  name: r730
  namespace: metal
spec:
  host:
    pool: r730
    bootInterface: enp129s0f1
  network:
    nodeIP: 192.168.120.70
    prefix: 24
    gateway: 192.168.120.1
    dnsServer: 192.168.120.1
    podCIDR: 10.244.0.0/16
    serviceCIDR: 10.96.0.0/12
  kubernetes:
    version: v1.36.2
    nodeLabels:
      node.kubernetes.io/instance-type: r730
      modelplane.ai/pool: gpu
    imageURL: http://192.168.120.249:8080/CENTOS_10_NODE_IMAGE_K8S_v1.36.2.qcow2
    imageChecksum: a6ecbde797e3b5592b4e500b9f8b6bb12bfd4b9cf4ca67bef6c023b439698496
    imageChecksumType: sha256
    imageDiskFormat: qcow2
  cni:
    ciliumVersion: 1.20.1
  gpu:
    driverInstallerURL: http://192.168.120.249:8080/NVIDIA-Linux-x86_64-580.178.04.run
apiVersion: metal.kube.farm/v1alpha1
kind: MetalCluster
metadata:
  name: r730
  namespace: metal
spec:
  host:
    pool: r730
    bootInterface: enp129s0f1
  network:
    nodeIP: 192.168.120.70
    prefix: 24
    gateway: 192.168.120.1
    dnsServer: 192.168.120.1
    podCIDR: 10.244.0.0/16
    serviceCIDR: 10.96.0.0/12
  kubernetes:
    version: v1.36.2
    nodeLabels:
      node.kubernetes.io/instance-type: r730
      modelplane.ai/pool: gpu
    imageURL: http://192.168.120.249:8080/CENTOS_10_NODE_IMAGE_K8S_v1.36.2.qcow2
    imageChecksum: a6ecbde797e3b5592b4e500b9f8b6bb12bfd4b9cf4ca67bef6c023b439698496
    imageChecksumType: sha256
    imageDiskFormat: qcow2
  cni:
    ciliumVersion: 1.20.1
  gpu:
    driverInstallerURL: http://192.168.120.249:8080/NVIDIA-Linux-x86_64-580.178.04.run
apiVersion: metal.kube.farm/v1alpha1
kind: MetalCluster
metadata:
  name: r730
  namespace: metal
spec:
  host:
    pool: r730
    bootInterface: enp129s0f1
  network:
    nodeIP: 192.168.120.70
    prefix: 24
    gateway: 192.168.120.1
    dnsServer: 192.168.120.1
    podCIDR: 10.244.0.0/16
    serviceCIDR: 10.96.0.0/12
  kubernetes:
    version: v1.36.2
    nodeLabels:
      node.kubernetes.io/instance-type: r730
      modelplane.ai/pool: gpu
    imageURL: http://192.168.120.249:8080/CENTOS_10_NODE_IMAGE_K8S_v1.36.2.qcow2
    imageChecksum: a6ecbde797e3b5592b4e500b9f8b6bb12bfd4b9cf4ca67bef6c023b439698496
    imageChecksumType: sha256
    imageDiskFormat: qcow2
  cni:
    ciliumVersion: 1.20.1
  gpu:
    driverInstallerURL: http://192.168.120.249:8080/NVIDIA-Linux-x86_64-580.178.04.run

As the pirate aboard Captain Phillips' boat famously said "Look at me, I am the Cloud now".

And in reverse, I delete the newly created MetalCluster API and the entire thing gets deprovisioned, cleared, turned off and ready to be claimed again.

Fun Along the Way

Being real, I hit a few snags going through this. Some of which were comedic at the time. One such example was during the creation of this project, my power went out. No big deal but the server was drawing off my battery power so I turned it off. But control plane gonna control plane, and the server sprung back to life as the reconciliation occurred.

Another interesting bit was that the iDRAC loves to fail and boot ordering is hard. It's a bit older and no longer officially supported by Dell, so I had to make some concessions. The machine would have issues with boot ordering during provisioning and deprovisioning. Thankfully Crossplane Operations saved me here by creating a job when it sees a BareMetalHostcome up or down to fix the ordering. And in cases where the ISO still fails it follows Dell's recommended "FIX" of rebooting the iDRAC. I would call their support if I paid for it.

Last was the support matrix between node images, their kernels, and NVIDIA's GPU Operator. The operator derives its driver image tag from the node's OS labels, and on CentOS Stream 10 it asked nvcr.io for tags like 580.178.04-centos10 that don't resolve. So I pulled the driver out of the operator entirely and install it from preKubeadmCommands at first boot, which keeps the node image stock and leaves whatever schedules the GPU up to the consumer. Ubuntu 24.04 on 6.8.0-138 was a separate fight I never won: the driver installed cleanly and the cards still would not initialize, with GSP firmware boot timing out every time. CentOS Stream 10 on 6.12 worked on the first try. This turned out to be the right split for Modelplane anyway, since it installs the NVIDIA DRA driver itself and just expects the kernel driver to already be on the host, so a node with nothing above the driver is exactly what it wants to adopt.

AI Fleet Management

Now that the machine was up and ready, all it took was two more objects for Modelplane to adopt it into my inference fleet (for now the R730 and an NVIDIA Spark). The first was the InferenceClass to describe the hardware shape.

apiVersion: modelplane.ai/v1alpha1
kind: InferenceClass
metadata:
  name: r730
spec:
  description: Dell R730, two Tesla V100 PCIe
  devices:
    - name: gpu
      driver: gpu.nvidia.com
      deviceClassName: gpu.nvidia.com
      count: 2
      attributes:
        productName:
          string: Tesla V100-PCIE-32GB
        architecture:
          string: Volta
        cudaComputeCapability:
          version: "7.0.0"
      capacity:
        memory:
          value: 32Gi
apiVersion: modelplane.ai/v1alpha1
kind: InferenceClass
metadata:
  name: r730
spec:
  description: Dell R730, two Tesla V100 PCIe
  devices:
    - name: gpu
      driver: gpu.nvidia.com
      deviceClassName: gpu.nvidia.com
      count: 2
      attributes:
        productName:
          string: Tesla V100-PCIE-32GB
        architecture:
          string: Volta
        cudaComputeCapability:
          version: "7.0.0"
      capacity:
        memory:
          value: 32Gi
apiVersion: modelplane.ai/v1alpha1
kind: InferenceClass
metadata:
  name: r730
spec:
  description: Dell R730, two Tesla V100 PCIe
  devices:
    - name: gpu
      driver: gpu.nvidia.com
      deviceClassName: gpu.nvidia.com
      count: 2
      attributes:
        productName:
          string: Tesla V100-PCIE-32GB
        architecture:
          string: Volta
        cudaComputeCapability:
          version: "7.0.0"
      capacity:
        memory:
          value: 32Gi

Followed by the InferenceCluster.

 apiVersion: modelplane.ai/v1alpha1
 kind: InferenceCluster
 metadata:
   name: r730
 spec:
   cluster:
     source: Existing
     existing:
       secretRef:
         name: r730-cluster-kubeconfig
   nodePools:
     - name: gpu
       className: r730
 apiVersion: modelplane.ai/v1alpha1
 kind: InferenceCluster
 metadata:
   name: r730
 spec:
   cluster:
     source: Existing
     existing:
       secretRef:
         name: r730-cluster-kubeconfig
   nodePools:
     - name: gpu
       className: r730
 apiVersion: modelplane.ai/v1alpha1
 kind: InferenceCluster
 metadata:
   name: r730
 spec:
   cluster:
     source: Existing
     existing:
       secretRef:
         name: r730-cluster-kubeconfig
   nodePools:
     - name: gpu
       className: r730

Apply those two and Modelplane takes it from there. It installs the DRA driver, the GPUs comes up as a ResourceSlice, and the R730 is in the fleet next to the Spark.

Finally, I applied the InferenceGateway, ModelCache, ModelDeployment, and ModelService so the models I apply will schedule and serve the appropriate endpoints.

What's Next?

It is possible for total integration of these APIs. Serve a single simple API that gets the machine inference ready, consumes the kubeconfig secret, and becomes an inference ready host in Modelplane.

About Authors

Jay Miracola

Subscribe to the
Upbound Newsletter

Subscribe to the
Upbound Newsletter

Subscribe to the
Upbound Newsletter

Related

Related

Posts

Posts

Sep 28, 2026

Drinking Our Own Champagne: What It Takes to Own Your Intelligence

Sumbry

Sep 28, 2026

Drinking Our Own Champagne: What It Takes to Own Your Intelligence

Sumbry

Aug 19, 2026

Upbound Insights told us 230 resources were broken. Hub told us why.

Sumbry

Aug 19, 2026

Upbound Insights told us 230 resources were broken. Hub told us why.

Sumbry

Aug 19, 2026

Announcing Upbound v3: one view, API, and governance model for every control plane you run

Upbound

Aug 19, 2026

Announcing Upbound v3: one view, API, and governance model for every control plane you run

Upbound

Get Started with Upbound Crossplane 2.0

Trusted by 1,000+ organizations and downloaded over 100 million times.

Get Started with Upbound Crossplane 2.0

Trusted by 1,000+ organizations and downloaded over 100 million times.

Get Started with Upbound Crossplane 2.0

Trusted by 1,000+ organizations and downloaded over 100 million times.