
Jay Miracola
AI
The Problem - I want MetalPlane
I wanted a way to drive my hardware from a simple API to make inference ready for Modelplane. I set out with the goal of having my Dell R730 not only be provisioned from a single API, but be complete end to end. Turn on, provision, bootstrap Kubernetes, and be inference ready. Then make it true of the reverse: remove the API which cascades a flow to tear down the machine and turn it off. This way my server can spin up and be inference ready in ~21 minutes, while saving me some power costs after teardown when unneeded.
The Tech
First I standardized on the control plane, which was Crossplane. But to be able to drive the machine towards Kubernetes I needed to add some extra flavor. First I went with Metal3 which if you know OpenStack is just a Kubernetes flavored Ironic to drive the iDRAC. Then layered in CAPI (Cluster API) to help drive the Kubernetes provisioning. I went with these because my use case was enterprise hardware. Ironic mounts the boot ISO over Redfish rather than PXE, so my control plane can live anywhere that can reach the BMC instead of having to share an L2 segment with the machine. There are a lot of great projects out there that could replace pieces of this stack like Talos, Tinkerbell, and more to accommodate PXE boot for other hardware.
Once the stack was tested, I layered in my Crossplane configuration:
https://github.com/jaymiracola/configuration-metal-k8s
Now I could drive the machine by picking it out of a pool of available hosts to provision like so:
As the pirate aboard Captain Phillips' boat famously said "Look at me, I am the Cloud now".
And in reverse, I delete the newly created MetalCluster API and the entire thing gets deprovisioned, cleared, turned off and ready to be claimed again.
Fun Along the Way
Being real, I hit a few snags going through this. Some of which were comedic at the time. One such example was during the creation of this project, my power went out. No big deal but the server was drawing off my battery power so I turned it off. But control plane gonna control plane, and the server sprung back to life as the reconciliation occurred.
Another interesting bit was that the iDRAC loves to fail and boot ordering is hard. It's a bit older and no longer officially supported by Dell, so I had to make some concessions. The machine would have issues with boot ordering during provisioning and deprovisioning. Thankfully Crossplane Operations saved me here by creating a job when it sees a BareMetalHostcome up or down to fix the ordering. And in cases where the ISO still fails it follows Dell's recommended "FIX" of rebooting the iDRAC. I would call their support if I paid for it.
Last was the support matrix between node images, their kernels, and NVIDIA's GPU Operator. The operator derives its driver image tag from the node's OS labels, and on CentOS Stream 10 it asked nvcr.io for tags like 580.178.04-centos10 that don't resolve. So I pulled the driver out of the operator entirely and install it from preKubeadmCommands at first boot, which keeps the node image stock and leaves whatever schedules the GPU up to the consumer. Ubuntu 24.04 on 6.8.0-138 was a separate fight I never won: the driver installed cleanly and the cards still would not initialize, with GSP firmware boot timing out every time. CentOS Stream 10 on 6.12 worked on the first try. This turned out to be the right split for Modelplane anyway, since it installs the NVIDIA DRA driver itself and just expects the kernel driver to already be on the host, so a node with nothing above the driver is exactly what it wants to adopt.
AI Fleet Management
Now that the machine was up and ready, all it took was two more objects for Modelplane to adopt it into my inference fleet (for now the R730 and an NVIDIA Spark). The first was the InferenceClass to describe the hardware shape.
Followed by the InferenceCluster.
Apply those two and Modelplane takes it from there. It installs the DRA driver, the GPUs comes up as a ResourceSlice, and the R730 is in the fleet next to the Spark.
Finally, I applied the InferenceGateway, ModelCache, ModelDeployment, and ModelService so the models I apply will schedule and serve the appropriate endpoints.
What's Next?
It is possible for total integration of these APIs. Serve a single simple API that gets the machine inference ready, consumes the kubeconfig secret, and becomes an inference ready host in Modelplane.
About Authors

Jay Miracola







