
Sumbry
AI
We ran two proofs of concept self-hosting inference on Modelplane, and the most expensive problems were the ones that never threw an error.
In Part 1, I explained why Upbound is self-hosting inference on Modelplane. We had three original use cases: chat, agentic coding, and CVE scanning & backporting. These requirements led to the first models & GPUs we picked. As a follow-up in drinking our own champagne we’re now going to talk about the corks.
Modelplane splits the world into three layers:
the platform team who owns clusters, hardware classes, and gateways
the model operator or inference engineer who owns deployments, caches, and services
the people and agents consuming the inference endpoints
We decided to deploy two POCs to test our assumptions and surprise, that second layer, the land of the model operator or inference engineer is where we learned the most. This is SRE for Inference Engineering.
An inference engineer is the person who picks a model, gets it onto the right GPU, keeps it serving, optimizes it, and crawls out of bed when it fails. This is the traditional role of SRE applied to a new kind of workload and technology, where the familiar disciplines of capacity, testing, reliability, observability, and incident response now have to account for scarce GPUs, enormous models, and requests that run for hours.
We've now successfully run two proofs of concept (POCs):
POC-1 on Modelplane v0.4, with two ~30B mixture-of-experts models on two GPU classes serving real internal users
POC-2 on a Modelplane 0.5 pre-release with a new gateway architecture and a newer, bigger generation of models
Both POCs ran on unmodified upstream Modelplane. When something didn't work, we either worked around it beside Modelplane, in a directory literally named temporary/, or filed it as an upstream issue.
If Inference Engineering is Platform Engineering evolved for Inference workloads, then all the old lessons still apply! We’ll now walk through the Inference Management Lifecycle, one stage at a time: pick, place, fit, expose, serve, observe, change, & fail and cover the lessons that our POCs taught us along the way.
POC-1: Learning the lifecycle the hard way
Pick: test the hard path on purpose
As covered in Part 1, we ran Qwen3-Coder-30B-A3B for coding on an RTX PRO 6000 Blackwell (96 GB) and GLM-4.7-Flash for chat on an L40S (48 GB). While both models fit on either card and the cheapest setup would have been two identical L40S cards, we ruled that out. We wanted to test a heterogeneous deployment. With interchangeable cards the device selector problem disappears, and the device selector was the capability we were testing.
It was the right call. When both models fit both cards, a bad selector doesn't fail. It silently inverts placement. GLM pinned architecture == "Ada Lovelace" and Qwen required memory >= 70Gi. We confirmed placement against what the NVIDIA DRA driver actually publishes, not against Modelplane's status.gpuPools, because that status field echoes whatever you declared, wrong strings included.
Place: quota is permission to ask
Our first-choice GPU, an H100 p5.4xlarge, never produced an instance, even with 768 vCPUs of approved quota. Then g6e.2xlarge was refused with InsufficientInstanceCapacity in three availability zones within 80 minutes. What finally launched was a smaller shape, g6e.xlarge, in a zone that had just refused the larger one. Same L40S but cheaper.
The lesson for the inference engineer is that each instance size is its own capacity pool, even when the GPU is identical. A provider’s "try another zone" hints were wrong every time we followed one. And there's no capacity API to save you: describe-instance-type-offerings tells you a shape is offered, not that you can actually get one.
Fit: "it fits on paper" is a startup-time argument
Our fit math was pessimistic. Qwen came up with 1.19M tokens of KV cache, about 9x concurrency at 131K context, so we believed memory was a solved problem. Then GLM ran out of GPU memory nine times in four days. It did this mid-request, on an engine that had been serving happily for days. It crashed asking for 300 MiB with 159 MiB free and 2.88 GiB "reserved but unallocated." The cause was allocator fragmentation past the profiled peak, and it was triggered by a single unbounded request that took the model down for everyone.
The fix was ordinary memory tuning (JVM, I’m looking at you): memory utilization from 0.92 to 0.88, expandable_segments, and --max-num-seqs=16. That last flag shrank the CUDA graph pool from 2.57 GiB to 0.19 GiB, and free VRAM went from 23 MiB to 4.6 GiB with zero restarts since. Knowing the fix was easy but discovering the cause was hard.
Expose: sane defaults aren’t always secure and won’t get you to production
Out of the box, v0.4's gateway is an internet-facing load balancer with no auth. The listener protocol enum was [HTTP, TCP], so TLS couldn’t even be expressed with Modelplane 0.4. So we fixed it. Envoy Gateway's own CRDs were already installed and unused since we couldn’t run a gateway on our control plane, so we spun up an additional AI Gateway and added a SecurityPolicy that validates a Google identity token and terminates TLS. Neither touched a Modelplane-owned object because Crossplane reconciles those objects and reverts any edit.
Some things we couldn't reach on an Upbound Cloud Spaces control plane at all. InferenceGateway and ModelService both need a data path on the control plane, which our hosted control plane correctly refused. So we reached the engines through the per-replica HTTPRoute that ModelReplica already deployed. One visible side effect: every ModelDeployment reports Ready=False for the duration of the POC while the model serves fine.
Serve: the silently truncated stream
Our stable /qwen/ and /glm/ paths were hand-written routes standing in for Modelplane's generated ones. We copied the shape and immediately failed. Envoy fell back to a 15-second default that bounds the whole response. Streaming answers returned 200, sent about 313 KB of valid tokens, and then stopped dead. To a client, that looks exactly like a model that's done.
Fixing that only raised the ceiling to 60 seconds. We measured a generation that ran 191 seconds cleanly straight against Envoy, and got severed at 60.7 seconds through the ELB while tokens were still arriving every 0.29 seconds. Modelplane was right both times, since its generated route already sets the timeout. The first mistake was ours (default timeouts). The second was an upstream gap: the load balancer settings can't be changed without editing a Crossplane-managed object.
Two things went much better than expected: vLLM serves the Anthropic Messages API natively, with streaming and tool calls, so our agentic-coding use-case drove our self-hosted Qwen end to end with no translation. We were about to build one. Because the route rewrites the path prefix, everything the engine exposes, metrics included, came along for free.
Serve++: auth that works for people and fails for agents
Google identity tokens are a good fit for people and a poor fit for agents. A browser can't attach a bearer token to a page load or it subresources. A correctly secured endpoint rejects the preflight and the browser never sends the real request. For browsers, we moved login to the gateway, since Envoy Gateway's SecurityPolicy supports OIDC natively.
That brought two more surprises: Envoy's OIDC can't refresh against Google. Sessions died silently every 60 minutes. Worse, layering gateway OIDC in front of our chat app with its own login caused an outage. When the outer session expired, the app's background fetch() calls couldn't follow the redirect to Google, and the browser reported ERR_HTTP2_PROTOCOL_ERROR against our own hostname. We removed the outer layer.
Agents fared worse. Coding sessions run for hours and the tokens last one. Our workaround was a small local proxy that holds and refreshes the token. It works, but every coding user has to run something locally, and that's the opposite of a platform.
Observe: is it available doesn’t mean it works
Our Prometheus configuration sat for days while reporting everything as healthy while asking for a 50Gi volume with no storage class, on a cluster with no default one. We lost days of real metrics. Behind that was a second blocker: it could only schedule on the system node, which three controllers had already filled to 99% of CPU requests.
The dangerous half was the metrics adapter. It registered the custom and external metrics APIs as AVAILABLE=True against a server holding no data. A metric-driven autoscaler wouldn't have failed loudly. It would have read a healthy API and gotten nothing. And even with Prometheus up, nothing scraped the engines: there was no ServiceMonitor and the engine port was unnamed.
Change: every change needs a human in the loop
Upgrading vLLM from v0.23 to v0.25.1 deadlocked on every engine. A GPU claim is exclusive, so the new pod couldn't schedule. And with one replica, maxUnavailable: 25% rounds down to zero, so the old pod never let go. The rollout stalled indefinitely without ever failing. The workaround is a manual patch to maxUnavailable: 1, which Crossplane reverts afterwards. We budgeted 6 to 8 minutes of downtime per engine, and rolled them one at a time. Upstream should emit Recreate for engines that hold an exclusive device.
Scaling was its own lesson. Editing nodeCount on a live cluster does nothing, because it only applies at creation, and the cluster autoscaler outranks the manifest anyway. Patching a pool to zero with a pending GPU pod took it from one instance to two.
Fail: the evidence disappeared
Back to the GLM OOM. The request that killed it came from our chat app inside the cluster, bypassing the gateway entirely, so the gateway access log showed nothing. And by the time we looked, 99.5% of the retained engine log was our own Prometheus scraping /metrics every 15 seconds. We recovered the stack trace only because kubectl logs --previous still held the crashed container.
Two takeaways. Anything that talks to an engine directly is invisible to gateway-level monitoring. And your observability can be the very thing that erases your evidence.
POC-2: What Modelplane 0.5 changed
POC-2 ran on a new control plane with a 0.5 pre-release, and we had two goals: see how many POC-1 workarounds we could delete, and move to a newer generation of models and hardware configurations.

Expose & serve: what are you trying to accomplish?
The biggest change is that the InferenceGateway now runs on the inference cluster, not the control plane. That closed POC-1's largest blocker. It's an Envoy AI Gateway that speaks OpenAI at /v1 and Anthropic at /anthropic/v1, routes on the request's model field, and terminates TLS itself. Models are named <namespace>/<service>, so callers send ml-team/mimo or ml-team/qwen to a single hostname. That retired POC-1's second gateway, its second load balancer, and the hand-written routes that caused the truncation bug.
Every organization handles certificates, TLS, and authentication different. Modelplane isn’t opinionated about this, but it means that Modelplane terminates TLS but can't issue the certificate. It reads a Secret from the control plane, so we just issued it on the workload cluster and copy it across. Renewal is manual for now and will inherit our organization-wide security policy once we get out of POC-land.
Modelplane by design also only supports API Key authentication. Keys that don't expire would actually solve the agent problem, but we kept OIDC as an additive policy as well. That had a side effect: leaving spec.auth unset also dropped Modelplane's allow rules for /healthz and the HTTP to HTTPS redirect, so the load balancer health check started failing with 401. Unset auth also means open, so the policy has to land before the gateway gets an address.
Now our identity now reaches usage records. Modelplane trusts an x-modelplane-caller header when it sits behind another authenticator. Our policy writes the token's email into that header and overwrites a forged one.
Pick: the newer generation was too big for the budget
In just a couple of weeks, the flagship "small" models had grown and changed. GLM-5.3-Flash is a 321B model, 328 GB at FP8. Pairing it with Qwen3.8-Flash-Next would have needed six RTX PRO 6000s, about $18.5k a month, against a budget of roughly $15k. So we chose MiMo-V2.6-Flash for coding and agents on two RTX PRO 6000s, and dense Qwen3.8-27B for chat on an L40S.
One small lesson cost us a wrong estimate: read the file size, not the dtype table. We estimated MiMo at ~315 GB by counting safetensors dtypes. The actual checkpoint is 177.7 GB, because the experts are MXFP4.
Place: capacity is a request, not an absolute
The plan was one g7e.12xlarge, two GPUs in one box. It was refused 17 times across two zones in about 22 hours. What changed from POC-1 is that multi-zone pools did fall back: the L40S launched in a zone that had refused a single-zone request three hours earlier.
So we changed the shape of our model deployment. MiMo now runs across two separate g7e.2xlarge nodes in two different AZs, as a Leader/Worker gang with pipeline parallelism over ordinary networking. It loads and serves 386K tokens of KV cache, and answers a short coding prompt through the gateway in 1.9 seconds. This was the shape we least expected to work. It's also the one that best shows what a Modelplane ModelDeployment can describe.
Getting there taught us the most expensive operational rule during the POC. EKS gives up on a node group after about 36 minutes and marks it CREATE_FAILED. The Auto Scaling group behind it keeps trying. Ours did eventually launch both of MiMo's instances, hours after the "failure." Our watcher had already stopped at CREATE_FAILED, and a pool rename deleted the node group and terminated both instances. That was about two hours of hard-won capacity, thrown away. The rule now: track instances and node joins, not node-group status, and list a pool's instances before you rename or delete it. There’s a lesson here about verifying against the cloud and not a controller and it will be incorporated into Modelplane 0.6.
Fit & change: the model outran our intent
Stable vLLM v0.30.0 registers MiMo's architecture and still couldn't load the checkpoint. Every gang start failed on a tensor shape mismatch (1856 vs 1792) in the FP8 attention sharding, and a model-specific vLLM image fixed it. The inference engineer's version matrix has three axes: model, hardware, and engine build. Supported in the release notes isn't the same as loads.
Two smaller findings as well: The gateway lags a ready replica by one to two minutes while the endpoint and route catch up, so "Ready" and "routable" are different moments. ModelDeployment still exposes no volumes, so /dev/shm stays at the container default. That's a risk for multi-GPU serving we haven't hit yet, but expect to.
What we (re)learned
#1 Silent failures
None of the expensive problems threw an error. The truncated stream returned 200. The metrics adapter reported AVAILABLE. The node groups reported Ready. The selector would have inverted placement without a word. The OOM never reached the gateway log. An inference platform's job is to turn failures like these into loud ones, and much of what we've filed upstream is exactly that.
#2 Capacity is a runtime condition
Across both POCs, we got the GPU we planned for roughly half the time. Design deployments so the shape can change, one big box or two small ones, one zone or any, without rewriting the model's definition. In POC-2, that flexibility is what got MiMo serving.
#3 The inference engineer owns a different kind of knowledge
Runtime GPU memory headroom, engine builds per model, rollout strategy for exclusive devices, generation limits per caller. None of it lives at the platform layer, and all of it caused an outage or a stall. Modelplane gives the inference engineer a place to express those choices, and the defaults for that role are where we see the most room to improve and optimize (JVM, I’m looking at you again).
#4 Additive workarounds keep the findings honest
Everything we built lives beside Modelplane, is labeled for one-command teardown, and has a README saying what replaces it. That's why POC-2 let us delete our biggest stopgaps (the second gateway, the hand-written routes, the self-signed TLS): nothing had been forked into place.
#5 Agents are the hardest clients
Hour-long sessions, streamed tool calls, large contexts, and no human around to click "sign in again." Every timeout, token lifetime, and truncation bug showed up first, and worst, in coding agents. If your inference platform works for agents, it'll work for everything else.
What's next
Upstream, we're implementing all the lessons learned: configurable gateway load balancers, certificate issuance, Recreate rollouts for exclusive devices, an honest Ready signal, engine scraping out of the box, and auth that suits long-running agents. In our environment, the next step is moving POC-1's users onto POC-2's gateway and adding the fourth use case, internal tools.
The champagne is better than it was a month ago. There are fewer corks, and we think we know where most of the rest are.
About Authors

Sumbry






