Sumbry

AI

Read time: 5 mins

Read time: 5 mins

Drinking Our Own Champagne: What It Takes to Own Your Intelligence

Drinking Our Own Champagne: What It Takes to Own Your Intelligence

We spent a decade building control planes for infrastructure. Now we're using that experience to learn how platform engineers become inference engineers.

Share

Share

Drinking Our Own Champagne: What It Takes to Own Your Intelligence

Drinking Our Own Champagne: What It Takes to Own Your Intelligence

“Ahhhhh, crap!”  These are the words that we collectively muttered when the realization actually hit us.  We have a technology for self-hosting inference. We weren't running inference.

I don't love writing that down, but it's the entire reason Project Champagne exists. Upbound created Crossplane over 10 years ago as a framework for building cloud native control planes and managing infrastructure & applications via Kubernetes style declarative APIs.  Upbound launched 0.1 of Modelplane 3 months ago that leveraged Crossplane to orchestrate models, serve engines, and manage AI infrastructure across cloud, neocloud, and self-hosted environments.  

We’ve always been the infrastructure experts, we were just now applying that to the world of AI which is just another complex infrastructure stack. We were leveraging all of our Crossplane experience to solve Day Zero, One, and Two challenges right out of the box.  The exception here was that we weren’t practicing what we preached. We needed to self-host inference ourselves, and solve our own internal use-cases if others were going to buy into our ultimate value proposition.

The goal for Champagne is simple: run Modelplane in our own environment, point our own workloads at it, and keep migrating use cases until the thing we ship is the thing we depend on.

The reality is that every platform team is about to inherit self-hosted inference. Someone in your organization has already decided the coding assistant can't send source code to a third party;  frontier model usage is too expensive; the website or support tool has to run against internal data, the list goes on.

That work is landing on the platform team (which I often refer to as the kitchen-sink team) and when it does, platform teams and engineers will be scrambling to understand the complex world of AI and models and get something going quickly that can meet enterprise scale.  

As Upbound navigates the world of self-hosted inference, we’re going to publish a series of blog posts outlining this journey: what we deployed, what we learned, and the mistakes we made along the way.  

Welcome to the Champagne Campaign.

Goals, budget, and timeline

We set three goals up front. First, run real internal workloads on Modelplane, used by real people. Second, run it the way a customer would, with the modelplane hosted on a managed control plane (via Crossplane or Upbound) and the GPUs in our own AWS account, so we test the split most enterprises will actually deploy. Third, find the gaps before customers do, and fix every one of them upstream rather than in a private fork. That last goal is the most important: A dogfooding project that ends in a pile of internal workarounds hasn't helped anyone.

We kept the budget deliberately modest, because a platform team getting started won't be handed a GPU cluster on day one either. We started with a $15,000/month budget that we would slowly grow into over time with ultimately no hard ceiling as we learned more.  Our strategy has always been to start small and optimize, then go big! The timeline was to get modelplane up and running in a couple of days and have our first customer workloads on it in under a week (surprise: we exceeded this timeframe by 50%).   We wanted it short enough to keep us honest and long enough to see the failures that only show up after days of real use.

Multiple use-cases, all pulling in different directions

Contrary to popular trends, we didn’t actually begin at the hardware or infrastructure layer.  We started with understanding our own use-cases. That was going to dictate all of our other decisions.

  • Chat

    • A chat style interface for conversing with our models for answering general questions, producing documentation and other engineering related artifacts

  • Agentic & Assistant Coding

    • Coding assistants and agents pointed towards an internal endpoint with very long context, heavy tool calling, streaming responses, and sessions that run for hours. This also has the clearest business case, since it's the workload where engineers are already pasting code into someone else's model.

  • CVE Scanning & Backporting

    • LLMs have exploded the amount of CVEs that we see on a daily and weekly basis, and we have automated pipelines that leverage agents to detect and patch this work automatically for us

  • Internal Tools

    • Upbound has lots of internal tools that leverage inference; customer health and satisfaction, support & triage, incidents & debugging, etc that are more general purpose but heavily leverage LLMs, multi-agent workflows & internal skills

We have many other internal use-cases at Upbound, and Internal Tools itself can be expanded to quite a few additional use-cases.  For the purposes of this campaign, we really decided to focus on the first three use-cases as we knew we’d learn a ton, and then could expand our Internal Tools use-cases over time with all of those learnings.

Requirements before models

The use cases wrote the requirements before we looked at a single model, and they split cleanly into what the workloads needed and what the platform needed. 

For coding & CVEs, the model had to handle long context without degrading and parse tool calls and skills reliably, because an agent that misreads a tool call five steps into a task is worse than no agent at all. It had to stream, and it had to hold up across sessions measured in hours, not seconds. Claude Code also needed an Anthropic-compatible API. 

For chat, the priorities flipped: low time-to-first-token, real concurrency across many users, and an OpenAI-compatible API for the tools already built against one.

The non-functional platform requirements were the ones we cared most about, because they're what a customer's platform team will be judged on. The endpoint needed real TLS on a real hostname and identity restricted to our own Workspace domain. We needed observability good enough to tell a healthy engine from a busy one. We needed the scheduler to put each model on the right hardware without us babysitting it. And every workaround had to be additive, living alongside what Modelplane owns rather than modifying it, so that anything we learned could go upstream cleanly.

Why AWS, and why that's just the start

Modelplane isn't tied to any one cloud. It runs inference clusters on any cloud, neocloud, on-premise environment, or any combination of them.  Modelplane can provision the clusters for you or let you bring your own. Under the hood that's Crossplane providers doing what they do best: in our own validation runs, applying an EKS InferenceCluster activated just the AWS kinds and scaled up only the four AWS providers, while GCP, Azure, Nebius and Vultr stayed at zero. Every provider is there when you need it and costs you nothing when you don't. 

We started on AWS because they’re the largest cloud provider and already where many of our customer workloads run in addition to their GPU workloads.  If you ever find yourself in prison, go after the biggest person in the yard first.  It was also the hardest version of the test we wanted to run: a control plane hosted in Upbound Cloud Spaces, driving GPUs in a separate AWS account we own, with all the IAM, networking, and capacity realities that come with it. 

AWS is the first stop, not the last.  That's where it gets interesting, because the real promise of a control plane isn't running inference on one cloud. It’s deploying inference across several clouds, neoclouds, and your own GPUs and moving a workload from one to another by changing a line of YAML, and routing across all of them at once.  This is the definition of multi-cloud.

The models picked the hardware

We decided to start with our first three use-cases: Chat, Agentic & Assistive Coding and CVE Scanning & Backporting. We landed on Qwen3-Coder-30B-A3B for coding & CVEs and GLM-4.7-Flash for chat. Both are MoE (mixture of experts) models at roughly 30B total and about 3B active, both served with vLLM at FP8.  We intentionally started with smaller models for a variety of reasons and to understand why, you need to grok some basics about modelplane and how it is structured across three layers:

  1. What the platform team creates: InferenceCluster, InferenceClass, InferenceGateway

  2. What the model operator creates: ModelDeployment, ModelService, ModelCache

  3. What the end-user consumers: ModelEndpoint, ModelReplica

At Upbound, due to our platform and infrastructure experience knew we’d have no problems or many new  learnings with #1.  We know infrastructure and how to deploy and operate it at scale.  With #3, there’s a low surface area here as this is basically what humans and agents consume and we’d had alot of experience here with our current usage of Frontier Models.

So our biggest unknown was with #2.  How do we deploy and operate these models at scale?  How do we tune them?  How do we deal with upgrades and changing models?  What are the things we need to do with regards to model reliability (think SRE)? What do we do when models crash and how do we debug them?  The list goes on, and these were all unknowns so by starting with smaller models and then tuning and changing them over time we hoped to develop more of this operational experience that has the added benefit of being input into our product roadmap.  You wouldn’t start developing a new application in dev by starting with the largest database, this was our same approach with starting with smaller models. 

One constraint that shaped everything was putting the workloads on different GPUs: Qwen on an RTX PRO 6000 Blackwell, GLM on an L40S. Both would have fit on a single GPU, and consolidating would have been cheaper. It also would have proven nothing. Hardware-matched scheduling is what a control plane should get right and you can't test it with one GPU. 

It worked the first time. The DRA device selectors did their job: GLM held its architecture == "Ada Lovelace" pin, and Qwen's memory >= 70Gi bound only to the 96 GiB card.

What we learned, and why it’s only the beginning

Two big lessons arrived before anything served a token:

Lesson #1: Capacity is not quota. 

We had quota for the instances we wanted, and AWS still wouldn't launch them. One GPU size was declined for capacity in three separate availability zones before we stepped down to a smaller one. Quota is permission to ask. Availability is a separate conversation, and while your model choice should happen before your hardware choice, it sometimes may be influenced by and sit downstream of it.

Lesson #2: The newest model isn’t always the best option. 

When we revisited this months later, the flagship small models were 320B and 80B parameters. MoE makes 18B active sound affordable, but every weight has to be resident, so 320B total means about 320 GB at FP8 and four of our largest cards. The 80B model's own docs ask for two GPUs and suggest cutting context if it won't start. Meanwhile, both of our existing models were running at half their supported context length. The biggest win available to us didn't require a new model at all.

What's next

Picking models and hardware was the easy part. Modelplane installed on Upbound Cloud Spaces and served traffic in about twenty minutes, and the serving layer has been solid ever since. Everything around it was a different story.

In Part 2, I'll walk through what happened once we ran it for real: streams that looked finished but were silently truncated, monitoring that sat pending for six days while reporting healthy, a GPU that "fit" right up until it didn't, and an authentication story that works fine for humans and browsers and falls apart for agents. Every one of those is a gap customers would have hit, and every one is landing upstream.

The champagne is drinkable. We're still finding the corks.  Let’s continue the Champagne Campaign!

About Authors

Sumbry

Subscribe to the
Upbound Newsletter

Subscribe to the
Upbound Newsletter

Subscribe to the
Upbound Newsletter

Related

Related

Posts

Posts

Aug 19, 2026

Upbound Insights told us 230 resources were broken. Hub told us why.

Sumbry

Aug 19, 2026

Upbound Insights told us 230 resources were broken. Hub told us why.

Sumbry

Aug 19, 2026

Announcing Upbound v3: one view, API, and governance model for every control plane you run

Upbound

Aug 19, 2026

Announcing Upbound v3: one view, API, and governance model for every control plane you run

Upbound

Jul 23, 2026

Anthropic is subsidizing our AI coding at 13x. How long will it last?

Bassam Tabbara

Founder and CEO

Jul 23, 2026

Anthropic is subsidizing our AI coding at 13x. How long will it last?

Bassam Tabbara

Founder and CEO

Get Started with Upbound Crossplane 2.0

Trusted by 1,000+ organizations and downloaded over 100 million times.

Get Started with Upbound Crossplane 2.0

Trusted by 1,000+ organizations and downloaded over 100 million times.

Get Started with Upbound Crossplane 2.0

Trusted by 1,000+ organizations and downloaded over 100 million times.