Dear Kubernetes: You need to get your priorities straight


Published on: Sep. 14, 2026 @ 11:00
Written with ❤️ by drmorr

Comic-book style graphic that says "Dr. Morr says NO!"

I spent way too much time making this image, I am extremely proud of it, and I am 100,000,000,000% turning this into a laptop sticker.

Editor’s note: I’ve been informed that the introduction to this post is convoluted and confusing. If you’d like, feel free to skip it go straight to the first subsection.

When I was in grad school, I spent a LOT of time thinking about mixed integer programming; in fact, my PhD thesis was primarily focused on developing new techniques for solving (certain classes of) integer programs). In short, if you’re unfamiliar with this concept, an “integer program” is just a system of equations that encodes certain “state” as variables that can only take certain (integer) values. The goal is to find the “best” set of assignments to those variables so as to maximize some objective or goal.

See, the thing is, though, integer programming is hard. Like, really hard. Like, as hard as a lot of other really hard problems. And so when you’re confronted with one of these problems, one of the things you might do is “try to make it easier”. For mixed-integer programming, one of the ways you can make things easier is called the “Big-M Method”. The idea is, “Well, OK, it’s really hard to find integer solutions to the equation , but maybe instead we can replace that equation with and for some other variable and some extremely large constant ”.

The problem with this approach is that has to be “big enough” so that your solver can’t “cheat” and find a solution to these equations where . BUT, and this is the important piece for the rest of this entire post, which I promise is relevant, the bigger you make , the harder the problem is to solve in practice, specifically because doesn’t have any meaning or relationship to the problem you’re actually trying to solve1. And therein lies the problem with Kubernetes.

Weight, weight, don’t tell me

So I’m pretty sure at least all of you are going “WTF is drmorr on about today?” and are ready to ragequit your browser. But if you bear with me2 I promise I’ll make it all make sense.

Let’s back up. I work on this extremely complex distributed system called “Kubernetes”. And because Kubernetes wants to be all things for all people, there are a lot of different ways that any individual user can configure Kubernetes. And in many cases it is impossible3 for the authors of Kubernetes to know how end users might want to configure Kubernetes; questions like “is this thing more important than that thing?” might reasonably have different answers for different people. And whenever we run into one of those questions, instead of doing the hard work of understanding our customers, we just do the distributed-systems equivalent of Big-M and say “Screw it, we’ll just slap a priority on that sucker and let the users decide.”

Here is a (non-exhaustive) list of where I’ve encountered this behaviour:

I’m sure there are others, but these are ones that I’ve personally interacted with (and been bitten by) in the past. The problem in each of these cases is the same as the Big-M problem in integer programming: these numeric weights a) have no meaning except that which is assigned by the users of the system, b) changing the value of the weights can have unintended consequences in distant parts of the system, and c) the temptation is to make these weight values really, really large to signal “Hey this thing is REALLY IMPORTANT, don’t touch it”, and bigger numbers = bigger problems.

Let’s look at a couple of examples.

Example 1: kube-scheduler node scoring

The first example of how this can bite you is in kube-scheduler; here, the weights are being applied to the different tests that kube-scheduler executes as part of node scoring. Let’s imagine that there are five different such tests. For our purposes, a test is a real-valued function that takes some arbitrary inputs about the state of the cluster and the node under consideration; in other words, in math-y lingo, .

Each test is then given a user-specified weight , and the score for a particular node in the cluster is calculated as the weighted sum of the tests evaluated for that node. Moreover, the selected node is the one that maximizes this weighted sum:

Now, let’s suppose as the administrator of the cluster, you want to ensure that pods always get scheduled on cheaper nodes before they get scheduled on more expensive nodes6. Let’s say that returns the cost of the node, so you set to the largest negative number you can think of, say, -100. Then you go about your day, and all of a sudden you get paged because some pod got scheduled to an expensive node before a cheap node! What! How can this be???

Well, in this example, you didn’t care about any of the other tests, so you just left . Each of the other tests returns the value 300, and since the expensive node of course costs $10, the final node score is 200. Turns out that for all the other (cheaper) nodes, the other tests returned 1, so the weighted sum for those nodes was -96.

“Well, ok, that can’t be right,” I hear you say from all the way over here. “Lemme just adjust some of those other weights real quick!” And bam! You have just entered the 12th circle of hell without even realizing it.

Example 2: pod deletion cost and balanced consolidation

In the previous example, there is a single hypothetical user that is controlling all of the numeric weights. So even though the weights don’t have any semantic meaning to the system, there is a single user that can assign some semantic meaning to their choices. This second example is more insidious, because it involves multiple users with competing incentives getting to pick the weights.

In Karpenter 1.14, the Balanced Consolidation disruption mode was introduced. The high-level idea here is that Karpenter shouldn’t disrupt nodes where the benefits don’t outweigh the drawbacks. Since Karpenter’s whole shtick is “save money”, the authors used “the amount you would save by disrupting the node” for the “benefit” side of that equation. On the “drawbacks” side of the equation is “how much disruption does this cause”, and the entire calculation boils down to comparing the “benefits / disruption” ratio to 2.

The problem is that the “disruption” part of this ratio is (essentially) just a weighted sum of the pods on the node, and the owners of the pods (who may or may not be on the same team, and who may or may not have an antagonistic relationship towards each other) get to set those weights, via the pod-deletion-cost annotation. So if engineers on one team see their pods getting frequently disrupted by Karpenter, they’re going to increase the weight of those pods so that some other team of losers gets their pods disrupted instead. And then that other team is going to play a tit-for-tat game until both teams have set their deletion costs to INT_MAX and are furious with each other, and congratulations, you’re back in the 12th circle of hell7.

Who could have foreseen this?

I’d like to argue in this post that the underlying cause of all these problems (including the Big-M problem at the beginning of the post) boils down to two factors:

  1. Allowing (essentially) unbounded range on your weights
  2. Using unit-less values for your weights

The first problem is somewhat obvious: allowing big numbers means there is always the temptation to just say “well THIS thing is really important, so lemme 10x the weight!” And then later on when the math fails you, you go and 10x some other really important thing, and the cycle continues ad infinitum.

The second problem I think is more fundamental. You’re trying to compare fundamentally incomparable things by assigning them numeric weights. You, as the external operator, give those weights some semantic meaning, but math doesn’t know what that semantic meaning is: all math sees is numbers, and numbers are easy to compare8! The “right” answer here is “don’t compare incomparable things” but I admit that isn’t always practical, so my second-place recommendation is to always make sure the things you’re comparing have units. As we all learned in physics class, unitless numbers are bad.

For example, in the Karpenter balanced consolidation case, an arguably “better” notion of disruption might be something like “time needed to get all the pods on the node back to Running”. Then you can consider the ratio of “dollars saved / time spent rescheduling”, and that lets you start to semantically-encode some of your external priorities into your policy engine. It might be fine to spend 2 minutes rescheduling all the pods if you’re going to save $100, but if you’re only going to save 50 cent, it might not be worth it9.

My last suggestion here is, if you do have to make up some unitless integer weights, consider dramatically constraining the range of values those weights can take on. Nobody can reason about the difference between something with weight 147938964 and 147938965, but people can actually think about the difference between, say, 1 and 3. And as a corollary to this, if you find that you only need a few weight values, maybe see if you can get that number down to two, because then you can replace it with a boolean: I claim that nobody ever needs to set a pod deletion cost in [-2147483648, 2147483647]. All anybody ever actually cares about is “delete these pods first, then delete those ones”, which is just two states.

In closing, just remember, kids: the next time you think about adding in a user-configurable integer weight value to your distributed system (or your integer program), “Dr. Morr says NO!”

As always, thanks for reading.

~drmorr


  1. And also in part because large values of start to introduce all kinds of floating-point errors in your calculations. ↩

  2. https://strikerless.com/wp-content/uploads/2016/07/bear-with-me.jpg ↩

  3. Or at least, it’s really hard and everybody is lazy. ↩

  4. More-or-less ↩

  5. By the way, this is how people typically configure “warm pools” in Kubernetes: use something like the cluster proportional autoscaler to schedule a bunch of “do-nothing” pods with extremely low priority. When higher-priority workloads come in, they can evict the do-nothing pods and start doing work immediately. The do-nothing pods then will trigger your node autoscaler to scale up and refill your warm pool. ↩

  6. This isn’t a real Kubernetes scheduler priority, but it could be, and it’s useful for illustrative purposes. ↩

  7. As a fun aside, you’ve also entered the NaN’th circle of hell, because the Karpenter balanced mode computation does a bunch of numeric divisions, and big numbers are murder on floating point calculations. ↩

  8. [Citation needed] ↩

  9. Although he does make some pretty good rap. ↩