Async Supervisor Upgrades
Kubernetes upgrades in vSphere were tied to vCenter releases, a slow and risk-heavy cycle. I designed the experience that lets platform teams upgrade Supervisor on its own, with clear upgrade paths, explicit system status and guided validation.
- Role
- Product designerLed UX direction
- Span
- 3 monthsCross-release initiative
- Scope
- Upgrade flowsSupervisor and workload management
- Research
- 9 platform operatorsJourney map, discussion, survey
- Prototype
- Open in Figma (opens in a new tab) ↗
§1Context
With Supervisor clusters, vSphere runs Kubernetes next to traditional infrastructure. Platform operators are responsible for stability, security and availability, and now also for giving application teams newer Kubernetes versions quickly.
Kubernetes releases often. Infrastructure upgrades are slow and careful, with validation, dependency checks and planned downtime. Because Supervisor upgrades were tied to vCenter releases, operators had to run full infrastructure upgrades just to reach a newer Kubernetes version.
Upgrading had become a high-friction, high-risk operation in which operators made complex decisions with little control or clarity.
§2Role
I defined the end-to-end experience for asynchronous upgrades: upgrade flows, system communication and interaction patterns across operational scenarios. That meant leading UX direction, facilitating cross-functional discussions with product and frontend and backend engineering, and turning research, technical constraints and requirements into one coherent experience.
Later I brought in a copywriting architect to keep terminology consistent with the mental models operators already had. A recurring part of the work was aligning strong technical voices on solutions that stay accurate without exposing low-level detail users do not need.
§3Research
The project started from direct customer feedback asking for more flexible upgrades. I mapped the full journey, from learning about a new release through preparation, execution and the days after. It surfaced emotional friction that the system view alone did not show.
When the standard route fails, find another one
I scheduled two open research sessions. Nobody signed up, because the profile was so specific. So I asked to join the weekly customer calls product management already ran, where the right people were present. Nine platform operators took part, through open discussion and a structured survey.
What we learned: operators treat upgrades as controlled, risk-sensitive operations and often delay them until stability is proven. They want control over when and how changes happen, not release cycles deciding for them.
§4Challenges
A new behavior in an old mental model
Supervisor could now be upgraded independently, but nothing in the interface said so. Users still thought in vCenter-driven upgrades. The change had to become explicit without adding complexity.
Visibility versus noise
A light alert would add little noise but could be missed in a high-risk context. I chose a persistent informational banner with a link to release notes.
- Why
- Awareness and confidence matter more than minimalism here.
- Cost
- A more visible element in an existing view.
A moving technical foundation
APIs changed during development and content library capabilities had to be extended. The interaction design stayed flexible while keeping flows consistent.
Words operators already use
New concepts needed terms that were technically accurate and matched existing platform language. Working with a copywriting architect kept operators' mental models intact.
§5The framework
An upgrade is a decision process, not a single button. I structured it in four stages so operators know when, why and how to upgrade.
- Stage 1Awareness
Independent Supervisor upgrades are visible and understandable.
- Stage 2Decision
Upgrade paths and version relationships support an informed choice.
- Stage 3Validation
Pre-checks and system feedback confirm compatibility and readiness.
- Stage 4Execution
A controlled flow with explicit status at every step.
§6Solution
Two starting points, one framework
Greenfield: operators enabling Supervisor for the first time, from content library assignment and configuration to activation. Brownfield: operators who already run Supervisor and come back to manage upgrades, content library changes and system updates. Each needed different entry points, hierarchy and status, on the same underlying framework.
Awareness
A persistent banner in existing upgrade contexts keeps the new option visible and links to release notes, so operators understand the change before they act.
Decision clarity
Available Supervisor versions, how they relate to the current environment and which upgrades are possible are shown in the product, without hunting through documentation.
Risk mitigation
Pre-checks validate dependencies such as vCenter and NSX before an upgrade starts, moving risk detection to the beginning of the process.
Guided execution
A step-by-step flow through selection, validation and confirmation, with clear status at each step to reduce load and prevent errors in a multi-step operation.
§7Edge case
During design I found a failure that was easy to miss: a content library assigned to the WCP service could be deleted by another user, leaving the system broken with no visible cause.
Instead of letting operators find out by trial and error, a warning surfaces the problem and points to the fix. In research sessions operators confirmed the scenario was rare but plausible, and that they would want it handled.
In enterprise environments, rare does not mean unimportant.
§8Outcome
Validated in customer sessions facilitated by product management. Operators consistently confirmed that upgrading Supervisor independently from vCenter would significantly improve their workflows.
- Incremental upgrades let infrastructure teams keep stability while DevOps teams get newer Kubernetes versions sooner.
- Clear upgrade paths, explicit status and built-in validation reduce uncertainty during upgrades.
- Upgrades move from forced, full-cycle infrastructure cycles to controlled strategies that fit real operations.
§9Reflection
I worked alongside experienced engineers with strong views on system design. There was a pull toward exposing low-level detail and system logic. My job was to keep the solution technically accurate and understandable for the people using it.
When the standard research approach doesn't work, find a different path. Don't wait.