All posts by Vikas Shitole

About Vikas Shitole

Vikas Shitole is a Senior Tech Lead at VMware by Broadcom, VCF division, India, where he leads system test efforts—including scale, stress, and resiliency testing—and drives product quality across VMware Cloud Foundation (VCF), Broadcom’s flagship private cloud platform. He is an AI and Kubernetes enthusiast, and is passionate about VMware customers and automation around vSphere and VCF. Vikas has been honoured as a vExpert for 13 consecutive years (2014–2026) for his sustained technical contributions and community leadership. He is the author of two VMware Flings, holds multiple industry certifications including VCF admin 9.0, and is one of the top contributors to the VMware API Sample Exchange, where his automation scripts have been downloaded over 50,000 times. Vikas has shared his expertise as a speaker at international conferences such as VMworld Europe and VMworld USA, and was selected as an official VMworld 2018 blogger. He also served as lead technical reviewer for the Packt-published books vSphere Design and VMware Virtual SAN Essentials. Beyond tech, Vikas is a dedicated cricketer, cycling enthusiast, and a lifelong learner in fitness and nutrition, with the personal goal of completing an Ironman 70.3

My learning journey: VMware Cloud Foundation & tips for VCF specialist exam

VMware Cloud Foundation (VCF) architecture is the backbone of VMware Cloud. I had worked on VMC on AWS, which is VMware’s hybrid cloud offering but I never had an opportunity to deep dive into VCF as a product, the platform for Multi-cloud and Modern apps. Recently I started exploring more on VCF and I thought to start my learning journey with VMware Cloud Foundation specialist exam as one of the milestones. Learning on why/how customers plan, design, deploy VCF (a SDDC stack)  & perform day1, day2 ops really helped me appreciate what VCF brings to the table. I loved how VCF orchestrates and simplifies SDDC stack deployment and its management. In this post, I am sharing super cool learning resources I went through so far, why I chose to write VCF specialist exam and tips to pass it.

Learning resources

Sr noDescriptionLink/URLsNumber of videosTotal time taken Comment
1VCF overview short videosVCF short videos1250 minIf you already know vSphere with Tanzu, skip 3 videos
2VCF technical deep dive with simulated UI hands onVCF deep dive880 minHCX and vRealize related videos you may skip as below we have detailed ones
3VCF life cycle workflows with simulated UI hands onVCF LCM workflows635 min
4VCF FAQsFAQsNA60 minThis is great read
5VCF solutions: HCX, vRA : simulated labs
vSphere with Tanzu session
vRA
HCX
Tanzu
130 minNote that HCX and vRA are simulated labs (UI hands on), kind of UI walkthrough. I really loved and only Tanzu is session
6VMworld 2020 and 2021 sessions VMworld 2020
VCF networking
VCF storage
Operational processes for VCF
VCF remote/edge clusters
VCF and vSAN stretched cluster
VCF and vVols
VMworld 2021
NSX design and operational recommendations for VCF with Tanzu
HCX with VCF
VCF tips and tricks from the trenches
9each video approx. 40 min 1. You will have to login with your free VMworld account
2. VMworld sessions definitely provide that real time best practices, challenges, customer perspective/view

7VCF official courses VCF plan and deploy : 2 days
VCF management and operations: 3 days
NA2-3 full days with labsThese courses will provide you comprehensive view of the VCF, super cool
8Setting up VCF yourself on a ESXi hostAutomated VCF Lab Constructor (VLC)
Part-1
Part-2
VLC session
NA1 dayTime taken is 1 day for entire preparation and then automatic VCF deployment
Download link: http://tiny.cc/getVLC and VLC slack channel: https://tiny.cc/getVLCSlack
9VMware official HOLVCF hands on LabNA3-4 hrsIf you can not setup your own lab as posted in #8, you can manage with HOL

Tips on passing this exam : my experience

  1. First thing you must be clear in your mind is that why do you want to write this exam. For me, getting comprehensive view of the VCF & hands on experience were aligned with my recent day to day focus in office. Also, VCF is one of the key focus areas for VMware as its the foundation for Multi-cloud, hence I thought let me target this exam to motivate & streamline my learning journey. I had written about Why I choose to target exams. I highly recommend you read that 7 line section.
  2. As mentioned in 1st point, idea is not only to get cool badge or get certified but also acquire VCF as a skill, hence I first started with top 5 sections/rows listed under Learning resources above. Overall it would take just somewhere around 200 min to 300 min of your time. Few things would be little repeated but it is worth to have that repetition.
  3. then I started with VMworld sessions listed in row #6. There are 9 sessions including 3 sessions from VMworld 2021 as well. You can choose to set the speed 1.25x or more as per your comfort
  4. At this stage, I was looking for getting more hands on experience deploying VCF on my own ESXi host and this is where VLC (VCF Lab Constructor) helped me immensely. I must say VLC team (specially Ben , Heath & SDDC commander) has done great work not only with building this automation tool but also building thriving VLC community. Beauty is VLC bits get updated with every new VCF release. Refer section #8 above for more articles around it.
  5. In each of the videos, you might see some VCF basic concepts are repeated but it is worth. You can always speed up the video or skip as needed.
  6. If you can not have your own lab, you can manage with VMware official HOL as listed in row #9. Some of this is already covered in simulated UI walkthroughs. However, playing around your own lab as usual gives you better learning experience.
  7. At this stage, I really started thinking about VCF exam and scheduled it around 2 weeks before . Please make sure you go through exam official pages here and here to clearly understand the certification requirements and exam objective/syllabus. You will be scheduling this exam from mylearn portal and then it will re-direct to “pearsonVUE” portal.
  8. There are 2 things: One is passing VCF specialist exam directly and another is completing all certification requirements to be “VMware Certified specialist : Cloud Foundation 2021“. You can decide how do you want to go about. In my case, I tried registering for VCF plan and deploy official 2 days course and I was lucky to get a seat being at VMware. Completing this course including labs made me ready for the exam and my confidence was high at this stage. There is another 3 days course as listed in row #7 but I did not register for it.
  9. I need not revise anything specifically before 2-3 days of scheduled date, it was all about the learning journey over 6-8 weeks made things fall in place. On 1st June, 2021 morning 7 AM IST, I had passed this exam. Of course, learning does not end with passing this cool exam but its continuous.
  10. Further, passing this exam did not award the VCF specialist badge as VCP 2021 was another requirement that I had to complete as my existing VCP certification was expired. I eventually passed VCP 2021 last week as well and earned this badge

Note that learning resources shared here are based on my own experience. There could be some other useful resources but I strongly believe these are gold resources for anybody who is just starting with VCF. In future I will update this blog post with VCF focused cool blogs & hands on lab resources as I experience them.

I hope this post was helpful. If it helped you, please make sure you share it with others. Also, please stay tuned for similar articles on VCP 2021 and vSphere with Tanzu exam/certifications.

Happy learning ! If you want to get updates on future posts, follow me on Twitter.

Deep dive: New REST APIs to manage VM service and Namespace self service

Couple of weeks back vCenter server 7.0 U2a (a monthly patch focused on vSphere with Tanzu) released with 2 super cool features. In this post, I would like to take you through new REST APIs introduced as part of these features & key notes around how these features/APIs behave.

Virtual Machine Service

As per me, this feature is one of the key features (like vMotion) in the history of vSphere. In brief, it enables managing virtual machines using Kubernetes control plane. With this feature, user can define desired state of VMs, virtual networks, virtual storage device. How cool is that if you can create vms (with customization) using simple kubectl command like “kubectl apply -f vm.yaml” ! To learn more about it, I highly recommend you to read this official deep dive blog and associated video.

Namespace Self-Service

Prior to this release, Supervisor namespace life cycle was completely managed by the vSphere admin. It was limiting the flexibility that k8s user had. With this feature, k8s user has ability to manage the life cycle (create/delete) of their own Supervisor namespaces while the resource constraints wrt cpu, mem, storage policy are still controlled by vSphere admin. Simply k8s user can run ” kubectl create ns ” to create their own namespaces in order to deploy k8s objects such as vSphere pods, guest clusters (aka Tanzu kubernetes cluster/TKC ) and now with this release VMs as well. Note that if this feature is not activated in your environment, vSphere admin can continue to manage Supervisor namespaces as it was prior to this release. To learn more about this feature, please go through this quick and cool blog post.

New vSphere with Tanzu APIs in action

If you are completely new to vSphere with Tanzu (specifically Supervisor cluster APIs exposed by wcpsvc service running on vCenter server) REST APIs, I highly recommend you to first read “Introduction to Supervisor cluster REST APIs ” post.

VM-class

VM-class is nothing but t-shirt sizes available for deploying vSphere pods/TKC/VMs under namespace created by k8s user or the traditional Supervisor namespaces (the ones created from H5C or REST API) created by vSphere admin. K8s user needs to pass these classes as part of TKC or VM creation yaml manifest file. Below is how we create custom VM class.

POST: https://{api_host}/api/vcenter/namespace-management/virtual-machine-classes
{
"cpu_count": 2,
"cpu_reservation": 0,
"description": "my vm class",
"id": "custom-class",
"memory_MB": 1024,
"memory_reservation": 0
}

Here is REST API super cool documentation for managing VM-class.

Associating VM-class and Content library to namespace

This is one of the super important operations must be performed by vSphere admin i.e. Every Supervisor namespace created either by k8s user through namespace self-service or namespace created by vSphere admin must have associated VM-class and Content library configured.
In order to configure this association, existing create namespace API is modified. Let’s see how to update existing namespace with a VM-class we created above and couple of existing content libraries that I had created already.

PATCH: https://{api_host}/api/vcenter/namespaces/instances/{namespace}
{
"vm_service_spec" : {
"vm_classes" : ["best-effort-xsmall","custom-class"
],
"content_libraries" : [
"62537201-8194-4d98-aea3-47b95f17077b",
"14268a2c-847b-4f84-9a1a-a97b524e263b"
]
}
}


“vm_classes”: They are simply the name (id) of the vm class. I have passed one custom vm-class and one default vm-class.
“namespace”: It is name of the existing Supervisor namespace to be updated.
“content_libraries” : These are content library ids that we can fetch using content library GET API. I have passed 2 content libraries one for VM service OVA images and another is for TKC/guest cluster OVA images

Here is REST API super cool documentation for managing supervisor namespaces

Key notes on VM-class/CL association

  1. VM class & Content library association with new supervisor namespaces is required for both TKC/Guest clusters as well as VMs created through VM service. However, TKC/Guest clusters content library can be configured at Supervisor cluster level as well (as it is from beginning)
  2. If Supervisor namespace is created prior to this release, all such namespaces will have all default VM-classes configured automatically, hence it will not impact any new or old TKC/Guest clusters. This automatic association happens as part of first k8s version upgrade at Supervisor cluster level
  3. In future releases, I personally expect associating VM-classes and content library gets further simplified so that vSphere admin is not forced to monitor new namespaces getting created and associate these mandatory constructs accordingly.

Namespace self-service workflows

As shown in this post, we need to activate this feature with setting controls such as cpu, mem, storage policy & users/groups. Let’s take a look at how to configure this using API. This is API doc for same i.e. Create a self-service template and further activate it.

POST: https://{api_host}/api/vcenter/namespaces/namespace-self-service/{cluster}?action=activateWithTemplate
{
"permissions": [ {
"domain": "vsphere.local",
"subject": "devops1",
"subject_type": "USER"
},
{
"domain": "vsphere.local",
"subject": "devops-group",
"subject_type": "GROUP"
}
],
"resource_spec": {
"cpu_limit": 10000,
"memory_limit": 20480,
"storage_request_limit": 204800
},
"storage_specs": [ {
"limit":104800,
"policy": "7327e17f-15fb-47a6-a82e-e7ec54839b59"
}
],
"template": "my-first-template"
}

“permissions”: Here we need to configure SSO users and groups. User/Group can come from vsphere.local or any custom identity source. Note that administrators group is configured by default, we do not explicitly need to configure any administrator.
“resource_spec”: These are the resource limits for all the namespace self service created by configured k8s users
“storage_specs”: The storage policies that namespace self service will get storage from.
“template”: Name of the template. Note that it must not have any space the string.

Here is REST API super cool documentation for managing self-service namespaces

Introduction of “OWNER” role

  1. There is new Supervisor namespace role is introduced i.e. OWNER. Earlier, we had only “EDIT” and “VIEW” roles. Note that this role is specifically introduced as part of Namespace self service feature. Users are not expected use it directly from H5C and even if they do, it will behave same as “EDIT” role from user standpoint.
  2. When k8s user creates supervisor namespace from kubectl after activating Namespace self service feature, every namespace by default will get this role. This enables create/delete these namespaces from kubectl itself.

Key notes on namespace self service APIs

  1. Currently H5C UI supports only one storage policy but using API you can configure multiple storage policies, so namespaces get multiple storage policies get configured automatically though Self-service template might show only one policy.
  2. If you see closely, there is storage storage “limit” param under “resource_spec” as well “storage_specs”. The limit in storage_specs applies to individual storage policy while limit in “resource_spec” is storage limit on all the namespace self service created by k8s user.
  3. Another important behavior is: above API can be used for 2 operations. One for initially creating the template and activating it. Second is updating the same template (only exception I see is we can not change the name of the template once created initially). Usually update operation is done via PATCH API but this API is an exception as both operations are done via POST API.
  4. Note that there are few more separate APIs for managing self-service namespace templates , here you can update the template with PATCH API.
  5. Currently, per supervisor cluster only one template is supported. This could be the reason name of the template need not be changed once its created. UI also does not provide an option for setting the name. If you activate this feature for the first time from H5C UI, template name would be “default”. In future release, we can expect support for multiple templates.
  6. One more important factor is that since currently only one template is supported, once you create template for the first time, there is no way you can delete the template but user can only update its configuration or simply deactivate this feature it.
  7. How do we know whether given namespace is created by k8s user or vSphere admin? When we make GET call on given namespace, it provides one Boolean property i.e. self_service_namespace
  8. When supervisor cluster is not in running state, user is not expected to activate/deactivate or update the template. Cluster may not be in running state when Supervisor cluster upgrade is is in progress or cluster is not in good state due to some issue. Even if you do, these ops will be will keep waiting for cluster to get into running state and the proceed as expected.

Automating above ops using SDKs

  1. It is important to notice that probably 70U2a release is first vSphere patch release, which introduces new APIs or modification to the existing API (specially vSphere with Tanzu REST APIs). Usually only major or update release had such changes.
  2. In order to write automation around above workflows, you must upgrade your SDK to the latest available on github.
  3. You can simply refer my post on Automating supervisor cluster operations through Java SDK. I highly encourage you contribute more samples around these features. Java SDK documentation for these ops are here: VM-class , Self-service namespace & templates
  4. Apart from Java, there is official python SDK as well.

VM operator

The VM service (even Namespace self-service) we explored above is specifically driven by wcpsvc service running on vCenter server. There is another key VM service kubernetes side component (runs as part of Supervisor cluster) as well i.e. VM operator. Beauty is that this critical component is completely open sourced. How cool is that!

Further learning

You can learn more around vSphere with Tanzu & its APIs here
Detailed post on VM service by Frank and Cormac here and here
Cool post by William here on how VM service capabilities can be used for cool use-case like Nested-ESXi.


If you have any query or comment, please feel free to post me on Twitter

Why wcpsvc service is running even when Supervisor cluster is not enabled on vSphere cluster?

Recently I got this question i.e. Why wcpsvc (Workload Control Plane) service is running even when Supervisor cluster is not enabled on vSphere cluster? I thought it is worth to share the answer with a quick post. Before we jump on to the list of reasons, let me share my understanding of what is the wcpsvc and what is its primary role. It is one of the services among several services running on vCenter server. It is primarily responsible for managing/orchestrating Supervisor cluster (which is key part of vSphere with Tanzu) workflows and functionality. In other words, it implements all the CRUD REST APIs for Supervisor cluster i.e. Enabling Supervisor cluster on vSphere cluster, updating Supervisor cluster to next available kubernetes version etc.

Now let us look at reasons why it should be up before enabling Supervisor cluster

  1. Since it orchestrates API for enabling Supervisor cluster, it has to be running before we call enable API.
  2. When you put host inside the cluster into maintenance mode, it has to check whether host being put into maintenance mode is part of Supervisor cluster (as one of the worker nodes). To understand what happens when we put host into maintenance mode, please have a look at my other article i.e. How to gracefully remove host from Supervisor cluster.
  3. Before enabling Supervisor cluster, there is REST API exposed by wcpsvc service to know whether given cluster is compatible or not (both Supervisor cluster with NSX-T and vSphere networking stack)
  4. If user wants to know how many distributed switches are compatible with NSX-T and in-turn wants to know whether NSX-T edge cluster is compatible with given distributed switch. There is API for the same as well.
  5. User wants to know what are the kubernetes versions supported by given vCenter server before enabling it.
  6. User wants to get an idea on sizing for the Supervisor cluster i.e. TINY, SMALL, MEDIUM, LARGE and network CIDR default sizes.
  7. H5C has some vSphere with Tanzu UI workflows and for it to dynamically showcase them, it has to make some API calls exposed by wcpsvc irrespective of whether Supervisor cluster is enabled or not

I hope you now got insight into why wcpsvc has to be running without even Supervisor cluster is not up. It might happen there are few more reasons wcpsvc has to be running, I will update this post as I understand more.

Further reading

1. Official documentation for vSphere with Tanzu is here
2. Automating supervisor cluster workflows using Java
3. Automation around Supervisor cluster

Automating key vSphere Supervisor cluster operations using Java SDK

Recently I got an opportunity to work as technical reviewer of the programming guide wrt vSphere with Kubernetes Configuration and Management (pdf here). As part of that, I had contributed few key Java API samples around vSphere Supervisor cluster into open source vSphere Automation Java SDK. I thought it is good to brief you about the same and also share key getting started tips around vSphere Automation Java SDK. If you are new to vSphere with Tanzu (aka vSphere with K8s) capability, I suggest you to go through below articles.

1. Introduction to vSphere Supervisor Cluster REST APIs
2. Python scripts to configure Supervisor cluster & create namespaces

Basically there are 4 Java API samples. If you are just looking for samples, just refer below links.

  1. Enable Supervisor cluster,
  2. Create Supervisor namespace and
  3. upgrading Supervisor cluster to next available kubernetes version.
  4. Disabling Supervisor cluster

Getting started tips : Enabling Supervisor Cluster with vSphere Automation java SDK

  1. It is assumed that you already have a base environment with NSX-T stack as specified here . Note that there is vSphere with Tanzu through vSphere network stack introduced with vSphere 70U1 as well. It is just that the sample I had contributed around enabling workload management feature (i.e. Configuring your vSphere cluster as Supervisor cluster) is applicable to NSX-T stack. I will contribute vSphere network stack based sample as well.
  2. Before you start, please understand this general REST API doc around Supervisor cluster enable operation. I love using this brand new developer portal. As a beginner, I would just spend around few minutes. It is fine if you did not understand few things, just move on.
  3. Now it is time to build your vSphere automation Java SDK environment in Eclipse. Steps for building it are explained here
  4. Once your eclipse environment is ready and added all samples into your eclipse project, you would also get all Supervisor cluster related samples I contributed into your eclipse project under “namespace_management” package. You can optionally spend few minutes, its fine if you could not, lets move on.
  5. In order to automate anything using vSphere automation Java SDK , it is important that you understand Java specific vSphere Automation doc. Note that this doc is different from the one mentioned in step 2 above. Again just spend few minutes understanding how it is organized and move on. As you spend more time, you will be comfortable referring it.
  6. Once you got fair idea around java specific doc as mentioned in step 5, now you can start looking at Java specific API doc for enabling workload platform i.e. Enabling vSphere cluster as Supervisor cluster. This is the doc I used to write this sample, just spend 1-2 min on this and move on.
  7. While you are going through doc in step 6, start co-relating it in parallel with actual java code sample around the same. Co-relating with doc and code will make your understanding better.
  8. Please closely look at what params are passed and EnableSpec is built.

Creating Supervisor namespace

  1. Most of the steps mentioned above applies here also. You can simply move to next step and take a look at this sample directly.
  2. This is the sample you need to understand for creating Supervisor namespace. Good thing is that same sample applies to creating namespace with either NSX-T or vSphere networking stack. With vSphere networking stack, there is one new param introduced i.e. networks on which you want to create this Supervisor namespace. Here is the Java Specific documentation around this API

Upgrading Supervisor cluster to next available kubernetes version.

  1. Upgrading Supervisor cluster deserves one separate post, which I will do as I get time but meanwhile you can learn more about it here
  2. Java sample for the same is here.
  3. Good thing is that same sample applies to upgrading Supervisor cluster with either NSX-T or vSphere networking stack.
  4. Note that we have REST API for upgrading multiple clusters in single API call as well. Refer doc here (upgradeMultiple method)

Disabling Supervisor cluster.

  1. Disabling supervisor cluster does lot more than usual disabling as it removes/deletes vSphere pods, Guest clusters/TKG also. Basically it makes cluster in a state that was just before enabling it. Hence be careful before calling this API.
  2. Java sample for the same is here
  3. Good thing is that same sample applies to disabling/removing Supervisor cluster with either NSX-T or vSphere networking stack.

This is all I have in this post. I hope it was useful post. If Java SDK is not your thing, you can simply automate Supervisor cluster ops using either python as I did or use official vSphere Automation python SDK or go for DCLI also. In addition, it is important to note that VMware has decided to deprecate vSphere Automation .NET and Perl SDK (VMware KB) to more focus on Java, python and Go SDKs.

How to gracefully remove host from NSX-T based Supervisor Cluster?

Recently got this question couple of times on how to gracefully remove host from existing Supervisor cluster, which is configured with NSX-T. I thought it is worth to write a quick post with detailed steps. Removal of the host from existing Supervisor cluster can be for multiple reasons, it can be for some maintenance or adding host into other supervisor cluster etc.

If you are still not aware of what is this Supervisor cluster, I would highly recommend you read this blog post.

Below steps are written assuming you have NSX-T based Supervisor cluster configured. In NSX-T based supervisor cluster, ESXi hosts are also working as kubernetes worker nodes, hence it is important they are properly removed from cluster. In new 70 U1 capability vSphere with Tanzu with vSphere (VDS) networking, ESXi hosts are no more worker nodes as there it is all about Tanzu kubernetes clusters (aka Guest clusters), hence removing host from VDS based Supervisor cluster is same as earlier.

Steps :

  1. Identify the host you would like to remove.
  2. Put the host in maintenance mode
  3. Since host is now going into maintenance mode, all the vSphere pods running on this node/host will be re-created on the other available hosts (DRS will take care of recommending proper host).
  4. VMs running on this host such as Supervisor control plane VMs or Guest Cluster VMs or any other normal VMs will be migrated to other available hosts as usual.
  5. Once host goes into maintenance mode, spherelet service on the host will be stopped but spherelet vib pushed to ESXi host will be still there. you can check spherelet vib and service status using below commands. SSH to host for running below commands.
  6. “esxcli software vib list | grep spherelet” & “/etc/init.d/spherelet status”
  7. When host goes into maintenance mode, in kubernetes term, it is called node is drained. i.e. this host is no more ready for taking any k8s workloads. When you run “kubectl get nodes”, it should be shown as “not ready”.
  8. NSX-T host transport node will be still as it is in configured state. You can confirm it from NSX-T UI or you can also check “nsx” vibs on host in maintenance mode.
  9. Now instead of removing this directly from inventory, move this host as standalone host into datacenter.
  10. As soon as you move this host as standalone host, spherelet vib on the host will be removed as well i.e. no spherelet service will be on it as well.
  11. Also, NSX-T host transport node will be un-configured as well automatically. i.e. All NSX-T vibs are removed at this stage. You can check it from NSX-T UI, it should be shown as “Not Configured”. This happens because while configuring NSX-T, we had applied host transport node profile on vSphere cluster level.
  12. “kubectl get nodes” or H5C UI (Cluster >> Monitor >> Namespaces >> Overview) should not show this host as worker node.
  13. Only reference now pending is, this host is still part of Distributed switch (DVS/VDS) you configured as part of Supervisor Cluster. You can remove it from usual DVS UI workflow.
  14. At this stage, this host is completely free to go for maintenance or add into another Supervisor cluster. You can remove it now from vCenter server inventory as needed.
  15. To add host back into same Supervisor cluster again, it is better to have host in maintenance mode, make sure you add this host back to DVS/VDS configured. Move this host into cluster, exit from maintenance mode and it will automatically install spherelet and NSX-T vibs and after few min, it will be ready for taking k8s workloads.

Further reading:
1. vSphere with Kubernetes (now vSphere with Tanzu) 101 is here
2. Official documentation for vSphere with Kubernetes is here
3. Automation around Supervisor cluster
4. Read about new 70 U1 capability vSphere with Tanzu with vSphere (VDS) networking
5. Automating supervisor cluster workflows using Java