Opens in a new tab
vmblog logo 2024 wht (updated)

Multi-cluster Service Mesh is Still Hard

Share: 

David Marshall | Published: September 29, 2021

By Lin Sun, Director of Open-Source, Solo.io

A year ago, I wrote about how service mesh is still hard for our users for a single cluster. While observing the great adoption of service mesh in the industry for the past year primarily for single clusters, I noticed many users evaluating or/and onboarding service mesh for services running on multi-clusters and VMs. This is the natural next step in the adoption journey of service mesh. For most enterprise users, you don’t have all of your workloads running on Kubernetes and you can’t afford an outage of your service simply because the single cluster where you run your service is down or overloaded.

If you are considering multi-cluster service mesh, I want to highlight a few key challenges for you to consider in this blog:

1. How are services discovered in multi-clusters?

As you grow your services to run in 1+ clusters, your services running in all clusters need to be discovered frequently as your services scale up and down so that endpoints for your services can be intelligently selected based on their locality and availability. Who is going to discover the services across different clusters for you and what access control are you willing to grant to enable the service discovery across clusters?

If you follow the instructions to configure Istio multicluster, you will notice that you need to give the remote cluster’s Kubernetes API server read access credential (https://istio.io/latest/docs/setup/install/multicluster/multi-primary_multi-network/#enable-endpoint-discovery) to the primary cluster(s) so that the primary cluster(s) can discover service endpoints on the remote clusters. If you have a simple 4 clusters environment where each cluster is a primary cluster where the Istio control plane runs, you will need to grant all remote Kubernetes clusters’ API server access to each cluster like below:

istio 

This diagram looks complicated even for 4 clusters, what if you have a few dozens of clusters? Our enterprise users have shared with us security concerns with this approach because they don’t want the Kubernetes API server access credentials to leave the cluster which could be maliciously used for other purposes. A service mesh management plane like Solo’s Gloo Mesh could be introduced to solve this problem:

istio-service-mesh 

2. Can you selectively expose your services outside of your clusters?

You don’t always want to have all of your services exposed out of the cluster and make them available to other clusters by default. The best security practice is to allow nothing by default and gradually enable access control as required. Do you have your services deployed the same way to different clusters or do you have different services deployed to different clusters? How do you selectively configure what services are exposed outside of the cluster? How do you import only the required services into your cluster? Another added complexity is that not all service mesh supports these intuitively and consistently. Using Istio as an example, there are different ways to do this:

  1. You may use the hidden cluster_local mesh configuration to indicate which services are scoped to the local cluster it runs.
  2. An experimental implementation of Kubernetes multi-cluster service API is evolving in Istio to allow users to export certain services out of the cluster for clients outside of the cluster to consume.
  3. Vendors build abstractions for service mesh such as Gloo mesh’s virtual destination which allows you to declare your global service for multi-clusters.

3. Do you need a global service name for your service?

After you figure out the right approach to declare the intention to export your service, the next question is how your service consumers (clients) can reach your service. Not only you would want to have the connection secured with mutual TLS, but also you will need to provide a global service name for your service regardless of where you run your service. Your service can run in the same Kubernetes cluster as the clients or different clusters, or on VM or bare metal.  Your clients don’t need to know or care about where your service runs, all that matters to them is to call your services from a single global service name. Similarly, you should not need to care about where your clients run, clients could run in Kubernetes or outside, with or without the sidecar proxy. Can you configure a global service name for your service along with service failover priority? For example, if your service is healthy in the same cluster as the clients, it should be called first.  If your service is not available in the same cluster as the clients, calls should failover automatically to your service running at the next closest location without you needing to do anything extra.

 

4. How can you observe your services running in multi-clusters?

Your services may have some unplanned outages and you need to debug where the problems are. For example, you could set up a Prometheus federated cluster for aggregated metrics for services running in multi-cluster Istio service mesh. The alerts that are valid for your single cluster may not apply as you expand your services to run on multi-clusters. You probably don’t want to be alerted in the middle of the night when your service is down in one or two clusters yet still up in a few other clusters. Kiali allows you to view dashboards easily for single clusters and even multi-clusters with the caveat that you can only view details of the local cluster well.  There may be lots of tools and resources available for your preferred multi-cluster service mesh but can you assemble them successfully and securely so you can have a global view, drill into a problematic cluster for the detailed view, and know exactly what metrics to search for?

Wrapping up

I hope the above multi-cluster service mesh challenges resonate with you. This is where innovation is happening in service mesh and I am super excited to work with folks in Istio and Solo to tackle these challenges. I invite you to comment below to share with me your thoughts and connect with me on Twitter/LinkedIn.

##

To hear more about cloud native topics, join the Cloud Native Computing Foundation and cloud native community at KubeCon+CloudNativeCon North America 2021 – October 11-15, 2021     

ABOUT THE AUTHOR

LIN SUN 

Lin Sun is the Director of Open-Source at Solo.io. She has worked on Istio service mesh since 2017 and serves on the Istio Technical Oversight Committee. Previously, she served on the Istio Steering Committee for three years and was a Senior Technical Staff Member and Master Inventor at IBM for 15+ years. She is the author of the book “Istio Explained” and has more than 200 patents to her name.