SOURCINGBLOX DEMake an appointment
Menu
← Back to Blog Zscaler Operations · Shortage of skilled workers

Zscaler Operations Under Skills Shortage: Seven Warning Signs of a Fragile Operating Model

Seven warning signs of an overloaded Zscaler operation—and how roles, runbooks, monitoring, and managed services create stability.

SourcingBlox GmbH · 5 min read
Zscaler Operations Under Skills Shortage: Seven Warning Signs of a Fragile Operating Model

A security platform can be set up in a technically correct way and still be fragile in everyday life. This happens when knowledge is in the hands of a few people, routine changes escalate and no one is clearly responsible for quality, monitoring or policy lifecycle.

The shortage of skilled workers is therefore not only evident in vacancies. It manifests itself in waiting times, exceptions, rework, and a growing dependence on individuals.

The following seven warning signs help to examine the operating model.

1. Every change ends up with the same person

If only an administrator can confidently assess rules, certificates, or connectors, there is a risk of concentration. Holidays and parallel projects immediately become a bottleneck.

The solution is not necessarily a larger team. Documented decision criteria, roles, the four-eyes principle and standardized change paths often help. Recurring activities become delegatable, while critical decisions remain with experienced people.

2. The Service Desk can only forward cases

A ticket contains "Access is not possible" but no information about the user, device, policy, destination, time or error pattern. The second level starts every diagnosis from the beginning.

A reliable model defines minimum information and initial test steps. The service desk doesn't have to become a Zscaler admin. But it needs enough context to identify known causes, capture data in its entirety, and escalate in a targeted manner.

3. Exceptions have no owner and no expiration date

SSL inspection bypasses, temporary releases and special groups grow over years. Without a documented reason and resubmission, the inventory becomes unverifiable.

Each exception should have at least a technical owner, technical reason, risk decision and review date. Unexplained entries are not deleted across the board, but examined as a priority.

4. Monitoring shows events, but no action logic

Many dashboards do not automatically mean operability. The decisive factor is whether a signal is followed by a clear next action.

Critical events require a threshold value, who is responsible, an initial review, an escalation path and a completion criterion. This is how telemetry becomes a process. Without this logic, monitoring remains another system that someone has to open occasionally.

5. Policy changes are technically approved, but not professionally approved

A rule can be syntactically correct and still miss a business requirement. If application owners or risk owners are not involved, the security team carries decisions that it cannot make on its own.

A good change model combines application, justification, test case, release, technical implementation and follow-up. There are templates for standard cases; Special cases deliberately remain visible.

6. Updates and certificates are only noticed in the event of malfunctions

Connector versions, client lifecycles, certificates, and integrations create repetitive work. If there is no predictive overview, plannable maintenance becomes an incident.

An operating calendar with owner, lead time and dependencies often creates more impact here than additional ad-hoc capacity. The connection to the test environment, rollback and communication channel is important.

7. No one can say if operations will get better

Without measuring points, optimization remains subjective. Suitable key figures are not only ticket numbers. More meaningful can be:

Key figures should support decisions, not reward activity.

From bottleneck to service design

A managed service is valuable if it closes these gaps in a targeted manner. A blanket promise "we will take over the operation" is not enough. Scope of service, roles, transfer points, service times, response logic, reporting, change procedures and exit rules are required.

SourcingBlox therefore starts with an operational analysis: Which platform parts are in scope? What activities actually occur? Where is knowledge and access? Which processes can be standardized? Which decisions must remain with the customer?

This is followed by a prioritized improvement plan, runbooks, governance and, where appropriate, a clearly defined managed operating model. Specific service times and response times are only promised after performance and contract review.

The goal is not to replace in-house expertise. It is to relieve experienced individuals of repetitive coordination and routine work so that they can focus on architecture, risk, and difficult cases.

A 30-day check for the existing operation

Before designing a new service model, it is worth taking a limited look at real work. Over 30 days, no people are evaluated, but work patterns are recorded:

This sample already shows whether the main problem is capacity, lack of context, unclear responsibility or unnecessary variance.

Standardize routine, maintain decision-making authority

Standardization does not mean rigidly automating every case. Rather, it separates three types of work.

First, there are recurring checks with a clear decision: Collect status, collect minimum data, check known error path. This work can be accelerated by running books or suitable tools.

Second, there are controlled standard changes. You need template, permission, protocol and, if necessary, release, but not a new architecture discussion every time.

Third, real risk and architecture decisions remain. They consciously belong to experienced managers. If routine is properly intercepted beforehand, more time is available for this.

A sensible service design remains verifiable

For each service module, the entrance, result and boundary should be visible. An example:

This clarity reduces friction between the internal team and the service provider. At the same time, it prevents "managed service" from being understood as unlimited responsibility.

What improvements should be visible after 90 days

After a transition period, results must be measurable. For example, you can expect more complete tickets, fewer unexplained exceptions, clearer maintenance leads, and a lower proportion of cases that escalate only because of a lack of context.

If these effects do not occur, more capacity will not simply be added. Scope, runbooks, accesses, telemetry and roles are checked again. A resilient operating model learns from actual work.

Classifying the next step together

Check your Zscaler operating model for bottlenecks and handoffs

Make an appointment for a consultation

Sources and further information