# How do you set an internal SLO for a platform team?

**An internal SLO** should describe a reliability promise for a user journey, not just the health of a component. Choose a critical action such as receiving a CI result, completing a deploy, querying observability data, or obtaining auth. Define a service level indicator from the user's perspective, set an objective based on consequence and cost, document exclusions, and decide how misses will affect priorities. The objective is useful only when it changes decisions.

        
## Why does a platform team need internal objectives?

        
Internal users build customer facing systems on the platform. When CI stalls, deploy fails, observability disappears, or auth is unavailable, product work and incident response stop. Calling the service internal doesn't make the impact minor. It changes who directly experiences it.

        
An objective creates a shared definition of acceptable reliability. Without one, every incident becomes a debate. Platform engineers may point to infrastructure uptime while users describe failed tasks. Leaders may request new features while operational risk accumulates. A clear objective connects these views and gives the team evidence for tradeoffs.

        
The goal isn't to create a reporting ceremony. The goal is to decide how reliable a journey needs to be, observe whether that promise holds, and respond when it doesn't. If nobody changes a plan after a miss, the objective is decoration.

        
## What should the SLO cover?

        
Start with a user journey that has a clear beginning and meaningful end. For CI, the journey might begin when a supported repository submits a valid change and end when a trustworthy result is available. For deploy, it might begin with an approved release request and end when the new version is healthy or safely rolled back.

        
A component measure can support diagnosis but may not represent success. The CI scheduler might be available while jobs wait too long. The observability store might accept data while queries return too slowly for incident use. The auth API might respond while incorrect policy blocks access. Measure what users need to accomplish.

        
Keep the first scope narrow enough to understand. Define supported users, workflows, and environments. Broad objectives with vague populations produce arguments about every exception. You can expand once collection and decision habits work.

        
## How do you choose a service level indicator?

        
| Journey | Possible indicator | User question |
| --- | --- | --- |
| CI | Valid runs with a trustworthy result in target time | Can I act on feedback soon? |
| Deploy | Supported releases that complete or recover safely | Can I release without expert rescue? |
| Observability | Queries returning current useful data in target time | Can I understand the service now? |
| Auth | Valid decisions returned correctly in target time | Can approved users act safely? |
| Service creation | Requests reaching a running owned service | Can I start without manual intervention? |

        
The indicator needs a valid event population, a success condition, and a measurement source. Review sampling, missing data, retries, and client behavior. An easy metric can be misleading if it excludes the failures users notice.

        
Pair aggregate numbers with examples. A strong monthly percentage may hide a daily period that blocks releases in one region. Segment by journey, user group, and time where consequence differs.

        
## How should the objective be chosen?

        
Begin with consequence. Ask what happens when the journey fails and how long users can tolerate it. A deploy path used many times each day may need a stronger objective than an occasional catalog update. An auth failure during an incident may carry greater risk than a delayed optional report.

        
Study current performance before announcing a target. A target far above the baseline may represent a useful investment goal, but it isn't yet a credible promise. Separate the present objective from an aspiration, and identify work required to close the gap. Otherwise teams will learn that the number has no operational meaning.

        
Reliability has cost. Greater redundancy, faster detection, simpler architecture, and more support coverage consume capacity. Discuss that cost with stakeholders. The right objective isn't the highest number. It is the level that matches user need and organizational willingness to invest.

        
## Which exclusions are legitimate?

        
Exclude events only when the boundary is understandable and doesn't hide ordinary failure. Invalid requests, unsupported environments, planned maintenance with suitable notice, and upstream events beyond the team's control may deserve separate treatment. Even excluded events can matter to users and should remain visible.

        
Be cautious with dependency exclusions. If product teams experience one platform journey, they won't care that a vendor caused the failure. The platform team may not control the cause, but it can own provider choice, fallback, communication, and recovery design. Track the full user impact even when contractual attribution differs.

        
Document exclusions beside the objective and review them after incidents. If a common painful event is always excluded, the objective is describing an idealized service rather than reality.

        
## How should error budget thinking guide work?

        
The gap between perfect performance and the objective creates room for ordinary failure and change. Treat that room as a decision signal. When performance is healthy, the team can take proportionate delivery risk. When it is consumed quickly, reduce risky change and prioritize causes that threaten the promise.

        
Don't turn the budget into a mechanical punishment. Freezing every change can prevent the repair needed to improve reliability. Classify changes by whether they reduce or add risk. Give service owners authority to choose the safest path and explain decisions to users.

        
Look at burn rate as well as a period total. A severe short event may require immediate action even when the monthly objective still passes. A slow pattern may call for planned improvement before it becomes a crisis.

        
## What should happen when the objective is missed?

        

          - Confirm the measurement and identify affected user journeys.

          - Communicate impact, current state, owner, and next update.

          - Stabilize the service and protect critical user work.

          - Review causes across technology, process, capacity, and ownership.

          - Choose corrective work based on repeated risk and consequence.

          - Explain which roadmap commitments change as a result.

          - Check whether the objective or indicator needs revision.

        

        
A miss should make priority cost visible. If the team adds reliability work, say what moves. If leaders choose to accept the risk and keep feature commitments, record that decision. Hidden tradeoffs make the platform team appear unreliable and leave users without context.

        
Use learning language, but retain accountability. A blameless review doesn't mean every choice was acceptable. It means the organization examines how choices made sense in context and changes the conditions that produced avoidable risk.

        
## How do you communicate an internal SLO?

        
Publish the journey, population, indicator, objective, time window, exclusions, owner, and current result in one accessible place. Add a plain explanation of what users can expect and where to report a problem. A dashboard alone isn't documentation.

        
Discuss the objective with representative users before finalizing it. Ask whether the measured success matches their experience. Product teams may reveal that a CI result is technically complete but arrives after their useful feedback window, or that deploy success hides manual cleanup.

        
Report trends and decisions, not only a colored status. Explain major misses, improvement work, and known limitations. Users can handle imperfect reliability better when the platform team is accurate, responsive, and explicit.

        
## How do you avoid common SLO mistakes?

        
Don't begin with every service. A large catalog of weakly understood objectives creates maintenance work without better decisions. Start with one critical journey, establish trustworthy data, and practice the response. Expand from a working habit.

        
Don't copy an external service target without context. Internal users may have different alternatives, timing, and consequence. Don't choose a target solely because current performance already meets it. The objective should express needed reliability, while the baseline shows the work ahead.

        
Finally, don't isolate reliability from product planning. An internal SLO is a compact agreement about user value and engineering investment. Review it alongside roadmap, support demand, incidents, and adoption so platform leadership sees one system.

        
## How should the first objective be introduced?

        
Run the proposed definition against recent events before publishing it. Ask whether known CI delays, deploy failures, missing observability data, or auth incidents would appear in the indicator. If users remember serious pain that the measure classifies as success, revise the population or success condition. Historical examples are a practical test of meaning.

        
Introduce the objective as a decision tool, not a judgment of individual engineers. Explain how the team will gather data, which choices a miss may affect, and when the definition will be reviewed. Invite users to report gaps between the dashboard and their experience. This helps build confidence without pretending the first version is perfect.

        
Set a trial period long enough to observe normal variation. During that period, practice communication and priority review even if the objective isn't yet formal. At the end, confirm the target, adjust it with evidence, or improve instrumentation before making a stronger promise.

## Frequently asked questions

### What is an internal SLO?

An internal SLO is a measurable reliability objective for a service or journey used by people inside the organization.

### Which platform journey should get an SLO first?

Start with a critical journey where failure has meaningful impact and where the team can gather a trustworthy indicator.

### Should an internal SLO promise perfect availability?

Usually no. The objective should reflect user consequence, realistic system behavior, and the investment required for greater reliability.

### How often should a platform SLO be reviewed?

Review performance regularly and revisit the objective when user needs, architecture, or business consequence changes.

### What happens when an internal SLO is missed?

Assess user impact, communicate clearly, address urgent risk, and use the miss to reconsider priorities and reliability investment.

Related: [managing a platform engineering team](https://isilta.com/blog/how-to-manage-a-platform-engineering-team/), [what a platform engineering team is](https://isilta.com/blog/what-is-a-platform-engineering-team/), [handling internal customers](https://isilta.com/blog/how-to-handle-internal-customers/).
