Uptime, latency, and SLOs
This page describes Temporal's current operational practices and does not create any additional commitment, representation, or warranty.
Temporal maintains multiple internal service level objectives (SLOs) and an internal alerting system based on those SLOs. The alerting system accounts for all errors, not only the service errors that count against the SLA, and also monitors latency. Temporal personnel receive an alert when an SLO is not being met, and on-call engineers are paged, which often means that issues are resolved before you notice them.
For current system status and recent incidents, see Temporal Status. For the regions you can run in, see Service regions; for the full set of rate, resource, and configuration limits, see System limits.
Service levels at a glance
| Measure | Applies to | Target | Type |
|---|---|---|---|
| Uptime | Standard single-region Namespaces | 99.9% | Contractual SLA |
| Uptime | Namespaces using High Availability | 99.99% | Contractual SLA |
| Uptime | All Namespaces | 99.99% | SLO |
| Latency | Worker requests, p99 per region | 200ms | SLO |
| Visibility API availability | List and Count requests, rolling 30-day window | 99.5% | SLO |
| Cloud Ops API availability | HTTP and gRPC requests, monthly | 99.99% | SLO |
| Recovery Time Objective (RTO) | Namespaces using High Availability, during a cell, region, or cloud outage | Under 20 minutes | Objective |
| Recovery Point Objective (RPO) | Namespaces using High Availability, during a cell, region, or cloud outage | Under 1 minute | Objective |
Uptime and the SLA
Temporal Cloud publishes two uptime numbers for each deployment mode: the service availability it operates to, and the contractual service level agreement (SLA) it guarantees against service errors.
| Deployment mode | Service availability (SLO) | Contractual SLA |
|---|---|---|
| Standard single-region Namespace | 99.99% | 99.9% |
| Namespace using the High Availability feature | 99.99% | 99.99% |
The SLA that covers normal Worker requests, meaning commands and polling, also covers Nexus requests in both the caller and handler Namespaces.
Temporal Cloud runs a cell architecture. Each cell holds the software and services needed to host a Namespace, and the components in a cell are distributed across at least three availability zones per region. Whether errors during an outage count against the SLA depends on the type of outage and the High Availability features the Namespace uses. See Outages and Recovery Objectives for each outage type.
How the service-error rate is calculated
Temporal Cloud captures every request that arrives in a Namespace during a five-minute interval and records the gRPC service errors among them. For each Namespace, the service-error rate is 1 - (count of errors / count of requests). Rates are averaged per month and reset quarterly.
Errors counted against the SLA are service errors, such as the UNAVAILABLE gRPC status code.
The following errors are not counted against the SLA:
ClientVersionNotSupportedInvalidArgumentNamespaceAlreadyExistsNamespaceInvalidStateNamespaceNotActiveNamespaceNotFoundNotFoundPermissionDeniedQueryFailedRetryReplicationStickyWorkerUnavailableTaskAlreadyStartedThrottling (resources exhausted; triggers retry)WorkflowExecutionAlreadyStartedWorkflowNotReady
Latency
Temporal Cloud has a p99 latency SLO of 200ms per region.
That SLO covers normal Worker requests, meaning commands and polling, and it applies to Nexus in both the caller and handler Namespaces.
Measured latency
Measured latency over a week-long period for starting and signaling Workflow Executions, and for starting Standalone Activities (StartActivityExecution),
was as follows:
August 2026
| Operation | p50 | p90 | p99 |
|---|---|---|---|
StartWorkflowExecution | 20ms | 32ms | 78ms |
SignalWorkflowExecution | 19ms | 42ms | 91ms |
SignalWithStartWorkflowExecution | 30ms | 47ms | 109ms |
StartActivityExecution | 13ms | 19ms | 45ms |
January 2026
| Operation | p50 | p90 | p99 |
|---|---|---|---|
StartWorkflowExecution | 14ms | 21ms | 69ms |
SignalWorkflowExecution | 11ms | 19ms | 46ms |
SignalWithStartWorkflowExecution | 19ms | 37ms | 95ms |
March 2024
| Operation | p90 | p99 |
|---|---|---|
StartWorkflowExecution | 24ms | 54ms |
SignalWorkflowExecution | 14ms | 40ms |
SignalWithStartWorkflowExecution | 24ms | 61ms |
Latency observed from the Temporal Client also reflects other components in the path, such as the Codec Server, an egress proxy, and the network itself. Concurrent operations on the same Workflow Execution can raise latency as well.
Custom persistence layer
Temporal Cloud runs a custom persistence layer rather than the persistence stores available to self-hosted deployments. Three components of that layer account for most of the difference in latency:
- Sharding: Distributes load across multiple databases and resizes them independently, so a traffic spike in one shard doesn't become a bottleneck for the rest.
- Write-ahead log (WAL): Batches updates in an append-only log before writing them to the database, which reduces write latency and database size.
- Tiered storage of Event History: Moves the Event History of closed Workflow Executions to cheaper storage, which keeps the primary database smaller and faster for running Executions.
Visibility API availability
What availability can I expect from the Visibility API?
Temporal Cloud targets 99.5% availability for Visibility read requests, measured over a rolling 30-day window. This covers List and Count operations against Workflow Executions, Activity Executions, Nexus Operations, Schedules, and Batch Operations.
This target is deliberately lower than the availability guaranteed for core Workflow Service APIs. Visibility is a search index intended for operational discovery, not a critical path for application logic.
This is a service level objective rather than a contractual commitment. Visibility API requests are not covered by the SLA, and no service credits are associated with this objective.
The objective measures whether a request returns a valid response, not how current the results are. For what it excludes and how to interpret it alongside Visibility's eventual consistency, see Visibility availability in Temporal Cloud.
Throughput
Each Namespace has a rate limit measured in Actions per second, and two capacity modes control it:
- On-Demand Capacity raises the Namespace limit automatically as usage grows.
- Provisioned Capacity sets the limit from the Temporal Resource Units (TRUs) you request, which suits predictable spikes and guaranteed throughput.
See Capacity Modes for how each mode sets your limit.
Related
- Outages and Recovery Objectives (RTO / RPO): Recovery targets for each type of outage, and how outages count against the SLA.
- High Availability: Replication and failover options that raise a Namespace's SLA to 99.99%.
- Higher throughput and lower latency: Temporal Cloud's custom persistence layer: How sharding, the write-ahead log, and tiered storage are built.
- Benchmarking latency: Temporal Cloud versus self-hosted Temporal: Measured comparison of the two deployment options.
- Replay conference talk: Custom persistence layer: Walkthrough of the persistence design.