Writing/Systems & capacity
RequestsReplicasDatabaseUseful work

A resource can scale.
The bottleneck can stay right where it is.

Autoscaling
is not
scalability.

What has to be true for adding capacity to actually help?

The original essay, with interactive illustrations. All models use hypothetical numbers.

“you can tell how inexperienced an engineer is by the degree in which they believe in autoscaling.”

Sam Lambert · X · August 15, 2026 · quoted in the original essay

I saw this post making the rounds on X in the middle of my two-week vacation in Europe and, naturally, prioritizing relaxation during my time off, decided to spend a whole bunch of time thinking and writing about it.

It’s designed to have that effect, in a way. The post succinctly makes a judgement about one’s experience with enough vagueness around autoscaling to encourage engagement.

So here I am engaging with it.

A SWE or SRE will read the original post and reasonably say:

“I’m encouraged to have autoscaling enabled. Why is that bad?”

And it isn’t, of course. Autoscaling is useful, but it’s not a panacea.

It’s also a great bit of marketing by Sam Lambert, the CEO of PlanetScale, a company whose product is specifically designed to absorb a lot of the difficult operational work involved in running and scaling databases.

Everyone who reads “believe in autoscaling” is going to bring their own definition of both “believe” and “autoscaling” to it, and that ambiguity is useful. The disagreement it creates exposes how much we tend to pack into the word autoscaling, and how easily we let it stand in for scalability more broadly.

Sam’s follow-up suggests that was at least part of the point: the simple experience of “it just autoscaled” can sit on top of a whole lot of engineering complexity, while still describing only one part of what makes a system actually scale.

A follow-up post by Sam helped to clarify:

A customer described a PlanetScale database that “literally just autoscaled.” Lambert’s response: “so proud of this. it means we got the complex bit right.”

Sam Lambert · X · August 16, 2026 · exchange quoted in the original essay

The PlanetScale customer isn’t wrong. From their perspective, traffic increased, the database continued to work, and PlanetScale handled the capacity change. It really did “just autoscale.”

And Sam isn’t wrong either, making that sentence true is the product.

A common mistake is in an assumption that can come next: “My service autoscales, therefore my system is scalable”.

It’s usually not that explicit. More often: autoscaling is enabled, the box is checked, and the capacity conversation ends before anyone has defined useful work, validated the scaling signal, or tested what happens at the ceiling.

Autoscaling is always scoped

A good infrastructure abstraction should allow its consumer to use a simpler mental model.

The consumer sees:

load increases
→ database autoscales
→ application continues working

The platform provider has to think about:

signals selection
→ scaling thresholds
→ minimum and maximum capacity
→ provisioning
→ placement
→ routing
→ replication
→ warmup
→ failover
→ recovery
→ cloud-provider constraints
→ cost

If every PlanetScale customer had to understand all of that before trusting the product, PlanetScale wouldn’t have successfully abstracted the problem. So ‘it literally just autoscaled’ is evidence that a complex implementation produced a simple experience, not that the problem was always simple.

The danger comes when we extend that experience beyond the boundary of the abstraction.

A PlanetScale database successfully autoscaling doesn’t mean that the application as a whole is scalable. It means PlanetScale successfully handled a changing capacity requirement within the part of the system it owns. A good abstraction can hide the complexity of scaling a resource, but it can’t make the scope of that resource disappear.

Even inside PlanetScale, “autoscaling” doesn’t refer to one universal behavior. PlanetScale Vitess can automatically change the number of VTGate processes based on CPU utilization, within configured bounds. Network-attached storage can grow automatically. PlanetScale Postgres, meanwhile, doesn’t currently autoscale CPU and RAM; customers deliberately select and resize that compute themselves.

The dividing line is state. A VTGate is a stateless proxy, so adding capacity means starting another process. Adding database capacity generally means moving or replicating state: rebalancing data, catching up replication, warming caches. Network-attached storage can grow automatically because nothing has to move; the state expands in place. State is the reason “my database literally just autoscaled” was a remarkable sentence in the first place.

Figure 1. A resource is not the request path.

Requests flow through the replicas to one database A request crosses the load balancer, a replica, the connection pool and the database. Replicas sit inside the autoscaler's scope. The load balancer, the pool and the database sit outside it. Rust marks requests that wait or fail. Autoscaler scope 600/s waitor fail 200/s waitor fail200/s wait for a connection 400/s wait for a connection 600 requestsper second 600/s Outside its scope RequestsLoad balancerReplicasDatabasePoolUseful work
Useful throughput of 1,000 per sec
Waiting or failing

The autoscaler scales on busy request slots per replica. Target: 5 of 10.

10 of 10 slots busy, above target

Desired replicas:

1 of 5

Follow one request.

A request comes in through the load balancer, runs on one of the service’s replicas, borrows a connection from the pool and queries the database. It isn’t done until every hop is.

2 of 5

What the autoscaler can see.

The autoscaler watches the replicas and changes how many there are. The load balancer, the pool and the database are outside its scope. However many replicas ask, the database serves at most 600 requests per second.

3 of 5

Start with the resource.

Four replicas can handle 400 requests per second. With 1,000 requests arriving, the application is the first constraint.

Every request slot is busy, so the autoscaler’s signal reads 10 against a target of 5. It asks for twice the replicas.

4 of 5

Then add capacity.

The autoscaler doubles the replicas. Application capacity reaches 800 requests per second, but the database can serve only 600.

Requests waiting for a database connection keep their slots. The signal still reads 10, so the autoscaler doubles again.

5 of 5

Follow the whole path.

Double the replicas again. Useful throughput stays at 600, and 400 requests per second now wait at the database, where sixteen connection pools compete for it.

The signal still reads 10. Only the maximum of 16 replicas stops the autoscaler. More capacity in one place cannot remove a constraint somewhere else.

Illustrative model. Each replica serves 100 requests per second and has 10 request slots; a request waiting for a database connection keeps its slot. The autoscaler targets 5 busy slots per replica, up to 16 replicas, and computes desired replicas as current × signal ÷ target, as the Kubernetes Horizontal Pod Autoscaler does.

With a CPU target of 60% instead, it would stop near 10 replicas, because requests waiting on the database use little CPU. The signal you choose decides where the loop settles.

Autoscaling isn’t the same as scalability

I think what’s at the core of Sam’s original post is that autoscaling != scalability.

But why, exactly?

Oftentimes, scalability, elasticity, resilience, and efficiency all get collapsed into “scalability,” despite being separate properties that answer different questions.

Elasticity: Can the system add or remove capacity as demand changes?

Scalability: Does adding capacity actually produce more useful throughput?

Resilience: What happens before capacity arrives, when a dependency saturates, or after the scaling limit is reached?

Efficiency: How much does each unit of useful work cost?

Autoscaling primarily addresses elasticity. It can contribute to the other three, but it proves none of them.

On its own, an autoscaler is a delayed supply-side control loop with a relatively narrow job:

observe a signal
→ compare it with a target
→ change the amount of capacity
→ wait for the system to respond

The loop is simple, but the assumptions are not: that the signal actually represents demand or saturation, that capacity will arrive before the system exhausts its headroom, that the application can use the added capacity, that dependencies can absorb the additional work, and that scaling won’t make the current failure worse.

The metric is itself a capacity model

CPU, memory, request count, and queue depth are all observable metrics, but none is automatically the correct representation of load. Requests per second may work well when requests have reasonably similar costs, but it becomes less useful when one request consumes ten times as much CPU, memory, database time, or downstream capacity as another. Queue depth may look like an obvious scaling signal, but a queue of 10,000 one-millisecond jobs is a very different problem from a queue of 10,000 ten-second jobs.

Meta describes throughput autoscaling in terms of useful work: for a conventional web service, useful work might be requests per second; for an ML ranking service, the number of elements ranked per second, because the cost of individual requests varied too much for request count to describe the workload accurately.

That’s a more useful way to think about metric selection:

A scaling metric is a claim about what constitutes work and how much capacity that work requires.

And that’s one of the things a managed infrastructure provider generally can’t define for us. A database platform can observe CPU, query latency, connection usage, rows read, and I/O. It can’t automatically know that one query supports a critical checkout while another belongs to a background analytics job that could safely wait.

PlanetScale’s Database Traffic Control documentation draws this boundary explicitly: it can block queries that exceed a configured resource budget, but the application still has to handle that rejection through fallback behavior, circuit breakers, retry policy, or some other deliberate response.

The provider can supply the mechanism, but the application still has to define what the work means.

Capacity is a property of the request path

An autoscaler also has a fundamental limitation: it can usually change only the resource it controls.

Your application autoscaler may be able to add pods. It probably can’t also increase the write throughput of your database, a third-party API quota, the number of partitions in your queue, downstream connection limits, or the concurrency another service can safely absorb.

This leads to what I think is one of the most useful lessons in the whole discussion:

Capacity is a property of the request path, not of an individual deployment.

The end-to-end capacity of a request is constrained by the tightest point in a larger dependency graph. Scaling one component may move that bottleneck. It may expose a new one. In some cases, it allows the component to reach the existing bottleneck faster.

OpenAI’s PostgreSQL engineering write-up from earlier this year provides a great real-world example.

ChatGPT runs on a single primary writer with roughly fifty read replicas. Every replica added increases read capacity, and it also gives the primary one more instance to stream write-ahead log data to. As the replica count grows, that fan-out puts increasing network and CPU pressure on the one component that can’t be replicated away, which is why OpenAI is building cascading replication before it becomes the binding constraint.

And replicas were never the whole strategy. The same write-up covers query optimization, isolation of high- and low-priority workloads, connection pooling, protection from cache-miss storms, rate limiting at multiple layers, sharding of write-heavy workloads, and constrained retries that could otherwise amplify an overload event. The database was an important part of the capacity model, but it was only one piece of the puzzle.

The same is true of any component we buy rather than build. It may solve its own layer extremely well, but it doesn’t control the application tier, every upstream producer, every downstream service, or the complete path a business transaction follows.

A component can successfully autoscale while the system around it becomes less stable.

Figure 2. More replicas stop helping at the database.

Successful requests per second 1,200 Compute per 1,000 successful requests replica-seconds 30 Traffic 1,000 Replica capacity Database limit 600 Database limit 900 Unused replica capacity 6 replicas 9 replicas 10.0 26.7 17.8 1481216

Throughput at 6 replicas: 600 requests per second

1 of 4

Add replicas and throughput rises.

Each replica serves 100 requests per second. With 1,000 arriving every second, each replica you add completes another 100.

2 of 4

Until the database is the limit.

At 6 replicas the service completes 600 requests per second, all the database can serve. The seventh replica adds nothing, and neither does the sixteenth.

3 of 4

After that, replicas add cost, not throughput.

Throughput stays at 600 from 6 replicas to 16, so the compute behind every 1,000 successful requests rises from 10.0 to 26.7 replica-seconds.

4 of 4

Raise the database limit and the knee moves.

Let the database serve 900 requests per second and replicas help again, up to 9. Past that the database is the limit once more. Capacity belongs to the whole request path, not to one part of it.

Autoscaling is only one part of overload management

The limits become especially visible when the autoscaler can’t obtain additional capacity at all.

PlanetScale experienced exactly thisduring the October 2025 AWS us-east-1 incident. At one stage it couldn’t launch new EC2 instances, and customers using diurnal VTGate autoscaling were approaching the US workday with substantially less capacity than they would normally have had.

The response involved much more than waiting for autoscaling to recover: the team delayed backups that would have required additional instances, stopped terminating existing capacity, packed VTGate processes more tightly onto the machines already available, and advised customers to shed nonessential database load by delaying queue processing or pausing ETL workloads.

This is a particularly useful example because it comes from the same company whose customer said the database “literally just autoscaled.”

When the supply path worked, the customer received a wonderfully simple experience. When it was blocked, resilience came from an entire set of controls that didn’t add supply: available headroom, workload prioritization, bounded queues, delayed background processing, load shedding, tighter packing of existing resources, and isolation between control-plane and data-plane failures.

Autoscaling belongs inside this larger overload-management system.

Queues absorb temporary variance, but they don’t create throughput. Backpressure limits how quickly new work enters the system. Admission control and load shedding reject work before it consumes scarce resources. Graceful degradation makes accepted work cheaper. Retry control prevents failures from producing even more demand. Isolation stops one tenant, feature, or workload from consuming every shared resource.

The failure modes run in the other direction, too. Many real autoscaling incidents are scale-in mistakes: flapping between sizes, scaling in just before a burst, cooldowns tuned for a traffic pattern that no longer exists.

PlanetScale’s own capacity guidance recognizes the timing problem. For an anticipated launch or event, it recommends temporarily increasing cluster capacity ahead of the expected traffic rather than assuming reactive behavior will always arrive in time. Reactive scaling handles unexpected variance while pre-scaling handles expected demand.

Headroom and overload controls help by protecting the interval between demand arriving and capacity becoming useful while also protecting the system after its scaling limit has been reached.

Here is one burst handled three ways. Each policy decides what happens to work that can’t be served yet:

The burst, the five-second client timeout and the autoscaler’s 30-second delay are the same in all three.

Figure 3. An unbounded queue turns overload into lateness.

  • On time
  • Served late
  • Rejected right away
  • Expired unserved
Chart: requests per second against capacity. Requests per second 06001,200 Capacity arrives after 30 s Demand Capacity Waiting requests 010,00020,000 0306090120150180 s Queue everything Cap the queue at 2,000 Drop latework Outcomes, 84,000 requests each Queue25,60058,400Cap58,60016,000Drop late70,000

1 of 4

A burst arrives.

Traffic jumps from 200 to 1,000 requests per second for a minute. The service handles 400 until the autoscaler adds capacity 30 seconds later. Every client gives up after 5 seconds.

2 of 4

Queue everything.

Every request waits its turn. The queue peaks at 18,000 and the longest wait is 19 seconds. 58,400 requests are served after their client has given up, so that work is wasted. Only 25,600 finish on time.

3 of 4

Cap the queue.

Accept requests until 2,000 are waiting, then reject new ones right away. 16,000 requests get a fast error they can act on, and 58,600 finish on time.

4 of 4

Drop late work.

Keep queueing, but skip any request that has already waited longer than its client will. 14,000 expire without using any capacity, and 70,000 finish on time, the most of the three.

The hidden failure signal of cost

When an autoscaler adds capacity, the bill follows automatically; that’s how usage-based pricing works, on every provider. The visibility is valuable, but it’s still a resource-level view of cost: it tells us how much capacity we consumed, not whether that capacity produced proportionally more valuable work across the entire application.

Consider a hypothetical release:

Before:
1,000 successful requests
10 application instances
healthy latency

After:
1,000 successful requests
20 application instances
healthy latency

A conventional availability dashboard remains green, latency remains healthy. The autoscaler did exactly what it was configured to do, but the software now requires twice as much compute to perform the same amount of useful work.

The autoscaler has quietly converted a performance regression into spend.

This is one of the ways elasticity can create false confidence in a rollout. The system remains available because the infrastructure absorbs the regression quickly enough that users don’t experience it as an outage.

That’s a good outcome for immediate availability, but it’s still an engineering regression. It has also quietly consumed headroom: the service now sits closer to its scaling and dependency limits, and the next burst will reach them sooner.

This is why requests, CPU, and memory shouldn’t be the end of the measurement model. We need to define what constitutes a useful unit of work for the service: a successful request, a completed job, a transaction, a ranked item, an inference.

We can then measure cost against it: CPU-seconds per successful request, database work per transaction, GPU-seconds per successful inference, or compute cost per thousand SLO-compliant requests.

Figure 4. The bill shows the regression first.

Successful requests per second 1,000 Replicas Maximum 24 Compute per 1,000 successful requests replica-seconds 25 Remaining headroom Release Not on the availability dashboard 00:0006:0012:0018:0024:00

At the 15:00 peak

1,000 requests per second
replicas: 10 of 24
20.0 replica-seconds per 1,000 requests, up from 10.0

1 of 5

A normal day.

Traffic rises toward an afternoon peak of 1,000 requests per second, and the autoscaler adds replicas to follow it.

2 of 5

A release lands at 10:00.

The new version needs twice the compute per request. Nothing else changes: same traffic, same users.

3 of 5

The autoscaler absorbs it.

Replicas double to keep up. Latency stays healthy and the availability dashboard stays green. Without the faint line to compare against, this looks like a busy afternoon.

4 of 5

The unit cost shows it.

Replica count rises with traffic and with inefficiency alike. Divide by successful requests and only the inefficiency is left: compute per 1,000 requests steps up at 10:00 and stays about double.

5 of 5

And the headroom is gone.

At the 15:00 peak the service runs 20 of its 24 replicas, up from 10. The next burst will reach the limit much sooner.

The FinOps Foundation describes this as unit economics: connecting technology cost to a meaningful unit of work rather than evaluating total spend in isolation. Cost per useful unit should sit right beside latency and error rate.

Test beyond the autoscaler

This framing also changes what it means to test autoscaling. A test that proves the replica count increased has mostly demonstrated that the controller is connected. A more meaningful test strategy asks:

Then cap scaling and keep testing. The most important behavior may be what happens after no more capacity can be added.

Explorer. Run your own burst.

A one-minute burst hits a service that starts with 4 replicas. Set how big the burst is, what the database can take, how far and how fast the autoscaler scales, how long clients wait, and what the service does with work it can't serve yet.

1,000 requests per second
2,000 requests per second
12
30 s
5 s
When work can't be served yet

Outcomes, 84,000 requests

  • On time
  • Served late
  • Rejected right away
  • Expired unserved

25,600 of 84,000 finish on time (30%). 58,400 are served after their client gave up.

Waiting requests

060120180 s018,000Autoscaler acts

Peak: 18,000 waiting. Longest wait: 19 s.

Cost: 68.8 replica-seconds per 1,000 on-time requests

The model is simplified on purpose. See About the model.

Google’s SRE guidance recommends testing not just until a service reaches its capacity limit, but through the resulting overload behavior and recovery. Its broader guidance treats queues, load shedding, graceful degradation, and retry control as complementary mechanisms rather than substitutes for capacity planning. Does the service reject excess work predictably while continuing to complete the most important work? Does it accumulate a queue it can never realistically drain, or overwhelm a downstream dependency? Do health checks remove already scarce capacity? Does it recover automatically when demand falls?

Those questions tell us more about scalability and resilience than whether a replica count moved from ten to twenty.

They also describe how a scaling policy should be treated from day one. Autoscaling configuration isn’t just infrastructure boilerplate that a service inherits and then forgets, it’s an executable statement of how we believe demand, capacity, performance, and cost relate to one another, and that statement is a hypothesis which should be continuously recalibrated with production evidence.

So, should we “believe” in autoscaling?

The PlanetScale customer who says, “It literally just autoscaled,” is describing a successful product boundary. Sam is also right to be proud of that response: a customer mistaking a difficult systems problem for an effortless product capability is often evidence that the platform team got the complex part right.

The mistake is still the one from the top: “My service autoscaled, therefore my system is scalable.”

A managed service can absorb a great deal of complexity within its boundary. It can’t automatically define useful work for our business or scale every dependency in the request path. It can’t decide which work should be rejected during overload, or whether preserving availability required an economically unreasonable amount of infrastructure.

This is why I would word it differently than Sam. Experience shouldn’t make belief in autoscaling decline; it should make it conditional. We can trust (much more than before) an autoscaler we have tested past its limits and wrapped in the overload controls that take over when capacity can’t arrive.

The position I have is neither blind faith in autoscaling nor reflexive distrust of it.

It is understanding:

Autoscaling is a behavior of a resource, scalability is a property of the request path.

Good autoscaling makes capacity changes feel boring.

Good systems engineering understands everything that still has to be true for that boring experience to remain safe, resilient, and economically sustainable.

About the model

Every figure and the explorer run on one small model. It’s a teaching model with made-up numbers, not a benchmark.