ScalerMOCK
AM
production / us-east-1

Scaling overview

Safe, focused control for native workloads and AWS SQS.

Simple by design

Native Kubernetes first · AWS SQS is the only external source · existing HPA ownership is protected.

Narrow scope
Active replicas
38
6 pending
Control sources
2
Native Kubernetes + SQS
Median API accept
420ms
Pod readiness tracked separately
Ownership guards
2
HPA workloads are read-only
Live operations

Current and recent operations

Timed attempts stay immutable; re-runs and reverts create new audit records.

2 running1 timed out
Runningop-0144
Scale · image-resizer
production · S3 event burst
48 replicas
3s elapsed15s limit
Retain 14 days
Older events purge nightly
Timed
Runningop-0142
Scale · events-worker
production · SQS policy
1824 replicas
7s elapsed30s limit
Retain 30 days
Older events purge nightly
Timed
Succeeded Reversibleop-0141
Scale · pdf-renderer
jobs · AM
08 replicas
3s elapsed10s limit
Retain 90 days
Older events purge nightly
Timed out Critical Slack sentop-0140
Scale · events-worker
production · Schedule
SCALE_UP_TIMED_OUT
1836 replicas
10s elapsed10s limit
Retain 14 days
Older events purge nightly
Scheduledop-0139
Schedule · pdf-renderer
jobs · Schedule · 23:30
80 replicas
Starts 23:3010s limit
Retain 30 days
Older events purge nightly

Workloads

Ownership is checked before every write

WorkloadReplicasSignalOwnerLast operationStateActions
checkout-api
production
12/ 12
CPU · 61%Native HPA
1.8s
18 sec ago
Stable
events-worker
production
18/ 24
SQS · 8.2k messagesScaler
2.4s
18 sec ago
Scaling
recommendations
production
8/ 8
RPS · 1.4kNative HPA
1.2s
18 sec ago
Stable
image-resizer
production
4/ 8
S3 · uploads bucketScaler
3.1s
18 sec ago
Scaling
fraud-alerts
production
3/ 3
CW Alarm · backlog-highScaler
1.6s
18 sec ago
Stable
ingest-pipeline
jobs
6/ 6
CloudWatch · DynamoDB WCUScaler
2.0s
18 sec ago
Stable
pdf-renderer
jobs
0/ 0
Schedule · 23:30Scaler
18 sec ago
Paused
Advisory

Add an SQS failure floor

Queue metrics failed twice this week. Hold six workers after three failed polls instead of relying on stale data.

events-workerfloor 6
After3 failed polls

Operation health

Rolling 60 minutes

99.3%
2 timeouts · 0 errors

Today’s schedules

All times in Asia/Jerusalem

08:52
Morning checkout peak
checkout-api · 12 → 20
17:45
Evening batch window
events-worker · 18 → 36
23:30
Render queue cooldown
pdf-renderer · 8 → 0
Suggested roadmap

Small features with clear ownership

Each addition keeps the controller focused instead of growing a generic trigger platform.

Now

Ownership guard

Detect HPA or KEDA before writes and create a transfer plan instead of competing for replicas.

Next

SQS failure policy

Choose hold-current or a safe replica floor when queue metrics become stale or unavailable.

Next

Readiness SLO

Measure API acceptance, scheduling, and Ready pods separately so latency claims stay honest.