Scaling overview
Safe, focused control for native workloads and AWS SQS.
Native Kubernetes first · AWS SQS is the only external source · existing HPA ownership is protected.
Current and recent operations
Timed attempts stay immutable; re-runs and reverts create new audit records.
Workloads
Ownership is checked before every write
| Workload | Replicas | Signal | Owner | Last operation | State | Actions |
|---|---|---|---|---|---|---|
checkout-api production | 12/ 12 | CPU · 61% | Native HPA | 1.8s 18 sec ago | Stable | |
events-worker production | 18/ 24 | SQS · 8.2k messages | Scaler | 2.4s 18 sec ago | Scaling | |
recommendations production | 8/ 8 | RPS · 1.4k | Native HPA | 1.2s 18 sec ago | Stable | |
image-resizer production | 4/ 8 | S3 · uploads bucket | Scaler | 3.1s 18 sec ago | Scaling | |
fraud-alerts production | 3/ 3 | CW Alarm · backlog-high | Scaler | 1.6s 18 sec ago | Stable | |
ingest-pipeline jobs | 6/ 6 | CloudWatch · DynamoDB WCU | Scaler | 2.0s 18 sec ago | Stable | |
pdf-renderer jobs | 0/ 0 | Schedule · 23:30 | Scaler | — 18 sec ago | Paused |
Add an SQS failure floor
Queue metrics failed twice this week. Hold six workers after three failed polls instead of relying on stale data.
Operation health
Rolling 60 minutes
Today’s schedules
All times in Asia/Jerusalem
Small features with clear ownership
Each addition keeps the controller focused instead of growing a generic trigger platform.
Ownership guard
Detect HPA or KEDA before writes and create a transfer plan instead of competing for replicas.
SQS failure policy
Choose hold-current or a safe replica floor when queue metrics become stale or unavailable.
Readiness SLO
Measure API acceptance, scheduling, and Ready pods separately so latency claims stay honest.