Capacity Autopilot
Every megawatt,
doing useful work.
Cirra is a capacity controller for NVIDIA GPU fleets on Kubernetes and Slurm. It increases useful compute per fleet and constrained megawatt when demand exists — and safely powers down the electricity and cooling cost of capacity that isn't needed.
Read-only baseline · operator-approved actions · every outcome measured
Powered nodes
112/112
Fleet utilization
52%
Power released
0 kW
The problem
Your datacenter is optimized by systems that never talk to each other.
Kubernetes and Slurm manage workloads. GPU managers expose accelerator state. Facility systems run power and cooling. No single layer understands workload intent, physical constraints, future demand, and economic value well enough to coordinate the full chain —so GPUs sit powered-on and underfed while queues grow and megawatts burn.
The dependency chain Cirra coordinates
Dual optimization engine
One engine. Two regimes. Your lead outcome.
Cirra models both value pools on every cluster, then leads each deployment with the single outcome its buyer owns.
When demand is high
Process more AI work with the fleet you already own.
Rightsizing, admission, placement, and consolidation policies that convert fragmented and oversized allocations into schedulable capacity — verified against completed work, queue throughput, and SLOs.
- ▸Completed work & queue time
- ▸ Capex deferred, overflow avoided
- ▸ Compute per constrained megawatt
When demand falls
Match powered-on capacity to real demand.
Predict the safe idle window, consolidate flexible workloads, hold your warm reserve, and release the remaining nodes to standby — with staged wake scheduled before demand returns.
- ▸Powered-node-hours & measured IT energy
- ▸Cooling load & peak-demand headroom
- ▸ Wake readiness, never sacrificed
0%
GPU allocations attributed to a workload
0s
cluster state freshness
0+
GPUs supported per cluster
10–0
measured canary runs per pilot
Evidence, not anecdotes
Not a recommendation engine. A control system that proves itself.
Cirra starts read-only, builds a live model of your fleet, and replays every policy against history before anything is touched. Then one reversible action class is executed repeatedly inside an operator-approved canary scope — so your pilot ends with a distribution of predicted-versus-measured outcomes, not a single success story.
Example pilot outcome
A frozen admission-and-placement policy increased completed batch work by 6.1% in replay while preserving inference SLOs.
The approved admission action executed 14 times over 10 days. No deadlines or SLOs violated. No rollbacks. 143 GPU-hours and 27 powered-node-hours released.
→ production policy enabled for the model-evaluation queue
Built for the whole capacity team
One auditable connection from application behavior to dollars.
AI Infrastructure
Know whether more work can be served before buying more hardware or power.
Platform Engineering
Receive safe, explainable placement and consolidation plans — never surprises.
ML Systems
See whether the workload or the infrastructure is constraining GPU output.
FinOps & Capacity
Translate technical inefficiency into auditable dollars, by evidence tier.
Prove it on your cluster.
A 30–45 day shadow-to-control pilot. Customer-hosted, read-only to start, and judged on one buyer-owned outcome you choose before installation.