Rent the GPU-second, not the GPU.
One authenticated request spins up a GPU, runs your model, and bills you for the seconds it actually burned — never for idle, never for the cold-start wait. Then it's gone. GPU inference metered and guarded by WAVE, behind the key you already have.
No cluster. No idle meter. No bill for the wait.
Every call is authenticated, quota-checked, and metered at the WAVE gateway before it ever reaches silicon — the same key you already use across WAVE. You're billed on the GPU's real execution time, in fractional GPU-hours at $0.50/GPU-hour — never on the cold-start wait. And it's fail-closed: no valid WAVE request, no GPU, no charge.
Federated — auth + metering federate through api.wave.online, so a GPU-hour shows up next to your transports, media, and data on one bill, behind one key.
Where it composes
GPU compute becomes one more metered WAVE rail. Inference lands on the same gateway, the same quota, and the same usage ledger as every other spoke. You stop operating GPUs and start calling them.