GPU power-performance curve saturates, creating optimization opportunity for memory-bound inference
The power-performance relationship in GPU clusters is non-linear and saturates because inference workloads are memory-bound — decode stages consume power reading weights into HBM while SMs…