Skip to content
Preptima

Cost & Capacity Management interview questions

GPU economics, spot and preemptible capacity, right-sizing inference, caching, and knowing the unit cost of a prediction.

3 questions in MLOps & ML Platform.

medium1

mediumConceptDesign

What does one prediction cost you, and how would you work it out?

Divide the all-in hourly cost of the serving capacity by the predictions it delivers in that hour at your real utilisation, then add the per-request cost of feature reads. Utilisation dominates the result, and most of what people call cost per prediction is fixed cost that one more request does not change.

5 minmid, senior, staff, lead

hard2

hardDesignScenario

How would you cut the cost of a GPU training fleet without slowing the team down?

Measure utilisation first, because most fleets pay for idle accelerators rather than for too few. Then right-size to whatever each job is bound by, move interruptible work onto spot capacity behind tested checkpointing, and pool the fleet behind one queue instead of reserving per team.

6 minsenior, staff, lead