Rank the ways you could lose model access, build portability you actually exercise, and buy committed capacity in proportion to the revenue at risk—before a vendor makes the decision for you.
Maintained in the open at github.com/Golden-Section-Tx/playbook · CC BY-SA 4.0
Once an AI feature is in your product, a third party you do not control sits inside your critical path. The question is not whether that is risky—it is which risk, and how much it is worth spending to cover.
The goal: Your product keeps working, in a known and degraded-but-acceptable state, when a model vendor raises prices, tightens limits, deprecates a model, or stops selling to you.
Founders reliably rank these risks in the wrong order. The usual instinct is to worry about a dramatic physical event—a regional grid failure, a data center going dark—and to under-weight the mundane commercial one. The order of likelihood is the reverse:
Two structural facts shape the response.
Your fallback does not have to match your primary. A smaller open-weight model running on modest hardware will not equal a frontier model, and does not need to. What it needs to do is keep the workflow functioning at reduced capability while you recover. Decide in advance what "degraded but acceptable" means for your product, and make sure your customer contracts permit it.
A failover that has never carried production traffic is not a failover. It is a configuration file and an assumption. The only way to know your second provider works is to keep sending it a small, continuous share of real requests.
Owning hardware is almost never the right first move at this stage. Buying accelerators means buying a depreciating asset with a three-to-six year economic life, plus the colocation, the spares and the person who owns it—to insure against a risk you have not yet experienced. It converts a margin question into a balance sheet question. Revisit it only when a vendor has actually restricted you, or when steady-state volume means the hardware would run at genuine utilization rather than sitting idle.
The ladder, in order of return on spend:
| Rung | What it is | Rough cost |
|---|---|---|
| 1 | Model portability: an abstraction layer, an evaluation suite, and continuous shadow traffic to a second provider | Engineering time, no capital |
| 2 | Committed or provisioned capacity bought from your primary provider, and a smaller commitment on the second | Low single-digit % of revenue |
| 3 | Reserved GPU capacity rented from a cloud provider on a term commitment | Materially more than rung 2 |
| 4 | Owned hardware in colocation | Capital, plus ongoing operations and a depreciating asset |
Most companies at this stage should be fully at rung 1, partially at rung 2, and nowhere near rungs 3 and 4.
Specific prices in this area move fast enough that quoting them dates the play. Two anchors that have held: committed capacity from a major provider is generally available at a low single-digit percentage of revenue for a company at this stage, and owned hardware does not beat reserved rental on a three-year horizon at small scale. Get three written quotes before budgeting anything.
Commercial deprioritization. Your vendor changes its terms, withdraws a capacity product, deprecates the model your prompts are tuned to, or simply serves larger contracts first when demand spikes. This is the common case, it arrives by email, and small accounts absorb it first.
Your fallback does not have to match your primary. A smaller open-weight model running on modest hardware will not equal a frontier model, and does not need to. What it needs to do is keep the workflow functioning at reduced capability while you recover.
Rank the ways you could lose model access, build portability you actually exercise, and buy committed capacity in proportion to the revenue at risk—before a vendor makes the decision for you.
Diligence the vendors underneath your vendor. For any hosting or infrastructure provider, ask for—and read: The utility service or interconnection agreement in the provider's own name . A reseller cannot produce one, because the power contract belongs to whoever actually owns the facility.
A failover that has never carried production traffic is not a failover. It is a configuration file and an assumption. The only way to know your second provider works is to keep sending it a small, continuous share of real requests.