DeepSeek V4 Flash 0731 is available from many inference providers, and I needed to evaluate which ones could handle our production traffic.

I sent 30 streaming requests concurrently to each provider across four workloads: 120 per provider and 1,080 total requests. Decode speed was measured as the combined rate of the 30 decode requests.

This is a one-time snapshot. DeepInfra felt faster earlier in the day, but got slower as US traffic was starting to ramp up at the time when I was doing this test.

Results

Provider Decode tok/s TTFT p50 / p99 Total p99 Cache hit Errors
Scaleway 1,650 733 / 818 ms 18.62 s 80.2% 18/120
TensorX 1,192 873 / 1,630 ms 25.77 s 73.4% 1/120
Fireworks 2,333 1,072 / 2,592 ms 13.17 s 17.6% 0/120
Baseten 3,980 774 / 2,979 ms 7.72 s 78.7% 0/120
Azure 596 618 / 3,029 ms 51.55 s 57.7% 0/120
DeepInfra 685 1,466 / 4,663 ms 44.85 s 86.6% 0/120
DigitalOcean 1,280 985 / 7,412 ms 24.00 s 26.5% 0/120
Nebius 2,542 948 / 7,563 ms 12.08 s 0.0% 0/120
Lyceum 934 8,275 / 9,354 ms 32.89 s 96.0% 0/120

Azure wasn’t running the new DeepSeek V4 Flash 0731 checkpoint, so its results aren’t directly comparable with the others.

Scaleway’s 18 failures were HTTP 429s from a token-per-minute quota which we need to increase. Other than the rate limits, successful requests to them had the best TTFT distribution. TensorX had one transient HTTP 502. Otherwise, all the providers were solid.