DeepSeek's fast, economical model. Handles both chat and reasoning (thinking) modes at very low cost.
Performance
Time to first token
441ms
↑ 2%vs prior 24h
Total response time
1553ms
↑ 1%vs prior 24h
Throughput
32.6tok/s
↓ 1%vs prior 24h
Inter-token latency
15.4ms
↑ 9%vs prior 24h