OpenAI's small and affordable model. Great for lightweight tasks.
Performance
Time to first token
—ms
—vs prior 24h
Total response time
Throughput
—tok/s
Inter-token latency