Roybench telemetry

Model speed history

Aggregate tokens_per_second is the arithmetic mean of output_tokens divided by the arithmetic mean of duration_seconds; it is not the arithmetic mean of per-attempt output_tokens / duration_seconds ratios. Effective output throughput includes proxy, network, provider queue, reasoning, and generation time; it is not raw decoder speed. Three Codex subscription GPT routes are post-response bounded at the declared 4,096 billed-output-token envelope. Eleven Cursor routes use provider defaults because the local agent exposes no documented output cap. Fifteen Devin routes send a 512-token request hint; throughput uses provider-reported output, and Devin documents no total-billed-output cap. The other eleven routes enforce a 512-token request output cap.

Leaderboard Provider uptime
Loading model speed observations…
Fastest latest
Waiting for data
Available now
Latest completed observation
Median latest
Effective output tokens per second
ModelsLoading selection… Filter by provider or model

Effective output throughput

UTC · pointer/touch · Left/Right and Home/End keys

A time series chart. Left and Right arrow keys move the shared active UTC bucket. Home moves to the first bucket and End moves to the latest.

Choose a time on the chart
    Current status and latest successful measurement for every configured model
    ModelRouteTPSDurationOutputSourceCurrent statusMeasurement observed at