
NVIDIA used the SemiAnalysis AgentX workload yesterday to evaluate the DeepSeek-v4-PRO 1.6T model, showing that Vera Rubin achieves 30x the throughput per megawatt of Blackwell.
Throughput per megawatt (Tokens per Megawatt, or tokens / MW) is a core energy-efficiency metric for measuring how many AI tokens an AI data center can produce within a fixed power budget.

According to the blog post cited by IT Home, SemiAnalysis AgentX focuses on evaluating agentic coding and reasoning workloads through four metrics:
End-to-end normalized interactivity measures the complete time from request submission until the final token arrives, reflecting user-visible output efficiency, including the wait for the first token;
Standard interactivity measures only the generation phase between the first token and the final token.
End-to-end latency measures the total time from request submission until the final output token arrives.
Time to first token (TTFT) measures the time from request submission until the first output token arrives.
On the DeepSeek-v4-PRO 1.6T workload, the Grace Blackwell server using GB300 NVL72 achieves 15x the throughput per megawatt of the H200 NVL8 Hopper solution.

Blackwell's cost per million tokens is approximately one-tenth that of the previous generation, allowing operators to support more agents within the same power and infrastructure budget, or maintain the same agent capacity at a lower operating cost.

Vera Rubin extends this advantage further. On the DeepSeek V4 Pro model, Vera Rubin NVL72 delivers approximately 30x the throughput per megawatt of GB300 NVL72.

Vera Rubin NVL72's cost per million tokens is one-thirty-fifth that of GB300 NVL72, enabling agents to run continuously at scale and meet customers' various workload demands.


