Benchmark Suite Results

MLPerf Endpoints

The modern way to measure AI system performance.
Built for buyers.

You are about to spend a large amount of money on AI infrastructure for your critical applications. You need to know what you are actually getting for your money. MLPerf Endpoints tells you, in a picture you can read in a minute.

The purpose of Endpoints is to help you understand, through empirical, industry standard, peer-reviewed measurement, how well a system serves a model under load. 


What does an Endpoints benchmark show me

Endpoints visualizes results of every benchmarked system. A plotted curve on that visualization represents one system – a combination of hardware, software, and a deployed AI model. Each point on the curve is a unique operating point that is measured across four critical dimensions that are all related: System Throughput or System TPS, interactivity, time to first token P95 (TTFT P95), and concurrency. Many buyers have dozens of applications which each demand separate operating points, and this curve lets you match your needs to what the system offers.

Each point on the curve is one real test, run at one level of concurrent load. As you pack more users onto the system, total throughput goes up and per-user speed comes down. That is the core tradeoff in serving AI, and the different curves show this clearly.

Concurrency is the number of simultaneous queries that the system can sustain in flight. This roughly corresponds to the number of users. For some applications, each query would correspond to one user, but more complex applications might require multiple queries for a user. In MLPerf Endpoints, the concurrency is fixed for each point on the curve but varies for different points to illustrate different load levels. As concurrency increases, system throughput tends to increase up to a plateau; the interactivity will tend to fall and TTFT P95 will tend to increase, especially in the plateau.

In our default view, the vertical axis is System Throughput or System TPS. This is the total number of tokens the system produces per second across all users. Higher means more capacity and better utilization of your system and greater efficiency. As concurrency increases, system throughput often increases until it reaches a plateau.

In our default view, the horizontal axis is interactivity, measured in tokens per second per user. This is how fast the AI system response streams back to each query. Higher means a snappier experience for a query to finish. As concurrency increases, the interactivity will tend to fall, especially around the plateau in system throughput.

Time to first token P95 (TTFT P95) is the time it takes for the first token (typically a word for an LLM) to come back to the user from a query, measured at the 95th percentile. Lower means that responses come back quicker and a better user experience. As concurrency increases, the TTFT P95 will tend to increase, especially around the plateau in system throughput.

In the default view, a higher curve or point means more total throughput at the same interactivity, or faster responses at the same throughput.

Other charts show system throughput, interactivity, or TTFT P95 on the vertical axis plotted against concurrency. For system throughput or interactivity vs. concurrency a higher curve or point is generally faster or more responsive. For TTFT P95 vs. concurrency, a lower curve or point generally means more responsive.


What do the numbers mean for me?

The default chart answers one question. What will this system do for my workload? The answer depends on what you are building. Here are three common cases.

You are building a customer-facing chatbot. People are waiting on the other end, live. Two numbers decide whether it feels good: TTFT P95 (how fast the first part of the answer returns to the user), and interactivity (how fast the answer streams, in tokens per second per user). You also watch concurrency. 

Define your latency requirements first (say, TTFT P95 under 400ms, at least 20 tokens per second after that), find the highest load that stays under your latency limit, then read the System TPS at the point where interactivity is 20. That is the traffic the system carries while every user still feels good. Pick the system whose curve stays highest while holding both limits.

You are building an internal tool and want to control cost. Internal users can trade speed for a smaller bill – perhaps you then prefer an open-weight model and you do not need the priciest hardware. The tradeoff runs in your favor: throughput peaks when the system serves many users at once, which lowers per-user speed, and your users tolerate that. Pick the lowest per-user speed they will accept, read the System TPS near the top of the curve, and divide by system price for throughput per dollar. Look for open-weight results (Llama 3.1 8B, GPT-OSS 120B) on hardware priced for the job, marked Available.

You are supporting a small team of demanding engineers. A few power users who want the best model and the fastest responses, budget second. Read the right side of the chart where per-user speed is highest, and favor systems that hold high interactivity while running a near-frontier model like DeepSeek R1. Newest hardware often shows up as Preview, so check the status and date.


How are they created? Who creates them? How is this different from MLPerf Inference?

Vendors run the benchmark on their own systems and submit the results. Every submission contains information about the system under test and the artifacts needed to reproduce it. You know exactly who produced each curve and on what.

Before a result is published, it goes through peer review for verification. The review committee includes anyone who has submitted in the last 6 months, which means direct competitors, customers, and selected experts from MLCommons and the broader community. The review committee examines every submission and can object or question the results. This is the same principle that governs scientific publishing. People who have every reason to find a flaw get the chance to look first and make sure the results are robust and verifiable.

How this differs from MLPerf Inference. MLPerf Inference defined a set of fixed scenarios meant to model different use cases. In the datacenter these were offline, server, and interactive. 

MLPerf Endpoints takes inference measurement from a few fixed operating points to a full characterization of performance. It measures a live endpoint the way a customer would actually call it, and sweeps the load to trace the entire performance curve. The result is the full speed-versus-capacity picture for the real serving stack: the hardware, the model, and the software that serves it. This essentially replaces the server and interactive scenarios. Endpoints builds on the MLPerf Inference rules and future versions will evolve the rules even more.

The models measured at this time include GPT-OSS 120B, Llama 3.1 8B, and DeepSeek R1.


Why can I trust these numbers?

Three things stand behind every published result.

They ran on the vendor’s real system. The submitter is named, and the benchmark tests a working endpoint. Our rules require that the results be reproducible, allowing you to verify performance claims.

They were peer reviewed. Competitors and independent reviewers check each submission before it is published. Objections are raised and resolved by a review committee to ensure reported results are valid.

MLPerf results are audited. MLPerf results are subject to audit by a third to ensure compliance with the rules, validity, and reproducibility.

MLPerf results that are verified have passed peer review. A provisional MLPerf result has been submitted to MLCommons and passed initial checks but is still under peer review and may change if issues are found.

Sometimes organizations will cite unverified MLPerf results, which have not been submitted for review at all and may not follow our rules and methodology and therefore may not be comparable to verified or provisional MLPerf results. For a purchase decision, look for verified.

Results also carry an availability label. Available means you can buy or rent the system today. Preview means it is coming within 180 days. Align the availability label to your procurement timeline.


How do I use this to procure new hardware?

Put it in your RFP. Ask each vendor for an MLPerf Endpoints result for the exact configuration they are proposing, in the standardized division, with an availability status that matches your purchase window. Ask for a verified result, or a commitment to submit for verification. Ask that they run MLPerf Endpoints on the system you buy or rent and submit it for verification to ensure that your deployed system meets your expectations and your vendor’s commitments.

Ask them to run your workload, not a generic one. The number that matters is the one at your operating point. Tell vendors the model you will run, the interactivity your users need, and the TTFT P95 you can accept. Ask them to run Endpoints in the hardware and software configuration you would actually buy, and to submit it. Or even better, ask them to run Endpoints on your system after it’s been deployed. If your workload isn’t here, you can ask for it, contact us at <SOME EMAIL PROXY>.

Buying from an OEM, a cloud provider, or a neocloud. The endpoint is the product. Ask for the curve for the specific instance or configuration you would purchase or rent, under the software they will actually serve you. Then place your requirement on the chart and read the capacity each option gives you at that point. Compare curves at the same point, never at each vendor’s most flattering point.

What if my vendor or solution has not been tested? Then you have no independent, peer-reviewed measurement of how it performs. Vendor benchmarks that have not gone through review are unverified by definition. You can still ask them to run Endpoints and submit. A vendor confident in their system has no reason to refuse.


Technical detail

For teams who want to go deeper.

Workloads. The current round measures widely used open models, including GPT-OSS 120B, Llama 3.1 8B, and DeepSeek R1. Each has its own reference implementation and its own accuracy and configuration rules.

How a result is built. A result is a pareto curve, not a single point. A submission includes at least 7 and at most 32 points. Each point runs for at least 600 seconds. Every point uses the same model, endpoint, and software stack, so the curve is internally consistent. Each point may be using a different configuration of the system, e.g., shifting resources from pre-fill to decode. Accuracy is validated for each distinct configuration.

Divisions and status. The standardized division requires model equivalence and full software disclosure with whitebox transparency. Every result is marked verified, provisional, or unverified, and labeled Available, Preview, or RDI.

Peer review and publication. Submitters run the benchmark, register their runs, and open a review request. Reviewers file objections, the submitter responds, and the result is published on a fixed cycle once review is complete. Objections to a published result can be raised later and routed through dispute resolution.

Governance. The rules are set by an open working group that operates by consensus and brings the whole community to the table. The methodology, the rules, and the review process are public.

Normalization. We are actively developing our normalization rules with our community, which we will incorporate into the visualization as soon as they are finalized. If you have a specific request for how you want results to be normalized that would support you in your business decisions, we want to know about it. Please contact us at <SOME EMAIL PROXY>. 

Links.


Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.