News Archives - MLCommons https://mlcommons.org/category/news/ Better AI for Everyone Wed, 29 Jul 2026 15:01:16 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.2 https://mlcommons.org/wp-content/uploads/2024/10/cropped-favicon-32x32.png News Archives - MLCommons https://mlcommons.org/category/news/ 32 32 MLPerf Endpoints v0.7: A Foundation Release https://mlcommons.org/2026/07/mlperf-endpoints-v0-7-release/ Tue, 28 Jul 2026 14:50:00 +0000 https://mlcommons.org/?p=4178 From demonstration at GTC to foundation release: new results, automated submission pipelines, and a roadmap to v1.0 with rolling submissions later this year.

The post MLPerf Endpoints v0.7: A Foundation Release appeared first on MLCommons.

]]>
MLPerf has been the standard-bearer for measuring AI system performance since its launch in 2018. During that time, MLPerf has tracked over a 100X improvement in inference performance per watt for large language models and over a 50X improvement in training speed [1] [2]. Over the past eight years, the AI industry has matured, with AI services now used daily by enterprises and consumers worldwide. Today, we are announcing an evolution of MLPerf that will make it even more useful for this maturing industry: we are launching the first version of MLPerf Endpoints

MLPerf was originally designed to help a relatively small number of cloud providers and downstream system builders inform their AI hardware purchasing decisions. Today, procuring inference compute is a salient business decision for companies of all sizes – and it means evaluating options across neoclouds, cloud providers, and managed services simultaneously. These buyers need reliable, comparable, and independent performance benchmarks to inform their purchasing decisions, and these benchmarks must be dynamic enough to keep pace with an industry that launches new models weekly. 

MLPerf Endpoints is designed to meet this more diverse and complex demand for benchmarks to inform AI system procurement decisions. Four key principles are guiding our development of enterprise-buyer-centric benchmarks:  

  1. Current: Results must keep pace with the market. Buyers can’t wait months for a benchmark round to include new hardware or models. 
  2. Comprehensive: Buyers need benchmarks that cover the many competing inference providers, systems, and workloads available for purchase. 
  3. Comparable: Buyers need to see apples-to-apples results across vendors, normalized for cost or power, to inform procurement decisions.
  4. Commentary: To support buyers in these decisions, Endpoints provides additional context for benchmark results through visualizations, data filtering, and analysis. 

MLPerf Endpoints v0.7 is a foundation release with initial results from Coreweave, Google, Intel, KRAI, and Nvidia. We want to congratulate these members on their excellent results, spanning several orders of magnitude in performance across 3 benchmarks. This is the infrastructure on which we are building a more dynamic, comprehensive, and comparable inference benchmark suite for the data center. Endpoints currently supports automated submission pipelines, continuous review tooling, and dynamic visualization of results that you can see at mlcommons.endpoints.com. We are also evolving our benchmarking rules towards more buyer-centric benchmarks.

Later this year, we will deliver MLPerf Endpoints v1.0 with more buyer-centric rules, normalization, and an expanded set of benchmarks including agentic workloads and then open the rolling submission process to our broader membership. The rolling submission process ensures Endpoints will provide current and up-to-date results that move at the pace of the market. 

We’re grateful to our 30+ supporters who have helped guide the development of MLPerf Endpoints, including AMD, Argonne National Laboratory, Broadcom, Core 42, Dell, HPE, Lambda, Oracle,  and Red Hat. Their partnership has been invaluable in stress-testing our rules, processes, and improving our infrastructure. With this release, we are just getting started. 

If you’re a system or service provider, now is the time to get involved, help shape the rules, and plan your submission at a time of your choosing. Complete this form to let us know you’re interested in submitting, and what hardware you want to submit on. The community-developed rules are actively being refined for the 1.0 release later this year, and early participants in this effort will shape Endpoints’ direction. If you’re an enterprise buyer, Endpoints is being built for you, and we want to know how we can best support you and your organization. Let us know by emailing alejandro@mlcommons.org.

We deeply value the trust our members place in us as the benchmark of record for inference. We are excited to build this future of MLPerf with all of you. 

[1] A. Tschand, A. T. R. Rajan, S. Idgunji, et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from µWatts to MWatts for Sustainable AI,” in 2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025. arXiv:2410.12032.

[2] D. Kanter, M. Ahmad, H. Kassa, and S. Rishab, “MLPerf Training v4.1 Results — Press Briefing,” MLCommons, Nov. 13, 2024. [Online]. Available: https://docs.google.com/presentation/d/1KSIJBvIV9OcswF1mVbhGRN0nUWbSgbYSoxu6dHCwajM/

The post MLPerf Endpoints v0.7: A Foundation Release appeared first on MLCommons.

]]>
MedPerf Meets Google Cloud Confidential Computing: Secure AI Benchmarking for Brain Tumor Research https://mlcommons.org/2026/07/medperf-google-cloud-confidential/ Wed, 22 Jul 2026 16:24:26 +0000 https://mlcommons.org/?p=4089 At Google Cloud Next 2026 in Las Vegas, MLCommons and Google Cloud demonstrated a powerful new capability for trustworthy medical AI - one that protects patient data, model IP, and benchmark integrity all at once.

The post MedPerf Meets Google Cloud Confidential Computing: Secure AI Benchmarking for Brain Tumor Research appeared first on MLCommons.

]]>
During Google Cloud Next 2026 in Las Vegas, the MLCommons Medical AI working group and Google Cloud announced the enablement of MedPerf, MLCommons’ federated benchmarking orchestrator, on Google Cloud’s confidential compute capabilities. Both teams demonstrated this integration on a compelling real-world clinical use case: brain tumor segmentation.

A brain tumor (glioblastoma) is a rare disease with devastating outcomes for life expectancy. AI has the potential to improve the diagnosis and prognosis of brain tumor patients through a multitude of automated processes – including boundary definition/segmentation, tumor classification, and tumor quantification. Thanks to the Federated Tumor Segmentation (FeTS) consortium, coordinated by Dr. Spyridon Bakas at Indiana University School of Medicine and the Response Assessment in Neuro-Oncology (RANO) cooperative group, the MedPerf team demonstrated evaluation of a clinically valuable AI model – designed by Dr. Evan Calabrese’s team at Duke University – on real-world brain tumor MRI data.

With MedPerf on Google Cloud’s Confidential Computing, we aim to advance medical AI research worldwide while securely evaluating models on medical data. Using Google Cloud’s Confidential Space, a Trusted Execution Environment (TEE), medical data, model weights, and benchmarks are fully protected throughout the benchmarking process.


The Problem: Life-Threatening Disease, Locked-Down Data

Protecting both AI models and patient data in the medical space is paramount, driven by privacy, ethical, sovereign, and regulatory considerations. At the same time, measuring the performance of AI models on real-world patient data is critical for enabling trust and adoption of AI among clinicians, patients, regulators, and payors.

The MLCommons community built MedPerf to address exactly this gap. MedPerf uses a federated approach – introduced to healthcare by Dr. Bakas’ team – in which AI models move to the medical data owners (such as healthcare organizations, hospitals, and data brokers), rather than requiring data to move to the benchmark operator. This allows data owners to keep their data private while sharing only summarized results with the benchmark operator.

However, while this federated approach protects patient data, it does not fully address concerns around model theft, intellectual property, and benchmark integrity. Without additional guarantees, there is no assurance that model weights won’t leak or that benchmark results won’t be tampered with at the time of execution – raising serious concerns about trust and participation in the benchmark itself.


The Solution: MedPerf on Google Cloud Confidential Space

MedPerf’s integration with Google Cloud’s Confidential Space solves exactly this problem: protecting both model weights and benchmark results while simultaneously safeguarding medical data.

This integration brings four key benefits to MedPerf users:

  1. A global test dataset for healthcare AI – enabling evaluation at unprecedented scale
  2. Patient data is protected – data never leaves the data owner’s secure environment
  3. Model IP is protected – model weights remain encrypted and inaccessible to unauthorized parties
  4. No data sharing required – results are summarized and encrypted before transmission

The Demo: A Step-by-Step Secure Benchmark

A benchmark authority first publishes a benchmark on the MedPerf server, describing the benchmark’s aim, data preparation guidelines, and evaluation criteria. Data owners and AI model owners then register metadata for their private assets with the MedPerf server and link their registrations to the benchmark.

Private assets are encrypted using Google Cloud’s Cloud Key Management Service (Cloud KMS) and uploaded to a Google Cloud Storage bucket. Access to Cloud KMS and the bucket is governed by a Google Cloud workload identity pool, which allows access only if the request originates from a confidential virtual machine with specific measurements – such as the benchmark container hash and input hash.

When it’s time to run, data owners execute the benchmark container inside a Confidential Space virtual machine. The container receives cryptographic evidence of the trusted execution environment’s configuration, which is exchanged for access tokens from the model owner’s and data owner’s workload identity pools. These tokens retrieve and decrypt the model and dataset, inference is executed, results are encrypted with the data owner’s public key, and the encrypted results are returned to the data owner’s bucket.

The video below demonstrates a brain tumor segmentation AI model running on a clinical dataset in Google Cloud:

▶ Watch the Demo


Voices from the Community

“My experience testing federated learning on Google Cloud has shown that the future of medical AI lies in secure, scalable, and collaborative cloud environments. Moving beyond the controlled lab setting to test these workflows in a production-ready infrastructure provided a unique opportunity to evaluate the performance and security of federated learning in real-world clinical applications. This collaboration between the FeTS community, MLCommons, and Google Cloud demonstrates how scalable cloud solutions can accelerate the development of high-precision diagnostic tools in neuroradiology.” – Dr. Yury Velichko, Northwestern University

“Federated Learning offers a paradigm shift in healthcare multi-site collaborations that expedites access to unprecedented amounts of knowledge from diverse patient populations, towards robust generalizable AI models with real clinical impact. Security and privacy are prime considerations to ensure that FL-related assets are always protected. Training and evaluating AI models on confidential compute infrastructure, using MLCommons’ MedPerf on Google Cloud, describes a complete exemplary solution for large-scale real-world clinically-relevant AI studies in healthcare, providing the ability to track how well an AI model continues to perform after it has received regulatory approval, for example, from the FDA.” – Dr. Spyridon Bakas, Indiana University School of Medicine

“One of the greatest barriers to trustworthy medical AI has been the tension between the need for real-world data and the need to protect it. The integration of MedPerf with Google Cloud’s Confidential Space is a tangible solution to that tension – one that the MLCommons community has worked hard to bring to life. We are proud to demonstrate what is possible when open collaboration meets production-grade security infrastructure.” – Alexandros Karargyris, MedPerf Lead, MLCommons


Join the MedPerf Community

MedPerf is an open-source effort developed by the medical AI community to address real community needs. Thanks to contributions from the MLCommons community, MedPerf has already supported impactful clinical AI studies, including the FeTS challenge and FeTS 2.0 for brain tumors.

We know we still have much more to do – and we need your help. We are looking for:

  • Research investigators with the ambition to lead global clinical AI studies
  • Computational experts who want to build disruptive AI on real-world medical data
  • Infrastructure engineers to improve real-world benchmarking technologies
  • Privacy and governance experts to provide input on medical data access controls and policy
  • Regulatory experts to provide feedback on bringing regulatory alignment to real-world benchmarking

Acknowledgments

Contributors to the MedPerf confidential compute integration include Google Cloud team lead Nelly Porter, Dr. Spyridon Bakas from Indiana University School of Medicine, Dr. Yury Velichko from Northwestern University, and Dr. Amber Simpson from the University of Alberta (formerly at Queen’s University).

The post MedPerf Meets Google Cloud Confidential Computing: Secure AI Benchmarking for Brain Tumor Research appeared first on MLCommons.

]]>
Agentic Inference for MLPerf Inference https://mlcommons.org/2026/07/agentic-inference-for-mlperf-inference/ Wed, 08 Jul 2026 14:40:00 +0000 https://mlcommons.org/?p=4143 A new multi-turn benchmark for measuring LLM serving systems under growing context, and closed-loop agent workflows.

The post Agentic Inference for MLPerf Inference appeared first on MLCommons.

]]>
Introduction

The MLPerf Inference benchmark suite must evolve alongside AI deployment patterns. Early inference benchmarks focused on image classification, object detection, speech recognition, recommendation, and single-turn language generation. Those workloads remain important, but they no longer cover one of the fastest-growing ways large language models are used in production: multi-turn agentic inference.

For example, a coding assistant is far more complex than a single query. The agent will read an issue, inspect files, run commands, observe failures, edit code, and iterate potentially many times. Similarly, a workflow agent gathers customer information, calls tools, interprets results, asks follow-up questions, and continues until the task is resolved. In both cases, the workload is a trajectory: a sequence of dependent turns where each request includes the conversation history up to that point.

This changes the serving problem in four concrete ways:

  • Context grows across the trajectory, so prefill and KV-cache pressure increase over time.
  • KV-cache reuse is now a critical serving optimization for performance and efficiency.
  • Output lengths vary widely, from compact tool calls to long reasoning traces.
  • Turn dependencies make throughput a measure of closed-loop progress, not independent request rate.

Figure 1. Illustration of an agentic inference scenario involving multiple turns.

The Agentic Inference benchmark adds this class of workload to the MLPerf Endpoints framework. It keeps the general MLPerf measurement principles intact while defining the workload-specific pieces: model choice, dataset composition, multi-turn load generation, output validation, and constraints for optimizations such as prefix caching and speculative decoding.

Terminology

  • Turn: one client-issued user or tool request and the model response generated for that request.
  • Trajectory: the ordered sequence of turns for one task or simulated user; later turns include the accumulated conversation history.

Model selection

The benchmark needs long-context, thinking-capable LLMs that stress the serving behavior of agentic applications: growing context, KV-cache reuse, variable output lengths, and strict turn dependencies. For this benchmark, we chose Kimi K2.6 and Qwen3.6-35B-A3B. Kimi K2.6 brings state-of-the-art coding capability and a large model profile representative of leading agentic systems; Qwen3.6-35B-A3B is more compact and brings a new Gated DeltaNet (GDN) architecture and exceptional coding performance for its size. This gives the benchmark coverage across distinct serving behaviors, architectural choices, and speculative-decoding paths. Both models are evaluated with the same methodology and dataset. To avoid ambiguity, each model produces its own result; the two models are not run together or combined into a single score.

ModelKimi K2.6Qwen3.6-35B-A3B
ArchitectureMoE + MLAMoE + Gated DeltaNet/Attention
Params1T / 32B active35B / 3B active
Context262,144 tokens262,144 tokens
SettingsThinking; temp=1.0; top_p=0.95; preserve_thinkingtemp=1.0; top_p=0.95; top_k=20; presence=1.5; repetition=1.0; preserve_thinking
Spec decodingnvidia/Kimi-K2.6-Eagle3 headNative MTP within the model

Table 1. Model metadata and benchmark settings

Dataset and task selection

The benchmark dataset combines two agentic domains that exercise different infrastructure bottlenecks. Together, they contain 613 multi-turn trajectories: 113 agentic coding trajectories and 500 agentic workflow trajectories. Across the reference dataset, these trajectories contain 30,335 client-issued turns and 30,328 assistant turns. By using real collected traces for benchmarking, the workload simulates real-world behavior in the serving stack, including how speculative decoding and expert-rank balancing behave under realistic multi-turn traffic.

DomainScalePrimary stress
Agentic Coding113 traj.; 15,981 client turnsDeep trajectories; growing context; KV-cache capacity
Agentic Workflow500 traj.; 4,316 client turnsLarge shared prompt; prefix overlap; prefix-cache efficiency

Table 2. Dataset composition and mean token statistics

The agentic coding traces come from the DeepSWE dataset by DataCurve (datacurve.ai), which contains software engineering tasks built around repository-level bugs and feature requests. A typical trace begins with a user request describing an issue, followed by an assistant that investigates the repository through a series of bash commands. The agent searches files, reads code, runs tests, observes errors, edits the implementation, and iterates. These traces are deep: the median trajectory contains dozens of turns, and the conversation history grows steadily as command outputs, file contents, and test logs accumulate.

The Workato agentic workflow traces come from enterprise customer-support and orchestration scenarios. They represent interactions where a simulated customer asks for help and the agent uses tools to retrieve orders, track shipments, check policies, escalate cases, or resolve account questions. These trajectories are usually shallower than the coding traces, but they include a much larger shared system prompt with many tool definitions and business rules. The agentic workflow traces were provided by Workato, the agentic control and execution plane for the enterprise. The data are synthetic traces modeled on Workato’s production experience orchestrating customer support agents for enterprise customers.

The two domains are intentionally different:

  • Coding traces grow aggressively over many turns.
  • Workflow traces begin with a large common prefix and grow more slowly.
  • Coding stresses KV-cache capacity and long-context scheduling.
  • Workflow stresses shared-prefix reuse and routing locality.
  • The combined workload prevents systems from optimizing for only one agentic traffic shape and stresses context-aware routing.

Client Design

We introduced multi-turn support to MLPerf Endpoints so the benchmark can drive full end-to-end agentic workloads against standard serving stacks. Submitters simply need to spin up an OpenAI-compatible endpoint using vLLM, SGLang, TensorRT-LLM, or another serving framework, then point the client at that endpoint for full end-to-end benchmarking.

The client handles the multi-turn behavior:

  • Closed-loop replay: one active conversation issues one turn at a time and waits for the complete model response before the next turn.
  • Target concurrency: the load generator controls the number of active users or conversations without breaking turn dependencies.
  • Inter-turn delay: dataset-provided waits before tool or user turns preserve realistic pacing without counting delay as model serving latency.
  • Conversation-aware routing: a stable X-Session-ID header is sent for each trajectory so routers can preserve KV-cache locality.
  • Cache salting: deterministic salt markers around the system prompt allow valid within-trajectory reuse while blocking invalid cross-trajectory reuse to ensure the benchmark is representative of production workloads.
  • Deterministic prompt reconstruction: future prompts are built from the pre-recorded dataset rather than from live model output, which keeps performance runs reproducible while still measuring generated outputs.
  • Generated token cache clearing: In order to enable fair benchmarking of all platforms, the generated tokens are cleared out of the KV cache by introducing a whitespace character – ensuring that the performance is independent of the system that was used for generating the traces.

Performance metrics

The primary performance chart is a Pareto curve where each point is a fixed-concurrency benchmark result. Submitters need to submit a full Pareto curve as defined in the MLPerf Endpoints rules.

  • Y-axis: output tokens per second per system, measuring aggregate serving throughput.
  • X-axis: output tokens per second per user, computed as output tokens per trajectory / (E2E time – total delays).

The two axes represent complementary views of the serving system. The x-axis, output tokens per second per user, captures how quickly an individual agent progresses through its task; moving right means faster task completion. The y-axis, output tokens per second per system, captures the total work completed across all active agents; moving up means greater aggregate throughput and higher system utilization. As concurrency increases, the system may complete more total work while each individual agent progresses more slowly. The Pareto curve makes this tradeoff visible.

Figure 2. Chart illustrating agentic inference performance

Accuracy metrics

Accuracy is enforced as a three-level hierarchy so the Pareto curve measures usable agentic progress rather than responses that are shorter, lower quality, or easier to serve.

  • OSL mean value: prevents systems from gaining throughput by truncating responses or shifting the answer-length distribution. The check requires generated outputs to preserve the expected mean output sequence length for the workload.
  • Inline accuracy: catches behavior drift during the exact performance configuration being reported. It operates on the outputs produced by the performance run itself and compares them against ground-truth values preserved in the dataset, including coding tool-call action checks and workflow intent-code checks, so no separate inference pass is required.
  • Standalone accuracy: verifies task-level capability beyond turn-level similarity. The standalone evaluation will use 200 tasks from the SWE-bench Verified dataset to check whether the submitted model configuration can solve representative tasks end to end.

All three levels must be run for every point on the Pareto curve, and each score must meet a pre-defined threshold. Thresholds are defined separately for each level and for both Kimi K2.6 and Qwen3.6-35B-A3B.

Accuracy levelKimi K2.6Qwen3.6-35B-A3B
Inline accuracy63.08%55.86%
OSL per-turn mean range[404,494][355,434]
Standalone accuracy76.5%67.0%

Table 3. Tentative accuracy values, subject to change.

Reference implementation

Submitters only need to spin up an OpenAI-compatible server using vLLM, SGLang, TensorRT-LLM, or another serving framework, then point the MLPerf Endpoints client at that endpoint to run the full benchmark. The MLPerf Endpoints Agentic Inference README includes the example configs, command lines, and setup steps.

Conclusion

Agentic inference is one of the fastest growing applications of AI today and it’s far more complicated than existing single-turn text generation. Agentic workflows have serving patterns with unique constraints: growing context, strict turn dependencies, tool-mediated delays, shared prefixes, long-tail trajectories, and output quality checks that must run on the same configuration used for performance.

The Agentic Inference benchmark brings those constraints into MLPerf in a reproducible form. By combining deep coding trajectories with shared-prefix workflow trajectories, it measures whether serving systems can make progress for real multi-turn users while still using hardware efficiently.

As agentic applications become a larger share of production LLM traffic, the industry needs benchmarks that measure the systems work that actually matters: preserving cache locality, scheduling long contexts, maintaining output quality, and balancing aggregate throughput with per-user progress. This benchmark is a step toward that measurement standard. If your team is building serving infrastructure for multi-turn agents, join MLCommons to access the reference implementation and submit results

References

The post Agentic Inference for MLPerf Inference appeared first on MLCommons.

]]>
The Benchmark Behind the Next Wave of Ultra-Low-Power AI https://mlcommons.org/2026/07/mlperf-tiny-v1-4-results/ Tue, 07 Jul 2026 14:39:00 +0000 https://mlcommons.org/?p=4041 MLPerf Tiny: Benchmarking AI at the Edge 

The post The Benchmark Behind the Next Wave of Ultra-Low-Power AI appeared first on MLCommons.

]]>
Machine learning (ML) is no longer confined to data centers and is transforming the world around us, adding more intelligence to our day-to-day lives. It now runs on doorbell cameras, hearing aids, factory sensors, and battery-powered wearables. These devices operate on a few milliwatts and must respond in real time. 

As that footprint expands, a hard question follows. How do you fairly measure their performance and efficiency when no two of these devices look alike? A single release can include anything from a 60 MHz microcontroller to a vector-enabled RISC-V core to a dedicated neural processing unit.

MLPerf® Tiny is the consortium-built answer. Developed by the MLCommons® Tiny Working Group with EEMBC, it provides an architecture-neutral way to compare ultra-low-power systems using the same workloads, models, and measurement methodology. 

This post covers what MLPerf Tiny measures, why energy is treated as a first-class metric, and what the v1.4 submission round reveals about the direction of edge AI.

The rise of TinyML and why it matters

TinyML refers to ML models that are small enough, typically under 2M weights, to run on microcontrollers and other constrained devices that draw sub-milliwatt to low-milliwatt power. That footprint matters because it opens up places that cloud or smartphone inference cannot go, unlocking new applications and capabilities. 

In 2026, that includes:

The appeal is practical. Inference on the device keeps data where it’s created instead of sending it to the cloud. That means lower latency, lower cost, better privacy, and, for many battery-powered systems, much longer operating life.

The software ecosystem has matured just as quickly. TensorFlow Lite for Microcontrollers, ExecuTorch, Edge Impulse, STEdgeAI-Core, NXP eIQ, and vendor-specific compilers such as AndesAIRE and RUHMI have made it much easier to deploy models on embedded hardware. But they’ve also exposed the challenge of comparing results across substantially different platforms in a robust and trusted manner. Vendors benchmark different models across different datasets and under varying conditions, so performance numbers rarely tell the whole story.

The benchmarking gap no one had filled

Building a fair benchmark for ultra-low-power ML is hard in itself. The 2021 paper introducing MLPerf Tiny identified four challenges:

  • Measuring energy fairly. Power consumption varies from one device to another. On top of that, vendors don’t always measure the same things. Some include peripherals, firmware startup, or I/O activity, while others measure only the inference itself. These differences can change the reported numbers.
  • Fitting within tight memory budgets. TinyML devices typically work in kilobytes rather than gigabytes, roughly 6 orders of magnitude tighter than smartphone ML. Reference models, harness overhead, and multiple quantization paths all have to fit in that resource-constrained envelope, which lacks many of the conveniences of cloud systems.
  • Software & Hardware heterogeneity. Devices range from general-purpose microcontrollers to neural processors, event-based architectures, and in-memory compute. Also, each vendor typically ships its own tightly coupled toolchain. That makes a portable benchmark non-trivial since that may require over-constraining the stack, which can erase the very performance the benchmark is meant to measure. 

Figure 1: Summary of the Tiny Machine Learning Stack. There is diversity at every level, which makes standardization for benchmarking challenging. 

For years, embedded developers had benchmarks, just not ones that answered the questions they actually cared about.

CoreMark became the standard for measuring Microcontroller Unit (MCU) performance, but it was never designed for ML workloads. MLMark moved closer by benchmarking ML inference, yet its reference models, including ResNet-50, MobileNet, and SSD-MobileNet, target edge AI processors rather than resource-constrained microcontrollers. It also stops short of measuring energy, which is often the limiting factor in embedded deployments.

That leaves a gap. Speed and energy efficiency aren’t the same thing. A model can run fast and still drain the battery. So in TinyML, both have to be measured. 

MLPerf Tiny was created to fill that gap. More than 50 organizations from industry and academia, including Harvard, Google, STMicroelectronics, Qualcomm, Syntiant, Renesas, Infineon, CERN, and Silicon Labs, spent roughly 18 months building the benchmark suite. It is based on compact reference models, representative workloads, and a standardized methodology that treats energy as a first-class metric. The first public results were released as version 0.5 in June 2021.

What MLPerf Tiny actually measures

MLPerf Tiny is designed to answer a simple question. If everyone runs the same ML task, which hardware executes it most efficiently?

For that, it evaluates three metrics, including accuracy, latency, and energy per inference. Each task carries a fixed quality target, so submissions are compared on speed and energy at equivalent accuracy rather than trading one against the other.

Energy is the differentiator. In TinyML, a device that runs an inference in 10 ms but at twice the joules of a competitor is often the wrong choice, because battery life is the binding constraint. MLPerf Tiny uses EEMBC’s EnergyRunner harness to perform calibrated power measurement on the device under test, making energy directly comparable across silicon.

MLPerf Tiny is built to be modular. Improvements can come from anywhere in the stack, whether that’s the chip, the compiler, the runtime, or the model itself. A fixed reference keeps every submission measured against the same bar. 

To make that comparison fair, each task includes a predefined model, dataset, and a minimum accuracy target. Vendors can’t swap in an easier model or lower the accuracy to improve their numbers. Instead, they focus on optimizing how the model runs on their hardware. Once the implementation is ready, MLPerf measures three things, including whether it meets the required accuracy, how long each inference takes, and how much energy each inference consumes.

Visual Wake Words shows how this works in practice. Every submission must determine whether a 96×96 image contains a person, achieving at least 80% accuracy on the official test set. Because every vendor solves the same problem under the same rules, the results are directly comparable. For example, ASYGN’s ColibriNPU completed each inference using just 22.2 microjoules. It’s low enough that a CR2032 coin-cell battery could perform one inference every second for more than three years.

That’s the key idea behind MLPerf Tiny. The workload, evaluation, and accuracy targets are all fixed. What changes is how efficiently each hardware and software stack delivers the result.

The five benchmarks and the real-world problems they represent

MLPerf Tiny started with four tasks in v0.5 (2021) and expanded to five in v1.3 (2025). Each task corresponds to a recurring TinyML deployment pattern:

  • Keyword Spotting (KWS): A small DS-CNN trained on the Google Speech Commands dataset, representative of wake-word and voice-command detection in smart speakers, earbuds, and hearables.
  • Visual Wake Words (VWW): A MobileNetV1 binary classifier on 96×96 images, mirroring person-presence detection in low-power vision sensors for doorbells, security cameras, and occupancy sensing.
  • Image Classification (IC): A ResNet-style network on CIFAR-10, representative of general low-resolution vision workloads on edge devices.
  • Anomaly Detection (AD): An autoencoder trained on industrial machine sounds (ToyADMOS), representative of predictive maintenance and condition monitoring.
  • Streaming Wake Word: Added in v1.3, a 1D depthwise-separable CNN that detects a target word in a continuous audio stream. Unlike other benchmarks, this one also measures the device while it’s idle and listening, which is how wake-word systems actually run most of the time. These are the conditions under which most production wake-word systems actually run.

No benchmark can cover every embedded AI workload. These five capture the patterns that appear repeatedly in commercial TinyML products, making them a practical baseline for comparison. 

A benchmark built for everyone in the stack

MLPerf Tiny has two divisions because not everyone wants to measure the same thing.

  1. The Closed Division fixes the reference model, dataset, and processing pipeline. That leaves hardware and low-level software optimizations as the only variables, making results directly comparable across platforms.
  2. The Open Division is more flexible. Participants can use their own models, training methods, and optimizations, as long as they solve the same task and meet the required accuracy target. That makes it a place to evaluate new architectures, compiler techniques, and research ideas without losing the ability to compare against established benchmarks.

Figure 2: MLPerf Tiny’s modular design allows direct comparisons and shows improvement over the reference. Each reference implementation is swappable: green components can change in either division, orange only in the open division. 

MLPerf Tiny serves three different audiences:

  • Hardware vendors use it to prove their silicon
  • Software teams to evaluate their compilers and runtimes
  • Researchers to validate novel architectures, and potential customers to evaluate candidate hardware and software platforms against a neutral standard.

The v1.4 round includes examples of all three. 

Audiencev1.4 exampleWhat they showed
Hardware vendorsAndes TechnologyBenchmarked five RISC-V configurations, ranging from the compact D23 core to the vector-enabled AX46MPV and an AX27 paired with the AnDLA I370 accelerator. 
Software teamsDeepGate (first-time submitter)Demonstrated its compiler across Arm Cortex-M microcontrollers and neural accelerators. 
ResearchersUniversity of Leeds (Open Division)Evaluated a custom batch-normalization kernel running alongside an Xilinx DPU on a Versal Adaptive SoC. 

What the results are already telling us

The v1.4 round closed with submissions from nine organizations, including Andes, ASYGN, DeepGate, Kai Jiang, Qualcomm, Renesas, STMicroelectronics, Syntiant, and the University of Leeds. Together, they submitted 25 system configurations across the Closed and Open Divisions. Five of the participating organizations were first-time submitters, more than doubling the participant count from v1.3.

These submissions highlight the real value of MLPerf Tiny. Every result is measured against the same task, model, accuracy target, and energy methodology, making improvements directly comparable instead of relying on vendor claims. 

A few signals stand out:

  • Dedicated accelerators are moving into MCU-class parts. STMicroelectronics reports that enabling the hardware signal processor on the STM32U3 reduces image-classification inference time by up to 76.0% and 23.3% lower power on ic2 versus the same Cortex-M33 configuration. Additionally, the STM32H7P preview, with its Neural-ART NPU, improves inference time by up to 96.0% on the same workload compared to its CM7-only baseline.
  • Sensing-hub architectures are advancing rapidly. Qualcomm’s submission on the Snapdragon 8 Elite Gen 5 Sensing Hub reports legacy workload latencies below 0.30 ms for keyword spotting, visual wake words, and image classification.
  • Power efficiency is winning on streaming workloads. Syntiant’s NDP120 runs the streaming wake-word benchmark at a 3.3% duty cycle, meaning the device is idle ~97% of the time, with the remaining capacity available for concurrent tasks like noise cancellation or beamforming. 
  • Open Division innovation is healthy. The University of Leeds reports a median throughput of 1,849.7 inferences per second for image classification on its Versal-based heterogeneous platform, with a custom on-fabric batch-normalization engine running at 313 MHz.

Together, the round shows that the field is maturing beyond CPU-only baselines. Hardware acceleration, streaming workloads, and heterogeneous compute have become the defining characteristics of competitive TinyML systems.

Conclusion

Standardization is what makes engineering progress measurable. By fixing the models, datasets, accuracy targets, and energy measurement methodology, MLPerf Tiny allows the entire stack, including silicon, compilers, runtimes, and models, to improve against a shared reference. That shared foundation is how responsible and efficient edge AI gets built. It depends not on the claims of any single vendor, but on a community-maintained benchmark that any submitter can reproduce and any practitioner can trust.

To explore the full v1.4 results, visit the MLPerf Tiny results page and read the supplemental containing details of all our submitters’ results. To contribute to future rounds, join the MLPerf Tiny Working Group.

The post The Benchmark Behind the Next Wave of Ultra-Low-Power AI appeared first on MLCommons.

]]>
MLCommons Releases MLPerf Training v6.0 Results https://mlcommons.org/2026/06/mlperf-training-v6-0-results/ Tue, 16 Jun 2026 14:56:24 +0000 https://mlcommons.org/?p=4039 New benchmarks and increased diversity of submissions reflect important changes in AI ecosystem

The post MLCommons Releases MLPerf Training v6.0 Results appeared first on MLCommons.

]]>
Today, MLCommons® announced new results for the MLPerf® Training v6.0 benchmark suite. The two new benchmarks added in this round, and the submissions received, highlight rapid and significant changes in the AI ecosystem.

“It’s an exciting moment for the community,” said Shriya Rishab, MLPerf Training Working Group co-chair. “We’re seeing strong convergence on a set of best practices for training AI models, but at the same time there is increasing technical diversity in the underlying frameworks and systems that are being used to host and run them.”

MLPerf Training v6.0 adds two new benchmarks, emphasizing sparse computation

The MLPerf Training benchmark suite comprises full system tests that stress models, software, and hardware for a range of machine learning (ML) applications. The open-source and peer-reviewed benchmark suite provides a level playing field for competition, driving innovation, performance, and energy efficiency across the industry. The suite’s benchmark collection is curated by a panel of experts from the AI community.

Version 6.0 adds two new benchmarks: DeepSeek V3 and GPT-OSS 20B, both highlighting the industry-wide shift to sparse computation as exemplified by a Mixture-of-Experts (MoE) architecture. Mixture-of-Experts is a model architecture that uses a smart “router” to send different tokens to specialized sub-networks (“experts”). This enables using a high-parameter-count model that is very efficient because training and inference only activate a fraction of the experts for any given token, reducing the computational cost. 

DeepSeek V3 is a large-scale pretraining model, utilizing an MoE architecture. It uses 671 billion total parameters, of which 37 billion are activated per token. It provides a standardized platform for evaluating the training efficiency of a leading open-weights MoE model at production scale.

GPT-OSS 20B, also an MoE model, uses a much smaller footprint: 21 billion total parameters, of which 3.6 billion are activated per token. This allows organizations to evaluate the complex routing logic and sparse computation patterns common to MoE architecture on hardware configurations as small as a single 8-GPU node.

“Sparse computation is a dominant trend in AI right now,” said Rishab. “Over the past two years, all of the major new generative AI models have utilized a sparse computation architecture, frequently MoE. We have introduced our new DeepSeek V3 benchmark to test large-scale sparse computation training systems, and in fact it is now the largest benchmark in our suite with 671 billion parameters. It also exercises the performance of critical innovations that are now standard in the industry, including Multi-head Latent Attention (MLA) and auxiliary-loss-free load balancing.

“On the opposite end of the spectrum, we’ve introduced the GPT-OSS 20B benchmark as an entry point for organizations that may not have the resources to train the largest-scale models, but want to build advanced capabilities. We’ve carefully designed the benchmark for this scenario, including training from randomized weights to avoid the overhead of multi-gigabyte checkpoint downloads; using the same dataset as existing benchmarks in the suite such as Llama 3.1 8B; and choosing a representative sliver of end-to-end training to reduce the cost of generating benchmark results without compromising on the quality of the benchmark.

“Both of these new benchmarks saw quick uptake, drawing many results. Stakeholders clearly see the importance of performance benchmarking for MoE architectures.”

Increasing diversity of submissions highlights new paths to AI training

Version 6.0 set new records for diversity of the systems submitted. Participants in this round of the benchmark submitted 95 unique systems, utilizing thirteen different hardware accelerators, 19 different host processors and a couple of different software frameworks. 60% of the systems were multi-node.

Notably, there are more than double the number of cloud systems submitted compared to the version 5.1 results six months ago, reflecting the emerging market for hosting AI training in the cloud.

“There are more ways of getting your AI training than ever before,” said Pavan Yalamanchili, MLPerf Working Group co-chair. “Several companies now offer training systems in the cloud, complementing the on-premises systems that continue to be built out at a furious pace. And we are excited to see so many competitive submissions from a variety of on-premises and cloud providers.”

At the same time, the submissions illustrate growing technical diversity, reflecting a robust, rapidly advancing ecosystem. For example, submitters used multiple different FP4-precision recipes, reflecting the current diversity and exploration across the industry. 

“The diversity of FP4 implementations we see in the submissions is not surprising,” said Yalamanchili. “Some implementations are more flexible than others, which allow them to be used in unique training scenarios. But here is where MLPerf’s benchmarking delivers critical insight and value: it allows stakeholders to understand which implementations deliver the best performance for their specific needs. In particular, because MLPerf benchmarks require submissions to meet an accuracy threshold, we shine a spotlight on the differences in performance that these kinds of hardware and implementation design choices can lead to.”

Record industry participation points to broad ecosystem, driven by generative AI

The MLPerf Training v6.0 round includes performance results from 24 submitting organizations: AMD, ASUSTeK, Azure, Cisco, CoreWeave, Dell, Fujitsu, GigaComputing, Google, HPE, Inventec, Krai, Lambda, MITAC, Nebius, Netweb Technologies India LTD, NVIDIA, Oracle, Quanta Cloud Technologies, SCITIX, Supermicro, tinycorp, TTA and Vultr.  “We would especially like to welcome first-time MLPerf Training submitters, Inventec, Netweb Technologies India LTD, TTA and Vultr ” said David Kanter, Head of MLPerf at MLCommons.

Robust participation by a broad set of industry stakeholders strengthens the AI ecosystem as a whole and helps to ensure that the benchmark is serving the community’s needs. We invite submitters and other stakeholders to join the MLPerf Training working group and help us continue to evolve the benchmark.

View the results

Please visit the Training benchmark page to view the full results for MLPerf Training v6.0 and find additional information about the benchmarks. To learn about each submitter’s results, read the supplemental.

About ML Commons

MLCommons is the world’s leader in AI benchmarking. An open engineering consortium supported by over 125 members and affiliates, MLCommons has a proven record of bringing together academia, industry, and civil society to measure and improve AI. The foundation for MLCommons began with the MLPerf benchmarks in 2018, which rapidly scaled as a set of industry metrics to measure machine learning performance and promote transparency of machine learning techniques. Since then, MLCommons has continued using collective engineering to build the benchmarks and metrics required for better AI – ultimately helping to evaluate and improve AI technologies’ accuracy, safety, speed, and efficiency.

For additional information on MLCommons and details on becoming a member, please visit MLCommons.org or email participation@mlcommons.org.

Press Inquiries: contact press@mlcommons.org

The post MLCommons Releases MLPerf Training v6.0 Results appeared first on MLCommons.

]]>
MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMs https://mlcommons.org/2026/06/mlperf-mobile-v6/ Mon, 15 Jun 2026 14:40:00 +0000 https://mlcommons.org/?p=4021 Test LLM inference natively on mobile devices with new standardized benchmarks and expanded NPU acceleration.

The post MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMs appeared first on MLCommons.

]]>
MLCommons® today announced the release of MLPerf® Mobile v6.0, introducing new generative AI benchmark tests for running large language models (LLMs) on Android devices. These tests join a comprehensive suite of benchmarks built into the MLPerf Mobile app, including tests for image generation, object detection, super resolution, and more.

NEW ON-DEVICE LLM BENCHMARKS
MLPerf Mobile v6.0 adopts these models for its new LLM benchmarks:

  • Llama 3.2 1B Instruct
  • Llama 3.2 3B Instruct
  • Llama 3.1 8B Instruct

The models are asked to process requests selected from the TinyMMLU and IFEval datasets to quantify  the performance and accuracy of on-device AI inference.

The LLM tests can run devices with sufficient memory via the CPU–without tailored acceleration. Additionally, this release supports NPU-accelerated execution of the Llama 3.1 8B Instruct model on Qualcomm Snapdragon 8 Elite Gen 5 SoCs. The working group plans to expand LLM acceleration support to more devices and platforms in the future.

EXPANDED SOC SUPPORT AND BROAD AVAILABILITY
In keeping with the MLPerf Mobile working group’s commitment to integrating support for new devices rapidly, the v6.0 release adds support for devices based on the MediaTek Dimensity 9500 Series.

Support for the following chips is also updated in this release:

  • Qualcomm Snapdragon 8 Elite Gen 5
  • Samsung Exynos 2600

Of course, the app already supports NPU-accelerated execution on a host of mobile devices.

The MLPerf Mobile app is available via the Google Play store, the Apple App Store, and via the MLPerf Mobile GitHub repo. More details about device support are available at any of these distribution points. The GitHub repo also contains the full, open-source code for the MLPerf Mobile app, released under the permissive Apache 2.0 license.

ABOUT MLCOMMONS
MLCommons is an open engineering consortium with a mission to make machine learning better for everyone. The organization develops industry-leading benchmarks, datasets, and best practices spanning cloud, data center, edge, and client AI systems. Its MLPerf benchmark suite is widely recognized as the standard for measuring machine learning performance.

The post MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMs appeared first on MLCommons.

]]>
The patch model is breaking. AI evaluation needs a new way to disclose what it finds. https://mlcommons.org/2026/06/responsible-disclosure/ Tue, 09 Jun 2026 14:50:28 +0000 https://mlcommons.org/?p=4030 - Why traditional vulnerability disclosure fails for open-weight models—and how we are building a new standard for AI evaluation.

The post The patch model is breaking. AI evaluation needs a new way to disclose what it finds. appeared first on MLCommons.

]]>
For about thirty years the security community has relied on a well-understood approach for handling dangerous findings. Coordinated vulnerability disclosure is a standard practice for a reason, and it can neatly solve hazard disclosure problems with transparency and technical rigor. A security researcher finds a flaw, reports it privately to the vendor, the vendor ships a fix, deployers update, and only once that window has closed do the details go public. This best practice approach works because of a core critical assumption: the affected system can be repaired, and repairing it ends the hazard.

That assumption does not survive contact with AI systems. We recognized this early, in the course of building our safety and jailbreak benchmarks, and it has since become one of the defining governance problems in evaluating frontier models. The findings an evaluation produces are valuable precisely because they describe how a system behaves under pressure. A fundamental challenge is that the value of that finding does not respect the boundary between defenders and adversaries. At MLCommons, we are addressing this challenge in two ways: designing our own disclosure practice for the benchmarks we run and helping write the standard that supports the whole field of AI evaluation.

Why coordinated disclosure breaks down for AI

Three properties of AI evaluation impact the traditional model for responsible disclosure.

The findings are dual-use by nature. A result that tells defenders, regulators, and users how a system behaves tells adversaries the same thing. It effectively shines a spotlight on which systems, which categories of input, and which failure modes are worth their effort. The risk isn’t usually that a finding exposes an otherwise secret capability; it’s that it lowers the cost of locating one. We are describing uplift – the reduction in effort, time, expertise, or resources an actor needs to accomplish a task. Uplift is most of what makes AI valuable for legitimate users, which is exactly why it is dangerous to hand to the wrong ones. There is also a subtler trap. If results are published by default and one category is quietly left out, the omission itself becomes a signal. The structure of a disclosure carries information independent of its content.

Telling the developer too much corrupts the test. A benchmark meant to be run more than once faces a tension that one-off vulnerability reports have not needed to deal with. The developer of a system under test needs enough feedback to improve the general property being measured, but not so much that they can target the specific items on the test. Hand over the exact prompts, and you get a model that scores better on a test without improving in practice, and, over enough cycles, the benchmark score stops tracking the thing it was built to measure. The discipline is to communicate the general case, never the instances, and never to accept self-attestation alone as proof that something was fixed. This is a general challenge with the legitimacy and reliability of benchmarking evaluations in AI, which MLCommons is addressing in multiple ways, for example with our continuous prompt stewardship work. 

You can’t patch a released, open-weight model. This is the property that breaks the assumptions of the prior model of responsible disclosure. Open-source software can be patched in place; a deployer updates and the hazard closes. An open-weight model cannot. A new version is a new artifact, not an update. Every copy of the prior weights stays operational, unmodified, and in the hands of anyone who retained them – indefinitely. A hazard identified in such a system persists in deployment even after a successor ships, and no defender is positioned to remediate it. If a CBRNE hazard is found in a prior model deployment, for example, that hazard now exists indefinitely. Findings, therefore, have to be pinned to specific versions, and in the most sensitive categories, results may need to be aggregated or uniformly withheld across systems — because granular per-model disclosure in those categories functions less like a public-interest report and more like a targeting map for systems nobody can fix.

From principle to standard

We are taking the action to codify a defensible response to this challenge. We’ve taken our approach and practices into ISO/IEC JTC 1/SC 42, the international body responsible for AI standards, for review and discussion. We are contributing these learnings and corresponding responsible-disclosure principles into the work on ISO/IEC TS 42119-8. The aim is a real, citable standard that any evaluator — first-party, second-party, or independent — can build on, rather than a patchwork of one-off policies. Coordination around this issue will be critical to ensure the most hazardous findings are addressed by good actors without being broadcast to bad ones.

Disclosure norms only work if they are shared, and shared norms come from standards bodies, not from any single lab or benchmark operator acting alone.

What this means for our jailbreak benchmark

When our jailbreak benchmark launches, it will ship with a documented responsible-disclosure policy built around these three considerations — protecting the public from harmful uplift, protecting the integrity of the evaluation over repeated runs, and protecting against hazards in systems that cannot be centrally remediated. That policy is deliberately aligned, in advance, with the standard taking shape in SC 42. We would rather launch already pointed in the direction the field is heading than retrofit a practice once the standard lands.

The patch model gave the software security world a common language for three decades. AI evaluation needs its own. We recently launched an agentic-focused security working group to tackle critical challenges in AI security, such as this one, in the coming years. We welcome you to join us on that journey by signing up to be an MLCommons member.

The post The patch model is breaking. AI evaluation needs a new way to disclose what it finds. appeared first on MLCommons.

]]>
Chakra Comes of Age: A Standardized Trace Ecosystem for AI Systems Benchmarking and Co-design https://mlcommons.org/2026/06/chakra-comes-of-age/ Tue, 02 Jun 2026 14:27:00 +0000 https://mlcommons.org/?p=4009 From a working group proposal in 2023 to a 40+ member industry effort with native support in PyTorch, NVIDIA NeMo, ASTRA-sim, and commercial simulation and emulation tools — MLCommons Chakra is reshaping how the AI systems community studies, reproduces, and co-designs training and inference platforms.

The post Chakra Comes of Age: A Standardized Trace Ecosystem for AI Systems Benchmarking and Co-design appeared first on MLCommons.

]]>

When MLCommons announced the Chakra working group in July 2023, the premise was simple but ambitious: AI systems are moving too fast for the traditional benchmarking and co-design playbook. Production workloads live behind walls of proprietary code and models. Simulators, emulators, and replay tools each invent their own representations. And the workloads driving the next generation of AI supercomputers — frontier-scale LLM training, sparse Mixture-of-Experts (MoE) models, disaggregated inference — change at a pace that the industry has never seen before.

On May 21, 2026, at the MLSys 2026 Industry Track, the Chakra working group presented a comprehensive paper on what that vision has grown into: an open, interoperable ecosystem for performance benchmarking and software/hardware co-design across the AI stack. The paper, MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces, is a milestone for an effort co-chaired by Srinivas Sridharan (NVIDIA) and Tushar Krishna (Georgia Institute of Technology), with active contributions from industry and academia, including but not limited to NVIDIA, AMD, Meta, Keysight, HPE, Scala Computing, Georgia Institute of Technology, and Harvard University.

The paper is available here: MLCommons Chakra MLSys 2026 Paper (arXiv)

The problem: a fragmented co-design loop

Designing modern AI platforms, with clusters of thousands of NPUs (NVIDIA Hopper/Blackwell, AMD Instinct, Google TPU, and others) connected by high-speed scale-up and scale-out fabrics, is an iterative loop. Teams observe workloads in production, reproduce them via representative benchmarks, design and evaluate next-generation systems with simulators and emulators, validate the silicon, and finally deploy at scale. Then the cycle starts again.

The Chakra paper argues that the tools driving this loop are deeply fragmented. Hyperscalers and cloud service providers can’t easily share proprietary models. Simulators and emulators from compute, network, and switch vendors each speak their own dialect. Open benchmarks like MLPerf evolve on release cycles that lag the pace of AI innovation. The result is siloed optimization, longer time-to-market for new platforms, and limited ability for academics and startups to participate in production-relevant co-design.

The Chakra Answer: An Open Execution Trace Ecosystem

At the center of Chakra is the Execution Trace (ET)—a portable, graph-based representation of distributed AI workload behavior.
Rather than exposing model weights, datasets, or proprietary source code, Chakra traces capture the performance-relevant behavior of workloads: compute operations, communication patterns, memory activity, dependencies, timing information, and parallelization strategies.
The idea is deceptively simple but powerful:
A software organization can share an execution trace with a hardware vendor. The hardware vendor can drive internal simulators, emulators, or replay frameworks using that trace. Insights flow back to improve performance, without requiring either side to exchange proprietary IP.
In many ways, Chakra serves as a common language for AI infrastructure benchmarking and co-design, enabling more agile collaboration across workloads, software stacks, silicon, networking, and systems research.

Enabling the Full AI Infrastructure Lifecycle

One of Chakra’s defining strengths is its ability to support the entire AI systems co-design lifecycle through its common execution trace abstraction. Chakra traces can be collected directly from modern AI frameworks such as PyTorch, NVIDIA NeMo, and vLLM, enabling faithful representations of real training and inference workloads. These traces can then be replayed on existing platforms for debugging, performance analysis, and bottleneck identification without requiring access to proprietary models or datasets; simulated or emulated to study future architectures, interconnects, and deployment strategies using realistic workload behavior rather than synthetic approximations; and used for hardware-in-the-loop (“shift-left”) validation, helping organizations uncover system bottlenecks and subtle interactions earlier in the development cycle before large-scale deployment.

From Working Group to Ecosystem

“What began as an idea to make AI workload behavior more portable and reproducible has grown into a vibrant community effort spanning hyperscalers, silicon vendors, infrastructure providers, tool developers, and researchers. Seeing the community rally around shared tooling and interoperable methodologies has been incredibly energizing. This milestone reflects a shared commitment to making AI infrastructure innovation more collaborative and accessible to the broader community.”
— Srinivas Sridharan and Tushar Krishna, Co-Chairs, MLCommons Chakra Working Group


Today, the Chakra Working Group spans a 40+ member collaboration across hyperscalers, full-stack AI vendors, silicon providers, networking companies, system integrators, simulation and emulation tool suppliers, startups, and academia.
The ecosystem has expanded significantly since its launch:

  • Native support in PyTorch and NVIDIA NeMo, with trace collection now officially supported in both frameworks
  • Integration into vLLM, enabling modern inference and serving trace collection
  • Native support in ASTRA-sim, the open-source distributed AI simulator widely used for AI system design-space exploration
  • Commercial adoption in tools such as Keysight AI Data Center Builder and Scala Computing platforms
  • Support within proprietary internal simulation environments, including AMD workflows and Meta, NVIDIA trace replay infrastructures

Part of the Chakra Team at MLSys 2026. From Left: Jinsun Yoo (GT), Winston Liu (Keysight), Hanjiang Wu (GT), Huan Xu (GT), David Kanter (MLC), Tushar Krishna (GT), Vinay Ramakrishnaiah (AMD), Spandan More (AMD), Brad Beckmann (AMD), Brian Coutinho (NVIDIA), Vijay Janapa Reddi (Harvard)

Chakra Presentation at MLSys 2026 by Brian Coutinho (NVIDIA)

Open Trace Library for the Community

The core goal of Chakra is to make AI systems benchmarking and co-design more reproducible and accessible to the broader community.
To support this vision, the Chakra Working Group is also releasing an Open Trace Library for the Community—a growing collection of representative execution traces from real AI workloads spanning diverse models, parallelization strategies, and deployment scenarios.
The initial trace releases were collected with support from Georgia Tech’s AI Makerspace and Hewlett Packard Enterprise (HPE), leveraging production-scale GPU infrastructure to capture representative workloads across GPT-3, Llama, Mixtral, DeepSeek-MoE, and other distributed AI systems. These traces enable researchers, startups, and infrastructure teams to benchmark systems, reproduce workload behavior, and evaluate future AI platform designs—without requiring access to proprietary production environments.
Explore the Open Trace Library here: Chakra Trace Library
Access the Chakra codebase here: MLCommons Chakra Repository

Voices from the ecosystem

“Chakra is an indispensable framework that enables to debug AI systems and optimize performance. Standardization of execution traces enables cross-workload optimizations and debug at datacenter scale. Real-time traces along with open analysis tools helps to understand behavior and execution bottlenecks of highly-parallel distributed workload. This technology can be applied at design phase of integrated datacenter-scale computer. We’re proud to support the MLCommons Chakra effort and the collaboration behind this work” Michael Kagan, CTO, NVIDIA

“Optimizing AI benchmarks and workloads across AMD Instinct and ROCm — from inference serving to fine-tuning and pre-training — means navigating a system design space where robust methodology and simulation-driven exploration accelerate time-to-performance. Chakra’s graph-based execution traces capture real AI workload behavior in a portable, vendor-neutral schema, giving the industry a common language for benchmarking and co-design and giving AMD a faithful representation of MLPerf-class workloads to drive internal simulation on AMD Instinct silicon. That interoperability turns architectural advantage into reproducible performance.” — Meena Arunachalam, Fellow and Director, AI Performance Engineering, AMD,  Board Member MLCommons

“I had the privilege of being associated with Chakra since its inception. My former teammates at Meta started Chakra to enable platform agnostic way collecting production workload traces. It greatly enhanced the capabilities of engineering teams at Meta to instrument and analyze potential performance issues at scale. In my current role at AMD I get to notice how Chakra is enabling performance analysis of our products and driving crucial roadmap decisions.”  – Shashidhar Gandham, Former Director of Networking at Meta, now Corporate VP, AMD

“Chakra has become a foundational component of our AI systems research at HPE, enabling high-fidelity performance emulation and large-scale design exploration for next-generation AI infrastructure. Our teams have leveraged Chakra to study training, parallelization, and resilience strategies for extreme scale AI models well before access to hyperscale GPU deployments. Its presentation at MLSys 2026 underscores the growing importance of open, standardized trace ecosystems in accelerating innovation across the AI systems community.” – Puneet Sharma, Fellow and VP, HPE

“One of the persistent challenges in AI infrastructure development is that system-level behaviors, particularly interactions between collective communication, memory pressure, and congestion control, often surface after clusters are deployed, when the cost of correction is higher. Shift-left validation addresses this, but only if workloads driving that validation are realistic. Chakra’s execution trace format gives the entire ecosystem a portable, workload-faithful representation that can be shared across organizational boundaries without exposing proprietary models or source code. Keysight has invested in this effort because we believe open, interoperable foundations like Chakra are essential for helping the industry close the feedback loop faster rather than operating in isolated silos.”  Ram Periakaruppan, Vice President and GM, Network Applications & Security business at Keysight. More perspective from Keysight here

“Chakra is a fantastic showcase of the role Georgia Tech plays in connecting academic research with real-world systems. We can bring together expertise spanning the full AI stack in really the only way that makes complex work like this possible”  Arijit Raychowdhury, Steve W. Chaddick School Chair of ECE, Georgia Institute of Technology. More perspective from Georgia Tech here

​​“Open benchmarks and open tooling have historically played a foundational role in accelerating progress in computing because they give researchers and practitioners a common way to evaluate ideas, reproduce results, and build on one another’s work. Chakra extends that philosophy to modern AI infrastructure by creating a shared representation of workload behavior, helping bridge insights from academia and industry while democratizing access to realistic AI systems research and enabling deeper collaboration across the ecosystem.” — Vijay Janapa Reddi, Harvard University; Board Member, MLCommons

MLCommons builds shared standards that move the entire ecosystem forward like MLPerf. Chakra brings that philosophy to AI systems co-design, giving researchers and companies a common framework for understanding and reproducing workload behavior across hardware, software, and infrastructure boundaries.” – David Kanter, Co-Founder and Head of MLPerf, MLCommons

Looking Ahead

The Chakra ecosystem continues to evolve.
Future efforts include support for larger and more efficient trace representations for frontier-scale models, tighter integration with benchmarking efforts such as MLPerf Storage, and emerging work on InfraGraph, a portable infrastructure description framework designed to complement Chakra’s workload traces with graph-based descriptions of compute, memory, and network topologies.
AI infrastructure is advancing at extraordinary speed, and no single organization can solve the systems challenges ahead alone.
Chakra represents an important step toward a more collaborative AI systems ecosystem—one where workloads, tools, and methodologies become easier to share, reproduce, benchmark, and optimize across the stack.
To get involved, contribute traces or tooling, or join the working group, visit the MLCommons Chakra community pages: MLCommons Chakra Working Group

—–

Chakra means “wheel” in Sanskrit — a fitting name for an ecosystem built around the cyclic, iterative nature of AI systems co-design.

The post Chakra Comes of Age: A Standardized Trace Ecosystem for AI Systems Benchmarking and Co-design appeared first on MLCommons.

]]>
Introducing the 2026 MLCommons Rising Stars https://mlcommons.org/2026/05/2026-rising-stars/ Tue, 19 May 2026 15:34:00 +0000 https://mlcommons.org/?p=3994 Fostering a global community of emerging leaders at the intersection of ML and systems research

The post Introducing the 2026 MLCommons Rising Stars appeared first on MLCommons.

]]>
The MLCommons Rising Stars program is designed to support and connect early-career researchers working at the intersection of machine learning (ML) and systems. Through this initiative, participants engage with a vibrant global community, connect with leaders across academia and industry, and further develop their technical and professional skills.

We are excited to announce the 4th annual MLCommons Rising Stars cohort, featuring 39 outstanding junior researchers from 26 institutions worldwide. Selected from a highly competitive pool of over 175 applicants, these individuals have demonstrated exceptional promise in ML, systems, and data systems research. They stand out not only for their current achievements but also for their potential to shape the future of the field.

This year’s cohort reflects both the depth and breadth of emerging talent in the field. A strong majority of Rising Stars are advanced PhD students (primarily in their 3rd-6th years) alongside a select group of postdoctoral researchers. Their research spans a wide range of topics, including large language models, ML systems efficiency, hardware-software co-design, trustworthy AI, multimodal learning, and applications in domains such as healthcare, cybersecurity, and scientific computing. Notably, many projects emphasize scalability, efficiency, and real-world deployment, supporting a growing shift toward practical, systems-driven ML innovation.

The cohort also reflects the increasingly global and interdisciplinary nature of the ML research community. Participants represent leading institutions across North America, Europe, Asia, and Australia, with a meaningful international presence that contributes to the exchange of ideas among regions. At the same time, the program continues to make progress toward broadening participation: this year’s cohort includes researchers from a range of gender identities and backgrounds, with women and gender-diverse researchers comprising 28% of participants.

As part of the program, we will host the Rising Stars Workshop at AMD headquarters in Santa Clara, California, on July 30-31. During the workshop, participants will present their research, explore emerging opportunities, participate in career development sessions, and build lasting connections with peers and mentors across sectors.

Collectively, this year’s cohort reflects the growing importance of research on machine learning systems in shaping the future of AI.

“MLCommons Rising Stars highlights the researchers helping shape the future of AI engineering across the machine learning systems stack, spanning algorithms, models, systems, hardware, and scalable infrastructure. The rapid evolution of AI increasingly depends on advances across these interconnected layers,” said Vijay Janapa Reddi, Vice President of MLCommons and professor at Harvard University.

We warmly congratulate this year’s Rising Stars and thank all who applied for their interest and enthusiasm.

We also extend our sincere appreciation to the Rising Stars organizers – Abdulrahman Mahmoud (MBZUAI), Akanksha Atrey (Nokia Bell Labs), Muhammad Husnain Mubarik (AMD), Sercan Aygun (University of Louisiana at Lafayette), and Udit Gupta (Cornell Tech) – along with the broader organizing and program committee for their dedication in assembling this exceptional cohort. Finally, we thank Dave Graham (MLCommons) and Ralph Witting (AMD) for their invaluable support in bringing the program and workshop to life.

ML & Systems Rising Stars 2026

Recognizing exceptional early-career researchers in machine learning and systems

Simeon Adebola
Simeon Adebola
University of California Berkeley
Simeon Adebola is a fifth-year PhD student in Electrical Engineering and Computer Sciences (EECS) at UC Berkeley, advised by Ken Goldberg and a member of AUTOLab and Berkeley Artificial Intelligence Research (BAIR). His research lies at the intersection of machine learning, robotics, and systems, with a focus on scalable infrastructure, datasets, and representations for real-world embodied AI. He previously earned a master’s degree from Middle Tennessee State University, where he worked with Lei Miao on robotics for fall detection and received the Outstanding Master’s Research Award from the College of Basic and Applied Sciences. He has also conducted research at UCLA and Vanderbilt University.
Yuetao Chen
Yuetao Chen
The Chinese University of Hong Kong
Yuetao Chen is a Ph.D. student in Computer Science and Engineering at The Chinese University of Hong Kong, advised by Prof. Hong Xu. His research lies at the intersection of machine learning and computer systems, focusing on efficient and scalable systems for modern ML workloads, including large language model inference, speculative decoding, distributed training, and high-performance computing. His work has appeared in top-tier venues such as MLSys, ASPLOS, PPoPP, and FAST. His paper on Tensor Core-based stencil computation received the Best Paper Award at PPoPP 2024. He has also conducted research at Microsoft Research Asia and contributed to academic services such as EuroSys artifact evaluation.
Valerie Chen
Valerie Chen
Carnegie Mellon University
Valerie is a Machine Learning PhD student at CMU. Her work bridges machine learning, natural language processing, and human-computer interaction to advance the design of collaborative AI systems. Her research has fostered close collaborations with major engineering and financial companies, with findings cited by leading model providers and deployed in industry products. Valerie has been recognized with the Rising Stars in Data Science award, CMU Presidential Fellowship, and the NSF Graduate Research Fellowship. Her research has also received various awards, including Best Paper at a NeurIPS workshop and Oral Presentations at ICLR and AAAI.
Yuzong Chen
Yuzong Chen
Cornell University
Yuzong Chen is a final-year PhD student in the School of Electrical and Computer Engineering at Cornell Tech, advised by Prof. Mohamed Abdelfattah. His research focuses on Algorithm-Hardware Co-Design for Machine Learning Acceleration, with a special focus on quantization numerics, FPGA architectures, and processing in-memory. He was named a ML and Systems Rising Star by MLCommons in 2026 and was a finalist for the 2024 Qualcomm Innovation Fellowship.
Jae-Won Chung
Jae-Won Chung
University of Michigan
Jae-Won Chung is a fifth-year PhD candidate in Computer Science and Engineering at the University of Michigan, advised by Professor Mosharaf Chowdhury. He builds efficient software systems for machine learning, with a focus on treating energy as a first-class systems resource to be carefully measured, optimized, and allocated alongside time. He created and leads the ML.ENERGY initiative, a cross-institutional effort, and his research and open-source work, including the Zeus library, have been recognized and adopted by NVIDIA, Google, Microsoft, the PyTorch Foundation, and GitHub.
Stefany Cruz
Stefany Cruz
University of Washington
Stefany Cruz is a Washington Research Foundation Postdoctoral Fellow at the University of Washington’s Paul G. Allen School of Computer Science and Engineering, where she works with Dr. Vikram Iyer. In her research, she builds on-device agentic AI for urban safety and sustainability, focusing on real-time, privacy-preserving sensing and decision-making. She received her PhD from Northwestern University’s Department of Electrical and Computer Engineering, where she was supported by the Ada Lovelace Microsoft Research PhD Fellowship and received the Best Dissertation Award in Computer Engineering.
Vasisht Duddu
Vasisht Duddu
University of Waterloo
Vasisht Duddu is a final-year Ph.D. candidate at the University of Waterloo, advised by Prof. N. Asokan. His research is on trustworthy machine learning, covering the design of novel attacks and defenses to mitigate them, studying deployment trade-offs in ML systems, and designing technical mechanisms to support governance. His work has appeared in various top-tier security, privacy, and machine learning venues (e.g., IEEE S&P, ACL, ACM CCS, TMLR, ICML, and PETS). He is a recipient of the IBM Ph.D. Fellowship (2024), a Distinguished Paper Award at IEEE S&P (2024), and a Best Paper Award at ACM CODASPY (2025).
In Gim
In Gim
Yale University
In Gim is a fourth-year Ph.D. student at Yale University, advised by Prof. Lin Zhong. His research focuses on systems for machine learning, with an emphasis on scalable and programmable abstractions that bind accelerators, models, and application logic together. He currently leads the development of Pie, a programmable LLM serving system that enables dynamic application logic to run natively within the inference engine. His first-author works have appeared at venues including SOSP, MLSys, MobiSys, HotOS, EMNLP, and AAAI.
Alicia Golden
Alicia Golden
Harvard University
Alicia is a fourth-year PhD candidate at Harvard University, advised by David Brooks and Gu-Yeon Wei. Her research sits at the intersection of computer architecture, machine learning, and systems, with a focus on designing efficient hardware systems for large-scale AI. Prior to Harvard, she completed her undergrad at Cornell University, where she received a B.S. in Electrical and Computer Engineering.
Yongjun He
Yongjun He
ETH Zurich
Yongjun He is a fifth-year PhD student in the Systems Group of the Department of Computer Science at ETH Zurich, supervised by Prof. Dr. Gustavo Alonso and Prof. Dr. Ana Klimović. Before 2024, he was supervised by Prof. Ce Zhang and Dr. Theodoros Rekatsinas. Prior to that, he earned his M.Sc. from Simon Fraser University and his B.Eng. from Nanjing University. His research focuses on building systems that democratize the customization and deployment of generative AI models across diverse computing environments, from personal devices to large-scale clusters. In parallel, he has also been working on specialized hardware accelerators such as FPGAs for machine learning pipelines.
Muyan Hu
Muyan Hu
University of Illinois Urbana-Champaign
Muyan Hu is a 3rd-year PhD student at UIUC, advised by Prof. Charith Mendis and Prof. Vikram Adve. His research interests focus on MLSys and compilers, especially mid-end compiler optimizations for AI models, including operator fusion and data movement optimization at the graph and kernel levels. His work has been published at top compiler and system venues, including OSDI and ASPLOS. He is also currently a High-Performance AI Intern at NVIDIA, where he works on agents for compilers and kernels targeting next-generation NVIDIA GPUs. He obtained his B.S. degree in Computer Science from Tsinghua University.
Lanxiang Hu
Lanxiang Hu
University of California, San Diego
Lanxiang Hu is a third-year Computer Engineering PhD student at UCSD, advised by Prof. Hao Zhang and Prof. Tajana Šimunić Rosing. He is also a research intern at NVIDIA. Lanxiang’s research focuses on efficient AI, particularly on developing efficient algorithms and systems for training and serving parallel-decoding transformers. His work also involves evaluations of multimodal agentic workloads.
Yafan Huang
Yafan Huang
University of Iowa
Yafan Huang received his Ph.D. in May 2026 from the Department of Computer Science at the University of Iowa, advised by Prof. Guanpeng Li. He has been a visiting graduate student at Argonne National Laboratory since 2021, where he works with Dr. Sheng Di and Dr. Franck Cappello. His research focuses on high-performance computing (HPC), with particular interests in data compression, fault tolerance, parallel computing, and compiler optimizations for machine learning systems and scientific applications. Yafan is the recipient of the 2025 ACM–IEEE CS George Michael Memorial HPC Fellowship and has received multiple best-paper finalist and award recognitions at systems conferences, including SC’22, SC’24, ICS’25, and LDAV’25.
Yiqiao Jin
Yiqiao Jin
Georgia Institute of Technology
Yiqiao Jin is a Ph.D. candidate in Computer Science at the Georgia Institute of Technology, advised by Prof. Srijan Kumar. His research focuses on reliable and efficient intelligent systems, spanning multi-agent systems, multimodal large language models, and efficient AI. His work studies how large language models can reason, collaborate, and adapt under limited supervision and computational resources. His research has led to over 20 publications in leading AI, ML, NLP, and data mining venues, including ICLR, ICML, ACL, EMNLP, The Web Conference, KDD, AAAI, and ICWSM.
Hyungyo Kim
Hyungyo Kim
University of Illinois at Urbana-Champaign
Hyungyo Kim is a 5th-year Ph.D. student in the electrical and computer engineering department at the University of Illinois at Urbana-Champaign and has worked at IBM Research, Intel, and Samsung as a research intern. He earned his B.S. degree in electrical and computer engineering from Seoul National University. His research focuses on system architectures for AI.
Hyunji Lee
Hyunji Lee
University of North Carolina at Chapel Hill
Hyunji Lee is a postdoctoral researcher at the University of North Carolina at Chapel Hill, where she works with Mohit Bansal. She earned her Ph.D. in AI from KAIST under the supervision of Minjoon Seo. Her research focuses on developing robust semi-parametric models that integrate external knowledge modules, combining the strengths of both parametric and nonparametric representations. She is currently focusing on developing a memory system as a form of nonparametric knowledge across diverse domains.
Wei Li
Wei Li
Carnegie Mellon University
Wei Li (Neway) is a Ph.D. candidate in the Department of Electrical and Computer Engineering at Carnegie Mellon University, advised by Prof. Shawn Blanton and Prof. José Moura. His research is driven by a vision of End-to-End Autonomous EDA, aiming to shift the human role from active “”in-the-loop” intervention to strategic “”on-the-loop” supervision. To realize this, his work integrates multi-modal LLMs to perceive design context, differentiable optimization for high-performance sub-tasks, and hardware testing to anchor AI-driven synthesis in physical silicon reality. Wei is a recipient of the Croucher Fellowship (2026), the Apple PhD Fellowship in Integrated Systems (2022, 2024), and the Qualcomm Innovation Fellowship (2024). His research has been recognized with Best Paper Awards at ASP-DAC (2021), ISSTA (2019), and ICTAI (2019), as well as a Best Paper Honorable Mention at ICLAD (2025). His contributions deliver proven industrial impact, with algorithms integrated into Apple’s industrial physical design flows and a joint patent filing with NVIDIA for differentiable global routing. Looking forward, Wei is committed to establishing the iDAT Lab to pioneer the next generation of intelligent design automation and testing.
Payal Mohapatra
Payal Mohapatra
Northwestern University
Payal Mohapatra recently completed her PhD in Computer Engineering at Northwestern University, advised by Prof. Qi Zhu, and is an incoming Staff Research Scientist at Arm. Her research builds machine learning systems for time-series sensing under real-world deployment constraints: robustness to distribution shifts, efficient multimodal fusion across heterogeneous modalities, and graceful handling of missing or noisy signals. Her work on MAESTRO (NeurIPS Spotlight 2025) replaces expensive pairwise cross-attention with sparse, scalable multimodal fusion, and her phase-anchored representation learning framework (TMLR 2025; Spotlight at TS4H @ NeurIPS 2025) generalizes across cross-domain nonstationary time series. She has applied these methods to industrial fatigue monitoring with Boeing and John Deere (PNAS Nexus 2024), silent speech understanding from surface EMG (ACL 2025), and on-device wearables (through internships at Meta Reality Labs and Mitsubishi Electric Research Labs). She has won grand challenges at IEEE ICASSP (2022) and ACM Multimedia (2023), and was a DAC Young Fellow (2021). Before her PhD, she was an IC Design Engineer at Analog Devices and earned her Master’s from IIT Madras.
Rui Pan
Rui Pan
Princeton University
Rui Pan is a PhD candidate at Princeton, advised by Ravi Netravali. His research lies at the intersection of systems and machine learning, with a recent focus on software infrastructure and algorithms for efficient large-language-model inference. In particular, he studies how to co-design novel models (e.g., hybrid LLMs, reasoning LLMs, diffusion LLMs) and the underlying runtimes that serve them, ensuring they scale and perform efficiently for emerging generations of AI innovations.
Vaidehi Patil
Vaidehi Patil
University of North Carolina at Chapel Hill
Vaidehi Patil is a PhD candidate in the Computer Science Department at UNC Chapel Hill and a Google PhD Fellow. Her research focuses on privacy, safety, and security in LLMs, VLMs, and agentic systems, with an emphasis on unlearning sensitive information, adversarial attack–defense evaluation in unimodal and multimodal settings, and privacy leakage and belief steering in multi-agent LLM systems. Her work has been recognized with a spotlight presentation at ICLR. Vaidehi was also the lead organizer of the ICML 2025 Workshop on Machine Unlearning and Generative AI.
Shvetank Prakash
Shvetank Prakash
Harvard University
Shvetank Prakash is a Ph.D. candidate in computer science at Harvard University. His research focuses on hardware-software co-design for ultra-low-power machine learning systems at the Edge and on developing AI agents for computer architecture design problems. His work has appeared in ASPLOS, ICML, and Nature, and he is a 2026–2027 NVIDIA Graduate Fellow. Prior to Harvard, he received his B.S. in computer engineering from Columbia University.
Wenjie Qu
Wenjie Qu
National University of Singapore
Wenjie Qu is a PhD candidate in Computer Science at the National University of Singapore, advised by Prof. Jiaheng Zhang, and working closely with Prof. Dawn Song. His research lies at the intersection of machine learning, security, and cryptography, with a focus on making large language models trustworthy and verifiable. He develops systems for auditing and proving the correctness of AI services, including zero-knowledge proofs for model inference and training. His work has appeared in top security venues such as IEEE S&P, USENIX Security, ACM CCS, and NDSS, and has influenced both academia and industry.
Derrick Quinn
Derrick Quinn
Cornell University
Derrick Quinn is a second-year Ph.D. student in the Computer Systems Laboratory (CSL) at Cornell University, where he is advised by Professor Mohammad Alian. His research philosophy is driven by the belief that today’s increasingly complex systems require holistic co-design in order to realize their full potential. Currently, he’s focused on co-designing novel algorithms and near-data processing architectures to accelerate dense retrieval and long-context inference. Derrick’s long-term goal is to develop generalized, scalable architectures for Neural Memory Systems, enhancing their capability, adaptability, and sustainability across diverse computing environments.
Md Mostafijur Rahman
Md Mostafijur Rahman
The University of Texas at Austin
Md Mostafijur Rahman is a Ph.D. candidate at The University of Texas at Austin, advised by Radu Marculescu. His research sits at the intersection of AI, biomedical imaging, and computer vision, with a focus on building efficient, reliable, and scalable AI systems for deployment in healthcare under real-world constraints. His work has been translated into practice through research internships at GE Healthcare, the National Institutes of Health (NIH), and Bosch Research. He has published over 20 peer-reviewed papers in venues including CVPR, NeurIPS, MICCAI, and ICCV, with several works selected for Spotlight and Oral presentations. His research contributions have been recognized by the NIH Summer IRTA Fellowship, the Texas Health Catalyst Award, and the Discovery to Impact Award.
Akshat Ramachandran
Akshat Ramachandran
Georgia Institute of Technology
Akshat Ramachandran is a third-year Ph.D. student at the Georgia Institute of Technology, advised by Prof. Tushar Krishna. His research lies at the intersection of computer architecture, systems, numerical formats, and AI/ML algorithms, with a focus on designing efficient hardware–software co-design techniques for emerging AI workloads. His work spans three key areas: (1) next-generation arithmetic and numerical representations for lower-overhead and adaptive computation, (2) post-training model compression methods that assign layer-wise quantization and sparsity attributes, and (3) flexible mixed-precision hardware architectures for efficient AI processing. Toward this research vision, he actively collaborates with both academic and industrial research teams, including NVIDIA Research, Intel Labs, and Samsung Research. His research has resulted in multiple patent filings, a growing publication record in premier computer architecture and AI venues, and several best paper and excellence-in-research awards.
Jie Ren
Jie Ren
Massachusetts Institute of Technology
Jie Ren is a Postdoctoral Associate at MIT CSAIL. He received his Ph.D. from Michigan State University and his bachelor’s degree from Tsinghua University. His research focuses on trustworthy, interpretable, and scalable AI, with particular interests in data and model protection in generative AI and efficient test-time scaling. His work has been published in top-tier venues including NeurIPS, ICLR, ICML, ACL, CVPR, and WWW. His recent research explores how to build AI systems that are both powerful and accountable, especially in the context of large foundation models and generative AI.
Yeonju Ro
Yeonju Ro
The University of Texas at Austin
Yeonju Ro is a Ph.D. student at UT Austin, co-advised by Aditya Akella and Atlas Wang. Her research sits at the intersection of computer systems and machine learning, with a focus on algorithm-system co-design for next-generation AI systems. She is also a contributor to the Learning-directed Operating System expedition and the Infra AI Center at UT. She is a 2024 IBM Ph.D. Fellow and a 2024 Qualcomm Fellowship Finalist.
Amit Samanta
Amit Samanta
University of Utah
Amit Samanta is a Ph.D. student in Computer Science at the University of Utah, advised by Prof. Ryan Stutsman and Prof. Rohan Basu Roy. His research focuses on building ML and HPC infrastructure that is fast, fair, and sustainable, treating carbon and water footprint as first-class metrics alongside performance. He has interned at Lawrence Livermore and Argonne National Laboratories. He is honored to be a 2026 MLCommons ML and Systems Rising Star, an Internet Society Pulse Research Fellow, and a Young Researcher at the Heidelberg Laureate Forum.
Efe Sencan
Efe Sencan
Boston University
Efe Sencan is a Ph.D. candidate in Computer Engineering at Boston University, advised by Prof. Ayse K. Coskun. His research lies at the intersection of machine learning and computer systems, with a focus on building practical ML methods for performance anomaly detection, bottleneck diagnosis, and telemetry-driven analytics in high-performance computing systems. His work aims to make large-scale scientific computing systems more reliable, efficient, and easier to diagnose. He has collaborated with national laboratories and supercomputing centers to apply ML techniques in production HPC environments.
Erfan Shayegani
Erfan Shayegani
University of California, Riverside
Erfan is a 4th-year Computer Science Ph.D. student at the University of California, Riverside. His research focuses on Multimodal Language Models (LLMs/MLLMs) and AI Agents, such as Computer-Use Agents (CUAs), with an emphasis on Alignment, Robustness, Safety, Ethics, Fairness, Bias, and Security/Privacy. He has also completed two research internships at Microsoft Research, and I’m currently interning at Apple, looking at misalignment issues in multimodal models.
Michael Shen
Michael Shen
Cornell Tech
Michael is an Electrical and Computer Engineering Ph.D. student at Cornell Tech and a researcher in the Cornell Computer Systems Laboratory, co-advised by Professor Udit Gupta and Professor G. Edward Suh. His current research interests are in computer architecture and computer systems for efficient machine learning. Recently, he has been focused on improving inference pipelines for Retrieval-Augmented Generation (RAG) and Agentic AI systems.
Sudipta Saha Shubha
Sudipta Saha Shubha
University of Virginia
Sudipta Saha Shubha is a PhD candidate at the University of Virginia. His research sits at the intersection of distributed systems and artificial intelligence (AI). In particular, his research has focused on developing distributed systems for cost-efficient AI inference serving at scale, including generative and agentic AI. His work involves networked infrastructure, operating systems, GPU architectures and kernels, and AI models. He has published his research at top-tier AI systems conferences, including OSDI, SIGCOMM, EuroSys, SoCC, and IPDPS, and his work has been shipped to production in industry through research internships at Microsoft and HPE Labs.
Saranya Vijayakumar
Saranya Vijayakumar
Carnegie Mellon University
Saranya Vijayakumar is a Ph.D. candidate in Computer Science at Carnegie Mellon University, advised by Christos Faloutsos and Matt Fredrikson. Her research focuses on AI security and safety with a particular interest in red teaming and multi-agentic alignment. Before CMU, she did her undergraduate studies at Harvard with a joint concentration in Computer Science and Government, working with Cynthia Dwork and Jim Waldo on algorithmic fairness, and spent three years as a data scientist in electronic trading at Goldman Sachs. She has held research positions at Anthropic, IBM Research, Inria, and Fujitsu Research, and is supported by the DoD NDSEG Fellowship. She founded Women in CSD at CMU.
Ryan Wong
Ryan Wong
University of Illinois Urbana-Champaign
Ryan Wong is a fifth-year Ph.D. student in computer science at the University of Illinois Urbana-Champaign, working with Prof. Saugata Ghose. His research interests are in the broad area of computer architecture, with particular emphasis on memory and storage systems, as well as accelerators for machine learning, scientific computing, and database systems. He is a Mavis Future Faculty Fellow and has won the UIUC CS Outstanding TA Award. He received his B.S. in Computer Science, B.A. in Chemistry, and M.S. in Electrical Engineering, all from the University of Rochester. For more information, please visit his website at https://rwong.cs.illinois.edu/
Sanjali Yadav
Sanjali Yadav
University of Maryland, College Park
Sanjali Yadav is a third-year PhD student at the University of Maryland, College Park, where she also earned her Bachelor’s degree in Computer Science. Advised by Dr. Bahar Asgari, her research operates at the intersection of computer architecture and machine learning, focusing on developing self-adaptive systems. By leveraging machine learning as a foundational architectural mechanism rather than just a target workload, she pioneers the design of systems that autonomously learn and adapt to changing demands throughout their lifecycle. Her work transcends rigid, pre-defined design heuristics in favor of autonomous systems that proactively self-optimize for latency, throughput, and energy efficiency, ensuring architectural agility against the high-variance demands of modern computing. This research philosophy is embodied in her work on Misam, an adaptive framework that earned Sanjali First Place at the ACM Student Research Competition at MICRO 2024. Her contributions, including both Misam and the Boötes framework, have been featured at top-tier venues such as MICRO, establishing a new standard for treating intelligence as a first-class citizen within the hardware stack. By integrating adaptability directly into the architectural fabric, her research continues to push the boundaries of how intelligent computing infrastructure can be effectively deployed in increasingly resource-constrained and dynamic environments.
Ruokai Yin
Ruokai Yin
Meta
Ruokai is a Research Scientist at Meta Reality Labs, where he works on improving the runtime efficiency of AI models on Meta’s custom silicon. He received his Ph.D. in Electrical and Computer Engineering from Yale University. His thesis focused on designing efficient computer architectures, systems, and algorithms for asymmetric AI workloads, including low-precision and sparse LLMs and neuromorphic deep learning models.
Hengrui Zhang
Hengrui Zhang
Princeton University
Hengrui Zhang is a fourth-year Ph.D. student at Princeton University, advised by Prof. David Wentzlaff. His research lies at the intersection of MLSys and computer architecture, with a focus on hardware-software co-design for efficient Generative AI serving.
Tong Zhou
Tong Zhou
Northeastern University/ Microsoft
Tong Zhou is a recent Ph.D. graduate from Northeastern University. Her research centers on secure and responsible AI systems, focusing on integrating protection and accountability directly into model design rather than relying on external safeguards. Her work spans model IP protection, usage control, privacy-preserving inference, and cryptographically verifiable watermarking for generative models. She has published in leading venues across machine learning, security, and hardware systems, including NeurIPS, ICLR, ICML, NDSS, ICCV, ICCAD, and DAC.
Terry Yue Zhuo
Terry Yue Zhuo
Monash University & CSIRO’s Data61
Terry is a final-year PhD student at Monash University, supported by the Data61 PhD Scholarships, IBM PhD Fellowship Awards, the Google Research Scholar Program, and the DAAD AInet Fellowship. His main research interest lies in LLMs for code generation and cybersecurity.

The post Introducing the 2026 MLCommons Rising Stars appeared first on MLCommons.

]]>
GPT-OSS 20B: A Sparse MoE Pretraining Benchmark for MLPerf Training v6.0 https://mlcommons.org/2026/05/gpt-oss-moe-training6/ Thu, 07 May 2026 13:23:00 +0000 https://mlcommons.org/?p=3973 How MLCommons engineered a stable, accessible Mixture-of-Experts (MoE) pretraining benchmark for MLPerf Training v6.0 that runs on a single 8-GPU node.

The post GPT-OSS 20B: A Sparse MoE Pretraining Benchmark for MLPerf Training v6.0 appeared first on MLCommons.

]]>
Motivation and Architectural Relevance

Pretraining large language models (LLMs) requires massive computational resources, and the MLPerf™ Training benchmark suite has reflected that reality. Pretraining tests like Llama 3.1 405B and Llama 3.1 8B focus on dense models that can require substantial multi-node infrastructure. This can create a barrier to entry for many organizations looking to participate in benchmarking. 

To address this, the MLPerf Training Working Group—led by a task force from AMD, NVIDIA, and NIT University—is introducing a new pretraining benchmark: GPT-OSS 20B. A modern, Mixture-of-Experts (MoE) alternative, GPT-OSS 20B allows researchers and organizations to evaluate the complex routing logic and sparse computation patterns common to MoE architectures—on hardware configurations as small as a single 8-GPU node.

Model Selection and Architecture

The task force identified GPT-OSS 20B as the ideal small-scale candidate with a modern, sparse architecture for three reasons:

  • Sparse Efficiency: The model features 21B total parameters, but utilizes an MoE design that activates only 3.6B parameters per token. This allows it to maintain a massive knowledge base with a computational footprint similar to less dense models.
  • From-Scratch Training: To keep the benchmark simple and avoid the overhead of multi-gigabyte checkpoint downloads, GPT-OSS 20B is trained from randomized weights. This makes it a pure test of a system’s ability to optimize a sparse model from its initial state.
  • Reference Implementation: The reference code is built on AMD’s Primus framework, a versatile new training library supporting both AMD and NVIDIA backends. Primary validation was conducted on AMD Instinct™ MI355X and NVIDIA B200 systems.

Dataset and Tokenization

GPT-OSS 20B leverages the C4 (Colossal Cleaned Common Crawl) dataset, specifically using the same pre-tokenized subset and Llama-3 compatible tokenizer as the Llama 3.1 8B benchmark. This makes it easy for submitters who already have the dataset setup for other MLPerf Training benchmarks to run GPT-OSS 20B as well. The training data consists of approximately 80 GB of pre-shuffled C4 shards hosted on MLCommons™ storage. To maintain stability, the benchmark evaluates against the first 1,024 samples of the validation set every 12,288 samples (768 iterations at GBS=16).

The Challenge of Statistical Variance (CV)

A primary objective in benchmarking is ensuring fairness. In large-scale training, this is measured by the Coefficient of Variation (CV), or the ratio of the standard deviation to the mean ($CV = \frac{\sigma}{\mu} \times 100\%$).

Why High CV Undermines Fairness

When a benchmark has a high CV, it suffers from “statistical noise.” If one run takes 170k samples and another takes 250k samples due to random chance, the results no longer reflect hardware or software superiority. For a benchmark to be a trusted industry standard, a user must be able to reproduce the results. High variance makes it impossible to distinguish between an engineering breakthrough and luck.

Engineering Stability: The Path to 5% CV

One common approach to reducing CV is starting from a pretrained checkpoint that has already passed through the unstable early stages of training. The task force wanted to keep the benchmark simple, so requiring submitters to download a multi-gigabyte checkpoint would add overhead and complexity. Instead, the task force implemented three critical technical interventions to reduce CV from approximately 15% to less than 5%.

1. Eliminating Validation Noise

Early tests showed massive, non-representative spikes in evaluation loss.

  • The Discovery: The validation dataset was being shuffled at every evaluation interval. In a sparse MoE model, where routing is highly sensitive to input distribution, this shuffling introduced artificial “jitter.”
  • The Fix: The benchmark now mandates evaluation on a static, unshuffled set of the first 1024 samples of the C4 validation set. This ensures the test remains identical across every step of every run.

2. Stabilizing the Optimizer 

While many modern models use an Adam epsilon ($\epsilon$) of $10^{-8}$, the task force found this caused excessive divergence in 20B-scale MoE training from scratch. By aligning to $\epsilon = 10^{-5}$ (the standard used in Llama 3.1 8B), the team provided the numerical stability necessary to prevent “unlucky” gradient updates from derailing the sparse experts.

3. Standardizing Initialization

Technical Configuration and Quality Metrics

To ensure all participants start with the same “statistical energy,” the task force strictly defined the weight initialization standard (init_method_std = 0.008). This prevents variance born from different starting points in the high-dimensional loss landscape.

The target accuracy for the benchmark is a validation loss (log perplexity) of 3.34. This target was chosen based on extensive sweeps on AMD MI355X and NVIDIA B200 hardware, representing a point of stable convergence that balances thoroughness with a reasonable runtime (~6.5 hours to convergence) using BFloat16 precision.

FeatureSpecification
Model TypeMixture-of-Experts (MoE)
Active Parameters3.6B per token
Sequence Length8,192
Expert Parallelism8
Target Loss3.34
Submission Requirement10 runs per configuration (to average out noise)

To optimize for their specific hardware, submitters are permitted to tune three hyperparameters: global batch size, learning rate, and learning rate warmup. 

Conclusion

GPT-OSS 20B brings MoE pretraining into the MLPerf Training benchmark suite. By identifying and neutralizing the sources of training variance—specifically validation shuffling, optimizer instability, and initialization inconsistency—the task force has delivered a stable, high-fidelity benchmark that remains accessible to all submitters. This ensures that MLPerf Training v6.0 scores reflect genuine hardware and software efficiency.

GPT-OSS 20B provides the community with a standardized way to evaluate sparse pretraining performance alongside the suite’s existing dense workloads. The reference implementation is available on the MLCommons GitHub repository.

For additional information on MLCommons and details on becoming a member, please visit MLCommons.org or email participation@mlcommons.org.

The post GPT-OSS 20B: A Sparse MoE Pretraining Benchmark for MLPerf Training v6.0 appeared first on MLCommons.

]]>
DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0 https://mlcommons.org/2026/05/deepseek-v3-training-v6-0/ Tue, 05 May 2026 13:37:00 +0000 https://mlcommons.org/?p=3961 How MLCommons is bringing large-scale Mixture-of-Experts (MoE) pretraining to the MLPerf Training v6.0 suite.

The post DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0 appeared first on MLCommons.

]]>
Motivation and Architectural Relevance

As Large Language Model (LLM) development increasingly adopts sparse computation, the benchmarks used to evaluate training performance need to keep pace. MLPerf™ Training v6.0 adds a large-scale pretraining benchmark built on DeepSeek-V3, a Mixture-of-Experts (MoE) architecture with 671B total parameters, of which 37B are activated per token. 

This benchmark captures the performance of critical innovations now standard in the industry, including Multi-head Latent Attention (MLA) and auxiliary-lossfree load balancing.

Technical Architecture & Implementation

DeepSeek-V3 introduces specific computational patterns that differentiate it from the dense models (e.g., Llama 3.1) currently in the suite:

  • Multi-head Latent Attention (MLA): Unlike standard Multi-Head Attention (MHA), MLA uses low-rank joint compression for Key-Value (KV) caches to reduce memory bandwidth bottlenecks during training and inference.
  • Fine-Grained Expert Segmentation: Each expert feed-forward network (FFN) is segmented into $m$ smaller experts. While a typical MoE might use a top-2 routing over 16 experts, DeepSeek-V3 expands this to 160 routed experts plus shared experts to capture common knowledge.
  • Multi-Token Prediction (MTP): The model is trained with a 2-token prediction objective. This requires a shared trunk and two dedicated output heads, increasing the compute-to-memory ratio during the backward pass.
  • Auxiliary-Loss-Free Load Balancing: To prevent routing collapse without the performance degradation of heavy auxiliary losses, a bias term is dynamically adjusted for each expert based on real-time load.

Benchmark Definition & Reference Setup

The task is defined as LLM Pretraining using a Mixture-of-Experts objective.

Dataset and Tokenization

  • Dataset: C4 (Colossal Clean Crawled Corpus)
  • Tokenizer: Llama-3 compatible tokenizer (128k vocabulary)
  • Sequence Length: 4,096 tokens

Convergence and Checkpointing

The task force identified that MoE models spend early training time in a state of token imbalance. Since the benchmark captures only a small slice of the full end-to-end training, this token imbalance state lasted for approximately 50% of the benchmarking time, which is not representative of steady-state MoE training. To ensure the benchmark measures steady-state hardware efficiency, the task force adopted a warm-start approach.

Since the benchmark uses a Llama 8B tokenizer instead of the original DeepSeek tokenizer, initializing from a Hugging Face checkpoint leads to significant token imbalance. To address this, the task force fine-tuned the checkpoint for 50 steps, bringing the token-per-expert distribution close to that obtained with the original DeepSeek tokenizer (Fig 1). The resulting checkpoint is in HuggingFace format and is hosted by MLCommons. This methodology ensures more than 98% of the benchmark run occurs in a balanced expert state, reflecting long-term training dynamics.

(Fig 1)

Global Batch Size (GBS) Selection

This benchmark requires a GBS of 15,360 or greater.While internal testing showed that the model maintains computational efficiency at lower batch sizes (e.g., 512, 1k, or 2k), the task force mandated a higher floor for three reasons:

  • Representativeness: The original DeepSeek-V3 pretraining used a batch size scheduling strategy that peaked at 15,360. To remain representative of the real-world, large-scale pretraining described in the paper, the benchmark targets the 15k-18k range.
  • Fairness in Benchmarking: Setting a minimum GBS of 15k ensures a fair and representative playing field for all submitters, preventing “hero runs” on tiny batch sizes that don’t reflect production-scale MoE training.
  • Convergence Scaling: The task force established a square-root scaling rule for the learning rate to maintain stability across these larger batch sizes:
    LR(GBS)=2.4×105×GBS16384LR(GBS) = 2.4 \times 10^{-5} \times \sqrt{\frac{GBS}{16384}}

Engineering Challenges and Validations

Development of the reference implementation (using NVIDIA NeMo Megatron-bridge) highlighted several critical requirements for convergence:

  • Expert Parallelism (EP): For a single layer, routed experts are uniformly deployed across 64 GPUs (8 nodes). Node-limited routing is enforced where each token is sent to a maximum of M=4 nodes.
  • Memory Management: Expert capacity must not be restricted to avoid out-of-memory errors (OOMs) caused by token imbalance during the initial steps after checkpoint loading, ensuring the benchmark remains representative of real-world training conditions. 
  • Target Metric: The benchmark targets a cross-entropy validation loss of 3.6 with a Coefficient of Variation (CV) of 1.5% (Fig 2).

(Fig 2)

Conclusion

This benchmark provides a standardized platform for evaluating the training efficiency of a leading open-weights MoE model at production scale. By defining clear requirements for convergence, batch size, and expert parallelism, the DeepSeek-V3 benchmark ensures that MLPerf Training continues to reflect the contemporary state of AI infrastructure.

The reference implementation is available on the MLCommons GitHub repository, and the task force welcomes submissions and feedback from the community. For additional information on MLCommons and details on becoming a member, please visit MLCommons.org.

The post DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0 appeared first on MLCommons.

]]>
AI Reliability Map: Rules and Circumstances  https://mlcommons.org/2026/04/airr-map/ Wed, 22 Apr 2026 16:09:09 +0000 https://mlcommons.org/?p=3941 A framework for understanding what AI reliability actually requires: consistently following the right behavioral rules - whether facing normal everyday use or an active adversarial attack.

The post AI Reliability Map: Rules and Circumstances  appeared first on MLCommons.

]]>
Regardless of the intended use of an AI system – be it healthcare, banking, energy, etc. – the complex, black-box nature of LLMs means we need to define the behavior we want and then evaluate how reliably the system delivers it. It’s only with an understanding of reliability that we can both manage risk (is 80 percent reliable enough for life-critical systems?) and cost (what amount of customer service mistakes are acceptable to keep if it reduces per-interaction cost?).

This is why reliability measurement of AI software is the focus  – and namesake – of the MLCommons AI Risk and Reliability Working Group.  Significantly higher AI reliability would both grow markets and protect society. We believe that to successfully increase AI reliability across the industry, we need a systematic and thoughtful plan. We then need to effectively and collaboratively implement, iterate on, and sustain that plan over time. 

As with any long journey, our planning should begin with a map – a map we expect to evolve and improve over time. 

We begin our mapping effort by focusing on pre-deployment testing of an AI system’s behavioral reliability. Reliability needs to be addressed over the AI application lifecycle: during development, at deployment, and during operation. Achieving reliability during each of these three stages involves a different focus: development processes, deployment testing, and operation monitoring. Our initial focus is on the testing that gates deployment because it gives us the most concrete opportunity for change. We further focus on the AI system behavior: the responses it gives and the actions it takes. There are existing approaches for managing the reliability of the hardware and conventional software substrate on which it runs.

The essence of AI reliability (AIR) is consistently adhering to behavioral rules across varied circumstances. We introduce the AI Reliability Map to relate these concerns to the essential concept of consistently following rules under varied circumstances as follows:

AI Reliability Map: Rules and Circumstances
Circumstances
Correctness:Obeying rules while following instructionsSecurity:Resisting rule violations caused by malicious actors
RulesFunctionalityTesting neededTesting needed
Data ProtectionsTesting neededTesting needed
Product SafetyTesting neededTesting needed
Frontier SafetyTesting neededTesting needed
Psychosocial Limits  Testing neededTesting needed

The rows express rules the system is intended to obey; the columns express the circumstances under which the system is expected to obey those rules. It is just as important to follow a functionality rule, such as a deployer instruction, when given normal instructions as when being manipulated by a malicious actor, either directly (through prompt hacking) or indirectly (through prompt injections or misinformation). Likewise, a normal instruction might tempt the system to violate a privacy rule just as much as an attack would, if violating that rule would lead to a “better” outcome for the user. This is why we must all follow all rules in all circumstances, from normal use to malicious action.

It is worth noting that AI Security cannot be tested in isolation: it must be tested by attempting to violate a rule, and the system’s behavior may vary substantially depending on which rule is probed; hence, it is tested across all system behaviors.

This AI reliability map helps us shape our definitions, but we need to be much more detailed to take action. Below, we provide additional details on both the categories shown for the rules and the circumstances, with subcategories. The subcategories are intended to address most of the known, salient concerns in pre-deployment testing of commercial systems. The entire matrix is extensible to accommodate missing or new concerns as they emerge or increase in importance or urgency.

It’s worth noting several things about this expansion. First, Functionality encompasses complying with regulatory and deployment requirements, with the former consistently taking precedence over the later. Second, Data Protection encompasses an individual’s expectations of data privacy as well as “discrete information management,” such as the proper use of corporate data and IP. Third, in our map, Frontier Safety encompasses CBRN and offensive cyber. 

The utility of the map in understanding the state of AI testing is shown with the colored cells. The yellow colored cells correspond to the scope addressed by most public capability testing. The green and blue colored cells correspond to the scope addressed by the MLCommons AILuminate Safety and Jailbreak benchmarks to-date. 

Any moderately advanced AI agent with a natural interface, regardless of purpose, is theoretically capable of failing across this entire scope. An AI system for personal finance could, in theory, offer bad financial advice, design a virus upon request, enable a hacker to access confidential corporate information, or deceive a user into an unintended purchase. There will be both specific risks (bad financial advice, unintended purchases) and general risks (viruses, hacking) associated with these AI tools across all verticals. Our challenge, as an industry and a field, is to develop a robustly structured yet constantly evolving approach to deployment testing that covers this map.

This is the work of the MLCommons AI Risk and Reliability Working Group, where industry, academia, and government collaborate to translate frameworks like this into actionable benchmarks, including AILuminate. As an open engineering consortium, MLCommons is uniquely positioned to lead this effort – bringing together the organizations that build AI systems with those that deploy, regulate, and are affected by them. If your organization is working to understand or improve AI reliability, we want to build this with you. Learn more and join the AIRR Working Group at mlcommons.org/working-groups/ai-risk-reliability.

The post AI Reliability Map: Rules and Circumstances  appeared first on MLCommons.

]]>