Skip to main content

MLCommons Adds End-to-End RAG and Edge Agentic Tests in MLPerf Inference v6.1

3 companies named across 4 categories, one of 12 articles referencing MLCommons. Previous coverage: MLCommons releases MLPerf Storage v3.0 benchmark results (Sep 2026).

Companies mentioned

Best suited for

Seniority
Analyst
Job function
Chief Information Officer
Buyer role
Decision Maker / Budget Holder
Buyer journey
Need to Buy
Adoption curve
Early Majority
Technology maturity
Operational Expansion
Industry
Information Technology / Software & Services / IT Services / Internet Services & Infrastructure

Our classification, not the publisher's statement. Best suited for, not only for.

MLCommons reported results for the MLPerf Inference v6.1 benchmark suite, adding two new tests for inference deployment scenarios that involve multi-step workflows and agentic behavior. The update matters to organizations evaluating inference performance across different system configurations, with published results covering datacenter and edge benchmarks.

MLPerf Inference benchmarks measure system performance in an architecture-neutral, representative, and reproducible manner, and the v6.1 release included submissions from 30 participating organizations. The suite also incorporated peer-reviewed performance results for several recently released or soon-to-be-released AI platforms, and the release described compound performance improvements for specific benchmark scenarios.

MLPerf Inference v6.1 introduced the End-to-End Retrieval-Augmented Generation (RAG) benchmark and an Edge Agentic Inference benchmark. The End-to-End RAG workload used an embedding model to convert queries into vectors, a retriever to pull candidate passages, a re-ranker to refine passages, and one or more LLMs to reason over the data to produce answers, including document-corpus ingestion and query-answering against a pre-built vector database. The Edge Agentic Inference benchmark targeted multi-turn agentic workloads using an edge model and quantization, with a single-stream coding workload, latency metrics, and a statistically robust accuracy gate.

MLPerf Inference 6.1 also added support for speculative decoding in the interactive scenario for two inference benchmarks and added support for a GPT-OSS task. In the results round, MLCommons listed five new processors or accelerators, including AMD Ryzen AI Max+ 395, AMD Instinct MI350P, and Intel Arc Pro B70 available at the time of the release, plus NVIDIA Rubin and NVIDIA Vera Rubin NVL72 in preview. The report also described the largest system submitted, which used 512 accelerators and two heterogeneous systems.

“We are working hard to ensure that the MLPerf Inference benchmark continues to reflect the scenarios that the AI community values most,” said Miro Hodak, MLPerf Inference working group co-chair. “We added the End-to-end RAG test because it’s clear that query-answering has evolved beyond simply an LLM trained on a corpus; stakeholders need to understand the real-world performance of the types of multi-step, multi-component pipelines that are being built today. Likewise, we added the Edge Agentic Inference test because complex inference systems with agentic properties are increasingly hosted on edge computing devices, creating a new set of performance challenges our customers face today. We are committed to providing timely and relevant performance data that reflects real-world production systems and the performance optimizations that are being deployed today.”

Looking ahead, MLCommons said MLPerf Endpoints would replace Inference in its family of datacenter benchmarks, and it characterized the API-centric harness adoption as supporting that transition.

Press release, provided by Globe Newswire on behalf of MLCommons. Read the original.