Inference is the execution of a trained artificial intelligence or machine learning model on new data to produce outputs such as predictions or classifications, used in enterprises to support production applications, automated decisions, and analytics within governed, monitored environments.
Alluxio introduced a distributed caching solution for AI workloads on Oracle Cloud Infrastructure, citing sub-millisecond access and up to 1 TB/s throughput.
Rackspace Technology and AMD signed an MOU outlining a multiyear partnership to build a governed Enterprise AI Cloud for regulated and sovereign workloads. The proposal integrates AMD Instinct GPUs and EPYC CPUs into a fully managed stack, covering private/hybrid deployment, inference runtime, and managed inference services with defined SLAs.
Aviz Podcast Episode 2 discusses AI factories and AI fabrics, arguing networks must support multiple fabrics for training and inference with cost models like token serving.
Hedgehog contributed OCP Accepted AI training and inference fabric reference architectures to the OCP Marketplace, available for deployment with Ethernet-based disaggregated networks.