Scality makes AI Inference Factory available for enterprise-owned AI inference
1st article in the last 90 days, one of 13 articles referencing Scality. Previous coverage: Scality and WEKA expand France partnership with joint customer support agreement (Jul 2026).
Companies mentioned
Best suited for
- Seniority
- EVP / SVP / VP / AVP
- Job function
- Data Center / IT
- Persona
- Platform / Infrastructure Engineer
- Buyer role
- Decision Maker / Budget Holder
- Decision stage
- Which to buy
- Adoption curve
- Early Adopters
- Technology maturity
- Emerging Exploration
- Industry
- Information Technology / Cloud & Infrastructure Software / Cloud Infrastructure, Automation & Platform Engineering / Cloud Management & FinOps
Our classification, not the publisher's statement. Best suited for, not only for.
Scality made AI Inference Factory available, an open-code software stack for deploying and operating AI inference on infrastructure owned by organizations. The update focuses on predictable AI inference costs and keeping model and data handling under enterprise control, rather than relying on cloud-based inference services.
Scality said that as AI moved into production, organizations sought more control over model lifecycles and more data sovereignty. It described hybrid deployments as a common approach, with critical AI processes running on-premises. The company positioned the stack as a supported alternative that avoids assembling and maintaining the full software stack end to end.
The stack brings together validated open-weight models, a disaggregated inference-serving layer that separates prefill from decode, and a control plane for authentication, metering, routing, and scheduling against SLA targets. Scality also included Scality AI Data Infrastructure (ADI) for policy-governed autonomous operations across the AI lifecycle, with ADI serving as high-performance object storage for models, enterprise data, and inference state.
Scality said ADI extends a shared key-value (KV) cache beyond limited GPU high-bandwidth memory by allowing GPUs to preserve and retrieve model context, including for reasoning models and agentic workloads where KV cache can exceed GPU memory. Scality said testing demonstrated results such as a 1.9-second load time for Gemma-3 27B, a 166 ms warm time-to-first-token from a 14K-token context restore, and 97% of network line rate for data transfer between GPUs and storage. “With AI moving into mission-critical production environments, organizations need greater control over where inference runs, how their models are managed and what happens to their data,” said Jérôme Lecat, CEO at Scality. “For 15 years, Scality has built data infrastructure that thousands of customers around the globe rely on to operate 24/7. AI Inference Factory brings that experience to on-premises AI, giving organizations the reliability and sovereignty they need to run critical AI workloads on their own terms.”
Press release, provided by Globe Newswire on behalf of Scality. Read the original.