🧔♂️ A friendly human may check it before it goes live. More news here
Arm, Meta team up to improve AI efficiency on cloud
Arm and Meta have announced a multi-year partnership to improve the efficiency of AI systems across cloud and edge devices.
The collaboration will see Meta’s AI software, including the PyTorch framework, optimized for Arm’s chip architectures to boost performance and reduce power usage.
Meta will use Arm’s Neoverse-based data center platforms for AI ranking and recommendation systems, which will help lower power consumption compared to x86 systems.
The partnership also covers joint work on open source tools like Facebook General Matrix Multiplication (FBGEMM) and PyTorch.
Their efforts focus on improving AI workloads for Meta’s platforms such as Facebook and Instagram, while contributing improvements to the open source community.
No financial details or specific performance targets were disclosed.
🔗 Source: Arm
🧠 Food for thought
Implications, context, and why it matters.
Meta moves ranking to Arm CPUs, leaves large language model (LLM) training on GPUs
- Meta is moving ranking and recommendation to Arm Neoverse CPUs (Arm’s server-class processor cores), not LLM training 1. This targets CPU inference for Facebook and Instagram, where power efficiency drives cost at 3+ billion users.
- Deployment runs on Nvidia Grace CPU systems (Nvidia Arm-based data center CPU), not custom Arm silicon 2. Each GB200 or GB300 NVL72 rack (rack-scale AI system that pairs Grace CPUs with Nvidia GPUs) holds 36 Neoverse-V2 CPUs 2. Meta tunes PyTorch plus Facebook General Matrix Multiplication (FBGEMM) for Arm vector extensions, moving from Intel and AMD AVX2 or AVX-512 2.
Vendors package Arm stacks as Meta code ships
- Build migration kits with PyTorch and ExecuTorch (Meta’s edge-inference runtime). Add KleidiAI (Arm’s optimized kernels and libraries for AI inference) for AWS Graviton, Google Axion, Microsoft Cobalt 2. Meta will open-source FBGEMM and PyTorch upgrades 1.
- Offer quantized large language model (LLM) deployment for cost aware enterprises. ExecuTorch with KleidiAI reaches 350+ tokens per second prefill and 40+ tokens per second decode on Arm devices 34. Gains up to 20% over non KleidiAI builds. Target teams running recommender engines, chatbots, or on-device AI.
- Publish Arm benchmarks and Docker containers before adoption firms up. As KleidiAI integrates with ONNX Runtime (an open-source engine for running AI models) and llama.cpp (an open-source LLM inference project optimized for CPUs) 5, early movers can win pilots from enterprises testing Arm for savings.
Recent Arm developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




