🧔♂️ A friendly human may check it before it goes live. More news here
Google rolls out powerful AI chip in challenge to Nvidia
Google will make its new seventh-generation Tensor Processing Unit (TPU), called Ironwood, available for public use in the coming weeks.
The chip was first introduced in April for testing and is built in-house to handle tasks like training large AI models and powering real-time chatbots.
Google says Ironwood can connect up to 9,216 chips in a single pod and is over 4x faster than its predecessor.
The launch is part of Google’s efforts to compete with cloud rivals Amazon and Microsoft, as well as chipmaker Nvidia, in the AI infrastructure space.
Anthropic, an AI startup, plans to use up to 1 million Ironwood TPUs for its Claude model, Google said.
Google also announced upgrades to its cloud services aimed at improving cost, speed, and flexibility.
The company reported Q3 cloud revenue of US$15.2 billion, up 34% year-on-year, and raised its annual capital spending forecast to US$93 billion.
🔗 Source: CNBC
🧠 Food for thought
Implications, context, and why it matters.
Ironwood’s value hinges on unconfirmed pricing and real-world performance versus Nvidia
- Google calls Ironwood its fastest TPU with near 2x power efficiency 1. Pricing is undisclosed, with no standard tests against Nvidia H100/H200 data center GPUs (mainstream enterprise AI accelerators).
- Past TPU rates put v5e at $1.20 per chip-hour on demand and v5p at $4.20 2. Buyers need MLPerf (industry-standard AI benchmarks) plus cost-per-token metrics (the cost to process or generate a unit of text) to gauge total cost of ownership (TCO).
- A 9,216-chip pod delivers 42.5 exaflops using floating-point 8-bit (FP8) precision 13. The Next Platform pegs Ironwood (TPU v7p) pods at $52 per TFLOPS to rent versus $21 to build 3, which could curb use outside long deals.
Third-party vendors can fill PyTorch/XLA optimization gaps for TPU migration
- Teams in MLOps (machine learning operations) and systems integrators (IT services firms that assemble and deploy complex stacks) can package migration help.
- The unified vLLM (an open-source large language model inference engine) TPU backend supports PyTorch and JAX models with no code changes 4. Throughput can rise up to 5x 4. This helps inference. Training still needs manual PyTorch/XLA (PyTorch on the Accelerated Linear Algebra compiler) integration 5. PyTorch/XLA docs call out multiprocessing hurdles and recompilation issues with dynamic shapes (varying tensor sizes) 67. Many teams will use cloud and AI consulting during pod-scale deployments (a ‘pod’ is Google’s tightly networked TPU cluster).
Recent Google developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




