Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

AWS, Cerebras partner for 10x faster AI inference

AWS and Cerebras said they will launch a disaggregated inference solution on Amazon Bedrock that pairs Trainium with Cerebras CS-3 over the Elastic Fabric Adapter to speed generative AI and LLM workloads.

The setup splits inference into parallel prefill and serial decode using Cerebras CS-3 and Trainium to reduce latency.

The systems will be built on the AWS Nitro System in AWS data centers and be offered on Bedrock, with open-source LLMs and Amazon Nova planned later this year, the companies said.

The solution, featuring open-source LLMs, will be available on Bedrock in the coming months.

🔗 Source: Amazon

🧠 Food for thought

Implications, context, and why it matters.

The IPO-bound company behind the AWS deal

  • This AWS tie-up goes beyond tech work. It sets Cerebras up for its planned IPO 1.
  • Cerebras landed a $1 billion funding round at a $23 billion valuation, plus a reported $10 billion supply deal with OpenAI. The AWS integration broadens distribution and lowers perceived investor risk 1.
  • The build aims at agentic AI workloads, including coding assistants, that can generate about 15x more tokens per query than conversational chat 2.
  • That pushes pressure onto the decode phase, when a model emits output tokens one by one after an initial prefill step. Decode becomes the main choke point and a cost sink for customers 3.

A system-level arms race changes how AI is bought and sold

  • The deal fits a shift away from single-vendor stacks such as Nvidia’s, toward a system-level arms race where cloud providers pull the pieces together 3.
  • AWS pairs its Trainium chips with Cerebras hardware to deliver a tuned setup. The choice cuts reliance on general-purpose GPUs 3.
  • The speed gains come with extra decisions. Teams may pick aggregated systems or a disaggregated design that splits prefill and decode across different hardware, depending on workload behavior 2.
  • Pricing also gets harder to track. Faster inference will sit on top of Amazon Bedrock’s multi-part model, which bills for model access plus options such as safety guardrails 4.

Recent AWS developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.