👩🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔♂️ A friendly human may check it before it goes live. More news here
🧔♂️ A friendly human may check it before it goes live. More news here
Rakuten develops lower-cost AI models for ecommerce
Rakuten Group is expanding its AI team and developing large language models with a focus on reducing costs.
The Tokyo-based ecommerce and tech firm, led by AI chief Ting Cai, now has a 1,000-person AI team and uses thousands of Nvidia chips for its projects.
Cai said the company’s new large language model, version 3, is designed to be 90% cheaper to run than other comparable models, though this claim has not been independently verified.
The company reported that AI contributed ¥10.5 billion (US$67 million) to its operating income in 2024, and said it aims to double this figure in 2025.
🔗 Source: Bloomberg
🧠 Food for thought
Implications, context, and why it matters.
Rakuten links MoE efficiency and internal trials to its 90% cost claim, with no independent validation
- Rakuten AI 3.0 uses a Mixture of Experts (MoE) setup with about 700 billion parameters 1. Only about 40 billion parameters activate per token (a chunk of text), with routing through one shared expert and eight specialized experts (specialized sub-networks) 1. This design lowers compute during inference (when generating outputs) versus dense models.
- Japanese MT-Bench comes in at 8.88, ahead of GPT-4o at 8.67 and Rakuten’s prior model at 6.79 1. The 90% cost cut comes from internal trials tied to Rakuten ecosystem services, with no published details on model size baselines, serving infrastructure (the hardware and software used to deploy a model), or quality at different price points 1.
- Training ran on an in-house multi-node GPU cluster in a secure environment, with all data kept internally 1. The cost edge may stem from owned infrastructure as well as model efficiency. Third-party tests on inference costs, context length capabilities (how much text the model can consider at once), and evaluations across varied tasks would confirm whether savings travel across deployments.
AI vendors targeting Japanese enterprises can serve growing demand for cost-optimized inference
- Japanese enterprises show cautious but growing AI use, with manufacturing at 33.2%, information and communication at 56.3%, and services at 33.5% 2. Accuracy concerns at 44.3% and security worries at 34.9% remain 2. Vendors that deliver inference optimization, serving stacks (software layers to deploy models at scale), and AI FinOps (financial operations practices to monitor and control AI spending) can ease these pain points.
- Per capita Claude usage sits at an Anthropic Usage Index (AUI) of 1.86 in Japan, while leaders like Israel reach about 7x expected usage 3. That gap opens growth as firms look for cheaper options than frontier models. Vendors focused on quantization (reducing numeric precision to shrink models and speed up inference), model compression, and spend tracking for Japanese language models can win early deals.
- About 72% of enterprises prefer purchase-led strategies over building from scratch 4. Companies want structured training at 35.5% and integration with local platforms 2. Turnkey inference optimization with compliance-ready features and strong Japanese language support fits this demand.
Recent Rakuten developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




