🧔♂️ A friendly human may check it before it goes live. More news here
DeepSeek founder proposes new AI training method to bypass GPU limits
Researchers from DeepSeek, a Chinese AI startup based in Hangzhou, and Peking University have published a technical paper outlining a new training method for large AI models that aims to work around GPU memory limitations.
The paper, released on January 13, 2026 introduces a “conditional memory” system called Engram, which the authors say allows AI models to expand parameters more efficiently by separating compute and memory processes.
The researchers tested Engram on a 27 billion-parameter model and reported improved performance on industry benchmarks, as well as better handling of long input sequences, a common challenge for advanced AI systems.
The team includes DeepSeek founder Liang Wenfeng and Peking University assistant professor Huishuai Zhang.
The industry is watching closely as DeepSeek is rumored to be preparing a new model launch in February, after its previous R1 and V3 models.
🔗 Source: South China Morning Post
🧠 Food for thought
Implications, context, and why it matters.
Engram claims need independent verification before judging impact
- Independent reviews say Engram stores vector embeddings for common n-grams with lookup tables, not neural computation, using a Zipfian distribution (a heavy-tailed pattern where a few items dominate) 1
- DeepSeek describes a U-shaped scaling law where devoting 20–25% of sparse parameters to Engram memory beats pure MoE experts (specialized sub-networks) across models with 5 billion to 27 billion parameters 1
- The Engram GitHub hosts a demo that focuses on core logic and mocks standard parts 2. Gaps include video RAM (VRAM) and bandwidth savings. Hardware examples such as Nvidia A800 or H800 GPUs or Huawei Ascend AI accelerators are unlisted, and direct tests against gradient checkpointing (recomputing activations to save memory) or CPU offloading (moving tensors to CPU memory) are missing
- Without third-party benchmarks, resource measurements, or production deployments, claims that Engram fixes China’s hardware limits remain unproven
CXL memory expansion vendors have near-term partnership openings if Engram shifts demand
- If Engram decouples memory from compute in practice, demand could move from scarce High Bandwidth Memory (HBM) to disaggregated designs, and Compute Express Link (CXL) Type 3 expansion now runs in production data centers 3
- Analysts peg the 2026 CXL Type 3 market at USD 1.8–2.5 billion 3. Module suppliers include Samsung, SK Hynix, and Micron 4. Switch makers are XConn and Astera Labs 4. Marvell provides controller silicon 4
- SMART Modular’s CXL add-in cards enable up to 4 TB of extra memory 5. With 64 GB Registered DIMMs (RDIMMs), systems reach 1 TB per CPU, and CXL attach rates could hit 30% by 2026 5
- Cloud providers and OEMs (original equipment manufacturers) testing memory-efficient AI designs should map which CXL vendors ship production gear to prepare for partnerships if Engram-style methods gain traction
Recent DeepSeek developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




