Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

DeepSeek keeps 18 core scientists behind R1 model

DeepSeek, an AI startup based in Hangzhou, has retained all 18 core scientists behind its R1 model, according to a newly updated technical paper.

The revised document adds 64 pages to the original and highlights contributions from the key research team, as well as many of the R1 project’s 176 contributors.

DeepSeek’s R1 model gained global attention in early 2025 for approaching the performance of top US AI models while being developed at lower cost.

The team composition, tracked through author lists in the technical paper and a peer-reviewed article published in Nature, shows that most core researchers remained with the company, despite strong competition for AI talent in China.

The number of contributors who left DeepSeek fluctuated throughout 2025, but the latest update indicates that the core team has stayed largely intact.

🔗 Source: South China Morning Post

🧠 Food for thought

Implications, context, and why it matters.

DeepSeek retention matters for R1 progress

  • DeepSeek-V3 used 2,048 Nvidia H800s (export-compliant accelerators) for 2.79 million GPU hours (one GPU for one hour) at $5.6 million 1, under comparable budgets 2. Efficiency came from DualPipe communication architecture (overlaps compute and communication across accelerators) 1. FP8 mixed precision (8-bit floating point) 1 needs high-performance computing co-design (jointly tuning model architecture with hardware) 3.
  •  R1 added reinforcement learning (RL) phases including pure RL training, cold-start supervised fine-tuning (a baseline pass on labeled data), and reasoning-oriented RL for step-by-step problem solving 3. It cost about $1 million 4.
  •  Keeping $0.56 per million input tokens versus OpenAI at $15 requires software efficiency gains 5. DeepSeek likely trails US frontier labs by six months 4.

Cloud teams can assess R1 now

  •  DeepSeek-R1 uses the MIT license, a permissive open source license 6. It allows commercial use and model distillation (training a smaller model to mimic a larger one) with no restrictions, unlike proprietary options from OpenAI or Anthropic.
  •  API pricing counts cache hits (reused prompts) at $0.07 per million tokens and cache misses at $0.56 5, yielding 75% savings for repeated prompts or system messages.
  •  The model supports 128K token context windows (how much text the model considers at once) 5 and a 64K token maximum generation length 6. That enables document processing and long code.
  •  Together AI is a model hosting provider 6. It lists serverless inference, dedicated endpoints, and monthly reserved capacity with FP8 quantization (low-precision weights). The page says it is not supported.

Recent DeepSeek developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.