Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

DeepSeek launches new open-source model to compress text

DeepSeek has launched DeepSeek-OCR, a new open-source multimodal model designed to compress text inputs using visual processing.

Available on Hugging Face and GitHub, the model uses a vision encoder to reduce the number of tokens needed to process large and complex documents, cutting computing costs for large language models.

The company said its approach allows for significant token reduction, between 7x and 20x, while maintaining information accuracy.

DeepSeek-OCR features two main components: DeepEncoder, which handles compression, and a Mixture-of-Experts decoder with 570 million parameters that reconstructs text.

In benchmark tests, it outperformed models like GOT-OCR 2.0 and MinerU 2.0 while using fewer tokens.

The model can also interpret structured visual content such as tables and formulas, making it suitable for finance and science applications.

🔗 Source: South China Morning Post

🧠 Food for thought

Implications, context, and why it matters.

DeepSeek-OCR (optical character recognition) compression claims need independent validation before enterprises commit

  • DeepSeek claims 7–20× token compression with 97% accuracy at 10× and about 60% at 20×, yet these numbers come from the company. Teams should secure outside tests on task quality and latency, plus cost checks against Anthropic’s Claude, Google’s Gemini or Retrieval-Augmented Generation (RAG) setups.
  • DeepSeek-OCR turns text into images, then back through vision tokens that aim to capture global patterns. Legal reviews need trials on domain accuracy and error carryover. Costs must include encoding overhead. OCR still struggles with complex layouts or handwriting or structure extraction, so results may vary beyond curated benchmark sets.

Software-as-a-Service (SaaS) providers can earn from compression as a service for API-heavy workflows

  • Machine Learning Operations (MLOps) and developer tooling vendors can sell compression as a service that pre-encodes long files. Teams should map context pricing from OpenAI, Anthropic’s Claude 200K, and Google’s Gemini to size 7–10× savings.
  • The model can produce over 200,000 pages of training data per day on one Nvidia A100-40G Graphics Processing Unit (GPU). This creates room for infrastructure firms to offer synthetic datasets in finance and science where DeepSeek-OCR parses tables, formulas, and diagrams. Vendors that handle scanned records can plug in visual compression to cut token use in LLM flows, but ROI must weigh two-stage encoding and decoding latency against direct text runs.

Recent DeepSeek developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.