Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

US-based Tiiny AI unveils world’s smallest AI supercomputer

Tiiny AI, a US-based AI startup, has launched the Tiiny AI Pocket Lab, which it claims is the world’s smallest personal AI supercomputer, verified by Guinness World Records for its size category.

The device, first unveiled in Hong Kong, is designed to run large language models with up to 120 billion parameters entirely offline and on-device, without relying on cloud servers or high-end GPUs.

Tiiny AI Pocket Lab is built for energy efficiency, operating at 65W and using an ARMv9.2 12-core CPU, with 80GB memory and 1TB SSD storage.

The company says the device supports several major open-source AI models and agent frameworks, and can install updates offline.

According to Tiiny AI, the Pocket Lab targets developers, researchers, and professionals seeking privacy and local processing for advanced AI tasks.

The company was founded in 2024 and recently raised a multi-million-dollar seed round.

🔗 Source: Tiiny AI

🧠 Food for thought

Implications, context, and why it matters.

ARM offline Large Language Model (LLM) inference lacks verified performance data

  • Tiiny AI says Pocket Lab runs up to 120 billion parameter models on-device at 65W using a system-on-chip (SoC) with a 12-core ARMv9.2 CPU and a dedicated neural processing unit (dNPU) 1. The launch lacks independent tokens-per-second results across model sizes and quantization levels 1.
  • llama.cpp and llamafile (open-source CPU-centric LLM runtimes) find wide swings on ARM CPUs, with Raspberry Pi 5 (ARMv8.2) seeing about a 10x FP16 boost while an Intel Core i9-14900K finished one task about 7x faster 2.
  • There is no third-party data for speed or latency or sustained thermals, which makes real ability hard to judge 1.
  • LLM spending could hit $35.4 billion by 2030, so buyers will expect transparent metrics aligned with MLPerf Tiny v1.3 (an industry benchmark suite for low-power ML) which lists 70 results across five tests for tiny neural networks 13.

ARM-optimized LLM runtimes can target the offline AI appliance market

  • Teams building ARM-friendly runtimes like llama.cpp can benefit, since ARM Neoverse-based AWS Graviton instances (Amazon’s ARM servers) deliver up to 4x better tokens-per-dollar than x86 options 4.
  • Support for ARM-specific optimizations can set products apart. llamafile maintainer Justine Tunney says new matrix-multiplication kernels deliver 30% to 500% prompt-eval speedups on CPUs with strong gains on ARMv8.2 using dot-product instructions and FP16 2.
  • Vendors shipping quantized model packs in GGUF (a binary file format used by llama.cpp for quantized LLMs) can ride demand, as Microsoft’s Phi-3-mini suggests that careful data curation yields smartphone-sized models that compare well with GPT-3.5 on standard reasoning benchmarks 56.
  • Integration shops that deploy local AI with tools like ClearML (an open-source MLOps platform that adds orchestration, autoscaling, and security) can win business as organizations seek privacy with lower costs over cloud reliance 4.

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.