Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

SandboxAQ unveils an open dataset for drug discovery

SandboxAQ has introduced SAIR (Structurally Augmented IC50 Repository), a dataset containing protein-ligand pairs with experimental potency data.

The company describes this dataset as the largest of its kind, aimed at helping researchers enhance AI models for drug discovery by improving predictions of protein-ligand binding affinities.

SAIR includes around 5.2 million synthetic 3D structures across more than 1 million protein-ligand systems.

It was developed using Nvidia’s DGX Cloud platform along with SandboxAQ’s AI Large Quantitative Model (LQM) capabilities.

The project reportedly achieved a twofold improvement in GPU utilization through collaboration with Nvidia.

The dataset integrates physics-based modeling and AI to boost reliability and speed in predicting molecular interactions.

🔗 Source: SandboxAQ


🧠 Food for thought

1️⃣ From hundreds to millions: the exponential evolution of binding data resources

The SAIR dataset represents a significant leap in the scale of protein-ligand binding data available to researchers, fundamentally changing what’s possible in computational drug discovery.

In 2006, BindingDB—then considered comprehensive—contained only about 20,000 experimentally determined binding affinities for approximately 11,000 small molecule ligands and 110 protein targets 1.

This new dataset of 5.2 million synthetic structures is roughly 260 times larger than what researchers had access to less than two decades ago, addressing what has long been a critical bottleneck in developing reliable AI models for drug discovery.

The transition from painstakingly collected experimental data to synthetically generated structures at massive scale parallels similar transformations in computer vision and natural language processing, where synthetic data augmentation enabled breakthrough performance.

The integration of binding data with structural repositories has been evolving since at least 2010, when RCSB PDB began incorporating binding constants like IC50 and Ki directly into their structure summaries, showing the field’s consistent movement toward more accessible and comprehensive datasets 2.

2️⃣ AI drug discovery moves from hype to validation with quantifiable benchmarks

The SAIR dataset provides a standardized benchmark that could accelerate progress across the entire AI drug discovery sector, which has been growing rapidly but often lacks comparable performance metrics.

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.