Tired of ads? Enjoy an ad-free experience by signing up.
Shadine Taufik · · 3 min read

Unleash the DragGAN

In art history, perspective took a long time to master for painters.

Across prehistory, antiquity, and the Middle Ages, scenes were formed using mere guesswork, ignoring conventions of depth and space. It wasn’t until the 1410s – during the Italian Renaissance when architect Filippo Brunelleschi painted the Florentine streets – when linear perspective was introduced into the mainstream artworld.

This is an integral part of artistic composition, with painters and photographers studying methods to obtain the most accurate, aesthetically pleasing results.

However, with DragGAN, even the most inexperienced of rookies can now articulate their vision through spatial manipulation.

Demo credit: DragGAN / Compiled by @_akhaliq on Twitter

Created by researchers from the Max Planck Institute for Informatics, Google, MIT CSAIL, and Saarbrücken Research Center, the program allows users to shift objects in photos and illustrations to warp perspectives.

This is done by plotting “handle” (red) points as well as “target” (blue) points on a 2D image and simply dragging on the handle pointsto spatially morph the image.

DragGAN supports a range of subject matter, including landscapes, humans, animals, and cars.

For portrait editing, it also allows users to change facial expressions, haircuts, poses, and lighting. Additionally, users can isolate areas in the picture using a masking tool, leaving the rest undisturbed.

As hinted at in its name, the tech uses generative adversarial networks (GANs), which can be used to generate realistic images, text, music, videos, and 3D objects from scarce data.

GANs work by pitting two neural networks – a “generator” and a “discriminator” – against each other. The generator is trained on purely real data and taught to generate new data similar to what it has learned. On the other hand, the discriminator is fed both real and generated data, and is taught to differentiate between the two. As both carry out their tasks, they become more sophisticated, leading to the generator creating output that is computationally indistinguishable from real data.

Photo credit: DragGAN

Initially introduced by Goodfellow et. al in 2014, the algorithm has since been used to power popular applications such as text-to-image generators DALL-E 2 and Midjourney. With this project, Google has yet another notch on its belt in its AI race against Microsoft or even Adobe, with its new Firefly offering.


Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

TIA Writer

Shadine Taufik

Fan of all things AI and art.