Fun with AI

Image transformers are pretrained on datasets with (at a minimum) millions of images, and in practice likely billions of images. That's what Google does with all of the images stored in (now-defunct) Picasa and stored on Google Drive. That's what Meta does with all of the images posted on Facebook.

There are several approaches to pretraining. A common one is to have the network learn to transform the image into itself.

Later on you use convolutions and things like Gram Matrices to effectuate the style you want. Basically you use that to tweak the pretrained transformer to make your style transformer.
 
Back