A model learns the features of a person's face, voice or body from data, then generates new frames or audio, matching facial expressions, lip movement and vocal timbre. In the GAN approach a generator competes with a discriminator; diffusion models gradually denoise toward a realistic output. Increasingly, multimodal models generate coherent audio and video.
A deepfake is itself a generative capability; as a phenomenon it drives the development of detection, watermarking and media authenticity verification methods.
Popularization of neural-network-based face-swapping techniques.