The cost of creating another image, video, or audio asset is approaching zero, while storytelling and creativity still require human input.
2
Advertising and e-commerce are early areas where generative media can support personalized, interactive content and virtual try-on experiences.
3
Video generation is growing quickly and could become 100 to 250 times larger than the image-generation market as models become cheaper and real time.
Summary
Gorkem Yurtseven traces generative media from DALL-E 2 and Midjourney through Stable Diffusion, Flux, Sora, and newer video models. He argues that the field became accessible quickly after OpenAI's early lead, as open models and new services let more people create and deploy generative media. The cost of producing another asset is approaching zero, although creativity and storytelling still matter. Yurtseven expects advertising and e-commerce to adopt the technology early, with personalized ads, interactive campaigns, and virtual try-on. He sees video as the largest future market. Video models already account for a growing share of FAL's platform usage despite their cost and limitations, and he estimates the generative video market could become 100 to 250 times larger than image generation. The talk ends with predictions about real-time generated video, interactive media, and more enterprise use of improved image models.
Generative media combines image, audio, and video generation
Yurtseven defines generative media as the generation of video, audio, or images. FAL operates a platform that runs many models through its inference engine and partners with closed-source model providers. He places the current market in a longer history of computer-generated art, including Harold Cohen's drawing systems, computer graphics, GANs, Google's DeepDream, and consumer avatar applications such as Prisma. The current wave differs in capability and reach, even though people have tried to make art with computers for decades.
Open models quickly narrowed the gap after DALL-E 2
Yurtseven remembers seeing Sam Altman's 2022 posts about DALL-E 2 as a startling moment, because the images appeared far ahead of what ordinary users could make. That lead did not last. Midjourney released an early beta as a Discord bot, and Stable Diffusion then open-sourced a model that people could run on home GPUs. Services grew around these models, followed by systems such as SDXL and Flux. He describes this sequence as the playing field evening out very quickly.
The cost of creating another asset is approaching zero
Yurtseven distinguishes the marginal cost of creation from the marginal cost of creativity. Once a creative setup exists, he says, producing the next new asset is approaching zero in cost. Storytelling and creativity remain important, but the production step becomes much cheaper. He expects this to affect social media, advertising, marketing, fashion, film, gaming, and e-commerce, with AI eventually affecting all content in some way.
Advertising can absorb large volumes of personalized content
Yurtseven expects advertising to be one of the first industries affected at large scale. Ads appear constantly, so the industry can use much more content than areas such as feature films. Generative systems could create many versions of an ad for different demographics, generate an ad for a person based on the website they visit, or make an ad interactive. He gives FAL's campaign for A24's Civil War as an example: users submitted a selfie and description, and the system created a green toy soldier with their face for display in Times Square.
Virtual try-on is an early product-market fit in e-commerce
Online shopping is visual, which gives generative media room to add interaction. Yurtseven identifies virtual try-on as one of the clearest early product-market fits he has seen in AI. Retailers and e-commerce companies are adopting it, and startups are being built around the technology. His expectation is that every retailer and e-commerce website could become a potential generative media user as online shopping continues to grow.
Video generation is growing despite higher cost and weaker quality
Yurtseven says Sora created another moment when OpenAI appeared far ahead, but his experience with DALL-E 2 made him expect other researchers to catch up. FAL's platform data illustrates the speed of adoption: video models accounted for barely any usage in October, reached 18 percent in February, and were around 30 percent when he gave the talk. He says this growth happened even though video models were expensive and still did not work as well as needed.
Video could become 100 to 250 times larger than image generation
Yurtseven estimates that video models require about 20 times more compute than image models. He also expects video to be more engaging and useful across more industries. Combining those assumptions, he predicts that the generative video market could become 100 to 250 times larger than the generative image market. He still expects substantial growth in image generation, but believes video will grow faster and eventually become much larger.
Real-time video could make generated media interactive
Yurtseven expects video generation to become faster and cheaper until generating one second of video takes about one second. That would allow systems to stream generated content to users. He sees this changing interaction with generative media, making more experiences interactive and blurring the line between games and movies. He also wonders whether live events in games such as Fortnite could become more lifelike and accessible to people who do not usually play video games.
Improved image editing will bring generative media into enterprise software
Yurtseven says newer image systems, including Flux Kontext and GPT-4o, have improved editing and text rendering. He rejects the idea that image models have reached their ceiling because each new capability opens more industry use cases. As the technology matures, he expects established companies to adopt these systems and incorporate them into enterprise workflows.
"Everything potentially becomes interactive. The line between games and movies gets blurred."14:48
Who should watch
You are building a media, advertising, or e-commerce product and want examples of where generative images and video are already fitting into products.
You need a short history of how image models moved from DALL-E 2 to open models and a growing set of commercial systems.
You are planning for video generation costs, adoption, and real-time interactive experiences rather than treating video as only an image-generation extension.