Pack · 8 talks · 5h 05m to watch, 48 min to read

Image generation beyond the prompt

Image generators are impressive until you need the same character twice, a product to keep its shape, or an edit to work for millions of people. The model mechanics come first, from FLUX to ComfyUI's denoising graph, where the pipeline can be inspected and changed. Better control builds on that view: GPT-4 with vision compares an image with the prompt, then revises the next DALL-E 3 attempt, while a Wind in the Willows project carries saved character references through a whole book. Google Photos shows what changes when a creative demo becomes Magic Editor, with narrow jobs, several candidate outputs and maintained benchmarks. Those benchmarks face an uncomfortable test in the final panel: people often reward the sharper, more saturated image even when it is less realistic or useful.

4
Yedri Kosinski, Comfy · 51:25 · AI Engineer World's Fair 2025
ComfyUI Full Workshop

Why here: Batifol presents generation and editing as model capabilities. Kosinski exposes the pipeline behind them as nodes for the model, text encoder, VAE, sampler and decoder. Once the workflow is visible, a person can replace or patch one part instead of hoping a better prompt fixes everything.

5
Logan Kilpatrick & Simón Fishman, OpenAI · 18:43 · AI Engineer Summit 2023
See, Hear, Speak, Draw

Why here: Kosinski gives a person direct control over the generation graph. Kilpatrick and Fishman hand part of the review back to the models: describe a target, generate an image, compare the two, then rewrite the prompt from the differences. Vernade turns that small loop into a longer creative job.

6
Guillaume Vernade, Google DeepMind · 1:17:14 · AI Engineer Europe 2026
Let's Go Bananas with GenMedia

Why here: Kilpatrick and Fishman test one image against one target. Vernade has to keep characters recognizable across a book. He extracts structured prompts from the source, saves character portraits and passes only the relevant references into each illustration. Ma takes that need for control into a consumer product.

7
Kelvin Ma, Google Photos · 20:28 · AI Engineer World's Fair 2025
Google Photos Magic Editor: GenAI Under the Hood of a Billion-User App

Why here: Vernade can tolerate retries and imperfect illustrations in a workshop. Ma cannot make that bargain inside Google Photos. Magic Editor narrows a broad promise into jobs such as moving an object or rebuilding a background, then uses constraints, evals and several candidate outputs to make those jobs dependable.

8
Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind · 56:59
SOTA Generative Media Panel

Why last: Ma argues that reliable editing needs benchmarks and constrained workflows. Brichtova, Erhan and Gu show why the benchmark itself deserves suspicion. People may choose the sharper, more saturated image even when it is less realistic, while real production workflows reveal mistakes that isolated samples never expose.

After this pack: Video generation →