# Veo 3 for Developers

Paige Bailey, Google DeepMind | AI Engineer World's Fair 2025 | 20:37

Source: https://www.youtube.com/watch?v=hlcAZ2lX_ZI
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/veo-3-for-developers
Published: 2025-06-21
Tags: developer-experience, image-generation, multimodal, video

## TL;DR
- Veo 3 generates video and synchronized audio natively from text and image prompts, including dialogue, music, sound effects, and background noise.
- Veo 3 improves visual and stylistic consistency across a sequence, while natural-language prompts provide control over characters, scenes, and camera movement.
- Developers can access Veo 3 through Google AI plans and Vertex AI private preview, with a short code sample for submitting jobs and storing generated video.

## Summary
Paige Bailey introduces Veo 3 alongside Imagen 4 for still images and Lyria 2 for music. She briefly reviews Veo 2 capabilities such as reference-powered video, style matching, camera controls, outpainting, object removal, character control, and first-to-last-frame interpolation. Veo 3 adds native audio generation, so dialogue, music, sound effects, and background noise are composed with the video rather than added afterward. Bailey focuses on improved prompt adherence, visual consistency, and contextual consistency, while also describing SynthID and visible watermarks. She compares recent text-to-video examples to show how quickly the quality has changed. A commercial recreation demonstrates the difference between a multi-step Veo 2 workflow and a one-prompt Veo 3 workflow. Developers can try Veo 3 through Google AI subscriptions or Vertex AI private preview, with Gemini API access and broader availability still developing.

## Key ideas
### Veo 2 already gives developers detailed control over generated video
[01:48](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=108s)
Before presenting Veo 3, Bailey reviews recent Veo 2 additions. Developers can use reference images to place a person in an environment or preserve a visual style. Natural-language camera controls include moving back, moving right, rotating up, and zooming in. Other API features include outpainting a partial scene, adding or removing objects, controlling a character's motion, and interpolating between first and last frames. Character control can use a script and voice tone, with lip movements mapped to the generated speech. Bailey says these capabilities are available through the API and tools such as Flow.

### Veo 3 composes video and audio together in one generation model
[06:16](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=376s)
Bailey describes Veo 3 as video coupled with audio. Audio is not pulled in later as a separate tool. The model composes tokens across multiple modalities natively. Its outputs can include dialogue, music, sound effects, and subtle background noises. She connects this design to Gemini's native audio output, where one system can work across text, code, images, and audio. The examples include a small llama and other generated scenes, with the audio carrying more than simple background noise.

### Veo 3 improves consistency when scenes contain movement and detail
[07:47](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=467s)
Video generation has historically struggled with prompt nuance and continuity. Characters can change from one frame to another, and backgrounds can break, with walls disappearing or revealing inconsistent spaces. Bailey says Veo 3 improves both stylistic consistency and contextual consistency. The model is built on years of research, including work she names as GQN and WALT. She also points to responsibility features, including visible watermarks and SynthID watermarks for synthetic images and video.

### Google's generative media models cover video, images, and music
[08:48](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=528s)
The talk also introduces Imagen 4 and Lyria 2. Imagen 4 generates still images while preserving realism, detail, varied styles, and typography. Bailey shows examples involving people, animals, local San Francisco references, and artist collaborations. Lyria 2 produces high-fidelity music and professional-grade audio with controls for tones and styles. Music AI Sandbox provides a visual music workflow, while MusicFX can compose beats from natural-language instructions. Bailey describes collaborations with artists and musicians, including Darren Aronofsky, Ross Lovegrove, Jacob Collier, and Toro y Moi.

### Recent model examples show how quickly text-to-video quality has changed
[12:08](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=728s)
Bailey compares a single prompt across several generations of video systems. The prompt asks for a raccoon wearing a black jacket and dancing in slow motion in front of the pyramids. She shows an early 2023 result, WALT, LTX Video from 2024, Kling 2.0 from 2025, Veo 2, and then Veo 3. The earlier examples are short or choppy and have weaker motion. The Veo 3 version produces a more stylish raccoon scene. She then shows image-to-video examples, where still images are turned into moving scenes through natural-language instructions.

### Prompt writing in Veo 3 expands a simple idea into a fuller instruction
[13:43](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=823s)
Veo 3 includes prompt writing that can take a short description and turn it into something more fully formed for the model to understand. Bailey returns to the raccoon example, starting with the simple sentence about dancing in front of the pyramids. She pairs this with Veo 3's sound generation, which can produce music, sound effects, and background noise. The generated examples are intended to preserve fine visual detail while following the expanded description and its requested audio.

### Veo 3 is available through Google AI plans and Vertex AI preview
[16:25](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=985s)
Bailey lists several access paths. Veo 3 is available through the Google AI Ultra plan and, for Google AI Pro subscribers, through the Gemini mobile app with a limited number of uses. It is also available in private preview through Vertex AI. Veo 2 is available through Vertex AI for the Gemini APIs, and Bailey says the team hopes to bring the models to AI Studio. She shows a short code sample with an output bucket, an optional input image, aspect-ratio settings, and model selection. The example is intended to show that basic integration takes a small amount of code.

### A Veo 3 commercial recreation removes several manual production steps
[18:34](https://www.youtube.com/watch?v=hlcAZ2lX_ZI&t=1114s)
Bailey recreates a Chick-fil-A chicken sandwich commercial to compare workflows. With Veo 2, she supplied the original video to Gemini, generated a detailed plan, divided the commercial into prompts because of the eight-second limitation, created a background track with MusicFX, and stitched the pieces together in Camtasia with transitions. With Veo 3, she supplied the original video, generated a description, and gave that description to the model. One prompt produced a recreation with the spoken line, visual performance, and audio together. Bailey says the Veo 3 result was produced from that single submission.

## Notable quotes
- "Video has the potential to be an incredible surface for communication, but also for education and for human creativity." (01:28)
- "V3 is video but coupled together with audio." (06:16)
- "The model actually able to compose together all of these tokens across multiple modalities." (06:16)
- "So again, one prompt and was able to produce this." (19:39)
- "That is not me actually waving at the camera. That is a static photo of me that has been Veo 3 animated." (20:24)

## Tools & references mentioned
- Veo 3
- Veo 2
- Imagen 4
- Lyria 2
- Gemini
- Vertex AI
- Flow
- Gemini APIs
- AI Studio
- Google AI Ultra
- Google AI Pro
- Music AI Sandbox
- MusicFX
- SynthID
- GQN
- WALT
- Runway Gen-4
- Kling 2.0
- LTX Video
- Camtasia
- Andre Karpathy
- Darren Aronofsky
- Ross Lovegrove
- Jacob Collier
- Toro y Moi

## Who should watch
- You are building a creative, education, advertising, or game workflow and want to understand what synchronized video and audio generation can replace or simplify.
- You need to test Veo models through Google APIs and want the current access paths, prompt shape, and basic integration pattern.
- You are comparing video-generation systems and want concrete examples of consistency, camera control, image-to-video generation, and audio production.

## Related talks

- [Prompt to Pipeline: Building with Google's Gen Media Stack](https://aietalks.com/talks/prompt-to-pipeline-building-with-googles-gen-media-stack) (Paige Bailey, Guillaume Vernade & Ian Valentine, Google DeepMind, 1:54:35)
- [See, Hear, Speak, Draw](https://aietalks.com/talks/see-hear-speak-draw) (Logan Kilpatrick & Simón Fishman, OpenAI, 18:43)
- [Voice agents with Realtime Video](https://aietalks.com/talks/voice-agents-with-realtime-video) (Sidney Primas, LemonSlice, 26:36)
- [Building an Agentic Video Editor for Mass Consumer](https://aietalks.com/talks/building-an-agentic-video-editor-for-mass-consumer) (Ekaterina Deyneka, Reelful, 12:45)
- [Generative Video at the Speed of Light](https://aietalks.com/talks/generative-video-at-the-speed-of-light) (Keegan McCallum, uRun, 08:43)
