Media generation
Everything an agent can create as media: images, short video with native audio, speech, transcripts, avatar videos, and maps.
Product shots for a launch, a narrated walkthrough, a transcript of yesterday's call, a map of the route. The media tools generate and edit that work directly in the thread, and everything made here becomes an artifact in the thread and your Library, ready to reuse or share.
Generated, not found
These tools create new media from your request. Finding existing images on the web is a separate always-available capability, and publishing pages and decks is Webpages and slides.
What an agent can make
Images, new or edited
Generate images from a prompt or edit ones you provide, across three models, resolutions to 4K, and a range of aspect ratios.
Short video with audio
Clips of four to eight seconds with native audio, kept consistent shot to shot with a first frame and reference images.
Speech and dialogue
A single narrated voice or a multi-speaker exchange with a distinct voice per speaker.
A transcript of a recording
Speaker labels, per-segment timestamps, and detected emotion, so a meeting or interview becomes searchable and quotable.
An avatar reading a script
A talking-head presenter for a spokesperson clip, explainer, or personalized message, from HeyGen avatars and voices.
Maps and street views
Geocoding, places, directions, distances, an interactive map, and a Street View image of a real location.
Media generation tools
| Media | Powered by | Capabilities | Limits / specs |
|---|---|---|---|
| Images | Nano Banana Nano Banana Pro GPT Image 2 | Generates one image or a set in parallel. Edits images you provide by restyling or combining them, or by creating variations for product shots, hero banners, and redraws. | Outputs 1K, 2K, or 4K images in square, 16:9, 9:16, and wider or taller aspect ratios. Nano Banana is the fast default; Nano Banana Pro prioritizes fidelity; GPT Image 2 handles crisp text, dense compositions, and complex multi-image edits. |
| Video | Google Veo 3.1 | Text-to-video, image-to-video, reference-guided generation, video extension, and native audio. A first frame and reference images keep a product, person, or style consistent across clips. | Four-to-eight-second clips at 720p or 1080p, in 16:9 or 9:16. Supports one first frame and up to three reference images. |
| Audio | Google Gemini Text-to-Speech | Single-speaker narration or multi-speaker dialogue with a distinct voice for each speaker. | Around 40 dialogue turns or roughly eight minutes per call. Generate longer pieces in parts and stitch them together. |
| Transcription | Google Gemini multimodal transcription | Structured transcripts with speaker identification, per-segment timestamps, language detection, and detected emotion. | MP3, WAV, AAC, OGG, FLAC, and AIFF files up to 20 MB. M4A is not supported. |
| Avatar video | HeyGen | Talking-head spokesperson clips, explainers, and personalized messages using avatars and voices from the connected HeyGen account. | Runs asynchronously and takes a few minutes. |
| Maps | Google Maps Platform | Geocoding, Places search, directions, distance calculations, interactive maps, Street View, weather, time zones, markers, routes, and street, satellite, or dark map styles. | Aerial View is available for US addresses. An uncached cinematic 3D view can take one to three hours to render. |
Media is not versioned
Unlike webpages and slides, generated images, audio, and video keep no version history. Regenerating doesn't replace earlier generations, it simply adds more artifacts to your thread.