# Media generation

Source: https://www.hyperagent.com/docs/tools/media-generation

> For AI agents: the documentation index is at https://www.hyperagent.com/llms.txt and the complete documentation in one file at https://www.hyperagent.com/llms-full.txt. Any docs page is also served as Markdown by appending `.md` to its URL.

Everything an agent can create as media: images, short video with native audio, speech, transcripts, avatar videos, and maps.

Product shots for a launch, a narrated walkthrough, a transcript of yesterday's call, a map of the route. The media tools generate and edit that work directly in the thread, and everything made here becomes an artifact in the thread and your [Library](https://www.hyperagent.com/docs/concepts/library), ready to reuse or share.

These tools create new media from your request. Finding existing images on the
web is a separate always-available capability, and publishing pages and decks
is [Webpages and slides](https://www.hyperagent.com/docs/tools/webpages-and-slides).

## What an agent can make [#what-an-agent-can-make]

Generate images from a prompt or edit ones you provide, across three models,
resolutions to 4K, and a range of aspect ratios.

Clips of four to eight seconds with native audio, kept consistent shot to
shot with a first frame and reference images.

A single narrated voice or a multi-speaker exchange with a distinct voice
per speaker.

Speaker labels, per-segment timestamps, and detected emotion, so a meeting
or interview becomes searchable and quotable.

A talking-head presenter for a spokesperson clip, explainer, or personalized
message, from HeyGen avatars and voices.

Geocoding, places, directions, distances, an interactive map, and a Street
View image of a real location.

## Media generation tools [#media-generation-tools]

| Media              | Powered by                                        | Capabilities                                                                                                                                                                                | Limits / specs                                                                                                                                                                                                                                   |
| ------------------ | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|  **Images**        | **Nano Banana****Nano Banana Pro****GPT Image 2** | Generates one image or a set in parallel. Edits images you provide by restyling or combining them, or by creating variations for product shots, hero banners, and redraws.                  | Outputs 1K, 2K, or 4K images in square, 16:9, 9:16, and wider or taller aspect ratios. Nano Banana is the fast default; Nano Banana Pro prioritizes fidelity; GPT Image 2 handles crisp text, dense compositions, and complex multi-image edits. |
|  **Video**         | **Google Veo 3.1**                                | Text-to-video, image-to-video, reference-guided generation, video extension, and native audio. A first frame and reference images keep a product, person, or style consistent across clips. | Four-to-eight-second clips at 720p or 1080p, in 16:9 or 9:16. Supports one first frame and up to three reference images.                                                                                                                         |
|  **Audio**         | **Google Gemini Text-to-Speech**                  | Single-speaker narration or multi-speaker dialogue with a distinct voice for each speaker.                                                                                                  | Around 40 dialogue turns or roughly eight minutes per call. Generate longer pieces in parts and stitch them together.                                                                                                                            |
|  **Transcription** | **Google Gemini multimodal transcription**        | Structured transcripts with speaker identification, per-segment timestamps, language detection, and detected emotion.                                                                       | MP3, WAV, AAC, OGG, FLAC, and AIFF files up to 20 MB. M4A is not supported.                                                                                                                                                                      |
|  **Avatar video**  | **HeyGen**                                        | Talking-head spokesperson clips, explainers, and personalized messages using avatars and voices from the connected HeyGen account.                                                          | Runs asynchronously and takes a few minutes.                                                                                                                                                                                                     |
|  **Maps**          | **Google Maps Platform**                          | Geocoding, Places search, directions, distance calculations, interactive maps, Street View, weather, time zones, markers, routes, and street, satellite, or dark map styles.                | Aerial View is available for US addresses. An uncached cinematic 3D view can take one to three hours to render.                                                                                                                                  |

## Media is not versioned [#media-is-not-versioned]

Unlike [webpages and slides](https://www.hyperagent.com/docs/tools/webpages-and-slides), generated images, audio, and video keep no version history. Regenerating doesn't replace earlier generations, it simply adds more artifacts to your thread.

## FAQs [#faqs]

Nano Banana is the default and the right pick for most images: Pro-level quality at Flash speed. Reach for Nano Banana Pro when the image has to be exactly right and you will trade speed for fidelity. Reach for GPT Image 2 when the image carries crisp text, a dense composition, or a complex multi-image edit.

Use the two consistency controls. A first frame image sets the exact opening frame each clip animates from, and up to three reference images guide style and character so the same subject holds together across clips. Feed in the images you already generated and a set of separate clips reads as one coherent piece.

M4A is not a supported format. Transcription accepts MP3, WAV, AAC, OGG, FLAC, and AIFF, up to 20 MB per file. Convert the M4A to one of those first and it will transcribe with speaker labels and timestamps.

Its supporting service isn't available to the account. Images, Video, Audio, Transcribe, and Avatar rely on a platform-configured service; when it's unavailable the tool reads as locked rather than off. You never supply keys for these yourself.

Every artifact lands in the thread that made it and in the Library, searchable across all your threads, so an image, clip, or transcript is easy to find and reuse later.
