Image & Video Capabilities
AdaL ships with full image and video capabilities built in — generation, editing, and analysis — available across all agent modes with no additional setup. Use them inline during any session: generate product shots, edit existing visuals, animate images into video clips, or analyze screenshots, all without leaving your workflow.
Image Generation
Ask AdaL to create images from a text description:
Generate a product shot of a wireless keyboard on a clean white desk, soft studio lighting.
Create a minimalist logo for a coffee shop with earthy tones.
Draw a futuristic city skyline at golden hour, wide-angle, cinematic.
Models
AdaL uses three image models depending on the task:
| Model | Slug | Best for |
|---|---|---|
| Nano-Banana-2 | nano-banana-2 | Default — best all-around, strong instruction following, text+image output |
| Nano-Banana-Pro | nano-banana-pro | Professional assets, text rendering, 4K resolution |
| GPT-Image-2 | gpt-image-2 | High-fidelity photorealistic, text in images (OpenAI) |
Nano-Banana models are powered by Google's Gemini image stack (gemini-3.1-flash-image, gemini-3-pro-image). GPT-Image-2 is OpenAI's model. Specify your preference in the prompt or let AdaL pick automatically.
Variants
Generate multiple variations from the same prompt:
Generate 4 variants of a logo design for a tech startup.
Create 3 different styles of a mountain landscape.
Files are saved as name_0.png, name_1.png, etc. Up to 4 variants per call.
Aspect Ratios & Resolution
Supported aspect ratios: 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, 21:9
Supported output resolutions: 1K, 2K, 4K
Generate a 16:9 banner image of an ocean wave at 2K resolution.
Create a 9:16 phone wallpaper with abstract geometric art.
Image Editing
AdaL can edit or transform existing images by passing them as input:
Remove the background from product_shot.png and place it on a white background.
Change the jacket color in portrait.jpg from black to deep navy blue.
Add a subtle lens flare to the sky in landscape.png.
You can also composite and merge images:
Combine hero.png and background.jpg — place the hero in the foreground with soft depth of field.
Pass up to multiple reference images to guide style or subject consistency across a generation.
Image Analysis
AdaL can read and analyze existing images — screenshots, diagrams, mockups, or any visual:
Read this screenshot and tell me what's wrong with the layout.
Analyze the architecture diagram in docs/diagram.png.
What does this error screenshot show?
Compare these two UI mockups and highlight the differences.
Paste or drop images directly into AdaL for instant analysis.
Vision Models
AdaL routes image analysis to the best vision model for the task. You can specify one in your prompt, or let AdaL pick the default.
| Model | Catalog key | Best for |
|---|---|---|
| Gemini 3.5 Flash | google-gemini-3.5-flash | Default — fast UI checks, simple descriptions |
| MiniMax M3 | minimax-MiniMax-M3 | Budget option — thorough, critical reviews |
| Gemini 3 Flash | google-gemini-3-flash-preview | Very cheap, decent quality |
| Gemini 3.1 Pro | google-gemini-3.1-pro-preview | Best multimodal understanding, large context |
| GPT-5.4 | openai-gpt-5.4 | Strong reasoning, complex UI/UX comparisons |
| Claude Sonnet 4.6 | anthropic-claude-sonnet-4-6 | Detailed text extraction, reading code in screenshots |
Supported Formats
PNG, JPEG, GIF, WEBP, BMP, AVIF, TIFF
Video Generation
Video is a capability, like browser use: you do not switch into a video agent. Ask for a video and AdaL loads the video capability on its own — generation, voice-over, captions, editing and Remotion scene rendering, added to whatever agent you were using. /capabilities shows whether it is on; /capabilities video turns it on ahead of time. Once on, it stays on until the session ends.
AdaL generates cinematic video clips with Google Veo 3.1 by default, or with MiniMax H3 (Hailuo 03) when you ask for it. Video generation runs as a background job — AdaL submits the request and polls for the result automatically.
Choosing a Model
Say which model you want in the request — "use minimax-h3" (or "use Hailuo") — and AdaL passes it to the video tool. Without a model name, Veo is used.
| Veo 3.1 (default) | MiniMax H3 | |
|---|---|---|
| Ask for it with | nothing — it is the default | "use minimax-h3" / "use Hailuo" |
| Resolutions | 720p, 1080p, 4k | 768P, 2K |
| Clip length | 4, 6, or 8 seconds | any whole number from 4 to 15 seconds |
| Aspect ratios | 16:9, 9:16 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 for text-to-video; image-to-video follows the image |
| Reference images | up to 3 | up to 9 |
| Video extension | yes (720p) | no |
| Native audio | yes | yes, stereo |
| Cost | ~$0.40/sec (4K ~$0.60/sec) | $0.08/sec at 768P, $0.13/sec at 2K |
Pick MiniMax H3 for clips longer than 8 seconds, for 2K output, or when cost matters. Pick Veo to extend an existing clip or for 1080p/4K delivery.
Use minimax-h3 to make a 12-second 2K clip of a paper boat drifting across a pond at sunrise, soft ambient birdsong.
Text-to-Video
Generate a video from a description alone:
Create an 8-second cinematic clip of a red sports car driving through a city at night, slow tracking shot, dramatic lighting, ambient street sounds.
Image-to-Video
Animate a still image into motion:
Animate product_shot.png — slow dolly push with gentle ambient light.
AdaL generates the source image first (if needed), then passes it to the video model as the starting frame. Both Veo and MiniMax H3 support this.
Frame Interpolation
Smoothly transition between two images:
Generate a dramatic weather-change video — start from sunny_beach.png and end at stormy_beach.png over 8 seconds.
The model generates a clip that starts at the first image and ends at the second. Great for before/after reveals, day-to-night transitions, and morphing effects. Works with both Veo and MiniMax H3.
Video Extension
Extend an existing Veo clip (Veo only — MiniMax H3 cannot extend videos):
Continue the camera movement from intro_clip.mp4 into the garden for another 8 seconds.
Reference-Image Consistency
Pass reference images to keep a subject consistent across scenes:
Generate a product video of the red sports car from car_ref.png driving through a rainy city street.
Up to 3 reference images per call with Veo, up to 9 with MiniMax H3. Useful for product demos, character consistency, and brand style.
Prompting Tips
- Camera: "slow dolly push", "tracking shot", "aerial pan", "handheld close-up"
- Lighting: "warm golden hour", "moody blue tones", "harsh midday sun"
- Style: "photorealistic", "cinematic film grain", "timelapse", "slow motion"
- Audio: both models generate native audio — be explicit: "with ambient street sounds", "dramatic orchestral score", or "silent, no audio"
Resolutions & Duration
| Model | Resolution | Best for |
|---|---|---|
| Veo | 720p | Drafts, mockups, fast iteration |
| Veo | 1080p | Final web/social delivery |
| Veo | 4k | Premium cinematic output (8s clips only) |
| MiniMax H3 | 768P | Drafts and longer clips at the lowest cost |
| MiniMax H3 | 2K | Final delivery, up to 15 seconds |
Duration: Veo takes 4, 6, or 8 seconds; MiniMax H3 takes any whole number from 4 to 15 seconds. Default is 8 seconds for both.
Video Analysis
AdaL can watch and analyze existing video files — reviewing quality, extracting timestamps, understanding content, and answering questions about what's on screen.
Analyze product_demo.mp4 — rate the visual quality 1-10 and note any artifacts or jitter.
Watch intro.mp4 and find the exact timestamp of the best hero frame for a thumbnail.
Review this clip and tell me if it matches the creative brief.
Models
| Model | Flag | Best for |
|---|---|---|
| Gemini 3 Flash | flash (default) | Fast, cheap (~$0.003/video) — general review, quality checks, content analysis |
| Gemini 3.1 Pro | pro | Best accuracy (~$0.01/video) — precise timestamps, nuanced judgments, complex analysis |
Use Cases
- Quality review: "Rate visual quality 1-10. Any artifacts, blur, jitter? Timestamps?"
- Frame extraction: "What timestamp shows the best hero frame for a thumbnail?"
- Clip extraction: "Find start/end timestamps for the product close-up segment."
- Content review: "Does this match the brief? What's the emotional tone?"
- Self-reflection: After generating a video, AdaL can review its own output before showing it to you.
Video analysis supports background mode — for large files, AdaL starts the analysis and polls for the result automatically.
Related
- Built-in Tools — disable or restrict image/video tools with
--disabled-default-tools - Input Methods — paste and drop images for analysis
- Agents & Modes — media tools are available across all agents