Video

AI Video Reference Images

AI video reference images give the model something concrete to follow instead of the prompt alone: the subject in the shot, the visual treatment, or the location. Upload a reference, write a short prompt for the action and the camera, and generate. The tool on this page takes one reference image; on the Astorie canvas you can combine several and keep them available to every later shot, so a sequence draws on one set rather than a fresh upload per clip.

Free to start. No card required.

Turn a reference image into a shot right now

Pick a reference image, say how the camera should move, and hit Create. The reference locks the look of the shot while the model animates it, on a real project you can keep editing.

Before and after examples

Before: original source image (Crane up example)Before
After
Crane up: before / after
Before: original source image (Slow push in example)Before
After
Slow push in: before / after

What this feature solves

Prompt-only video generation returns something different every run. The same brief gives a different face, a different palette, and a different room, which is fine for one clip and unusable for a sequence meant to read as a single piece of work. A reference image narrows how much the model has to invent.

References also come in types that do different jobs. A subject reference says what should be in frame, a style reference sets the treatment, and a scene reference fixes the location. Tools that accept one reference per generation make you choose between them. Supplying them separately helps the model apply each to the right part of the shot.

Models read references differently, so the useful question is which model handles the kind of reference a shot depends on. Vidu and Kling O3 in reference mode are strong on subject and scene; Seedance 2 tends to hold product and packaging detail better. Comparing them on the same reference set is faster than guessing.

How the workflow works in Astorie

  1. 1

    1. Upload the references

    Add the subject, style, and scene images you want the model to follow. Keep each one focused on its own job rather than compositing them into a single picture.

  2. 2

    2. Pick the references this shot needs

    A dialogue close-up may only need the subject reference; an establishing shot may need the scene and the style. Fewer, more relevant references usually beat more of them.

  3. 3

    3. Choose a model that reads that reference type

    Vidu and Kling O3 in reference mode are the usual choice for subject and scene work. Seedance 2 is the stronger option for product and brand detail.

  4. 4

    4. Prompt for what the references do not show

    Describe the action, the camera move, and the mood. Re-describing what the reference already contains tends to pull the result away from it.

  5. 5

    5. Review, then sequence and export

    Check the take against the references before generating the rest of the shots. Order the cuts in the sequence builder and export to Premiere Pro, DaVinci Resolve, or Final Cut Pro.

What you need to know

Accepted files
JPEG, PNG and WebP. Photos in HEIC or HEIF are converted to JPEG in your browser before they upload. GIF, animated WebP and SVG are not accepted.
Image size
Both sides must be at least 300 px. PNG and WebP files must be 30 MB or smaller. Larger JPEG files, and very large images of any accepted type, are resized in your browser before they upload.
Uploading before you sign in
You can attach an image before you have an account. It stays in this browser, up to 25 MB, and is uploaded only when you start a generation.
People in a reference image
A reference showing a recognisable real person is refused before the generation starts. Using a real face requires a verified asset and moderation review, which this page does not cover.
What you get back
One 5 second clip at 720p, in the aspect ratio of the image you attach.
Batch generation
This tool runs one generation at a time, and each run returns a single result.
What it costs
Uses up to 65 credits for a 5 second 720p shot. Plan discounts and included allowances may lower this.
Free allowance
A new account starts on the free plan with 40 credits a month. Some accounts are asked to claim that allowance once, on the credits page, before it can be spent. See pricing for what the paid plans include.
Your files and results
Uploads and generated results belong to the project they run in. This demo creates that project in your own account rather than in a team workspace. How long they are kept, and how to have them deleted, is set out in the privacy policy.
Commercial use
What you may do with a result is governed by the terms, and the model you pick may carry its own restrictions. Read both before you publish or sell a result.

Why Astorie is different

Astorie holds references as canvas nodes rather than per-clip uploads. Add the subject, style, and scene images once, then connect whichever ones a given shot needs to its video node. The references stay available to every later shot, so a five-cut sequence draws on the same set instead of being re-uploaded each time.

The same reference set can drive more than one model. Send it to Vidu, Kling O3, and Seedance 2 in parallel on a difficult shot and keep the take that follows the references most closely, then finish on the canvas: trim, crop for vertical, add audio, and export the sequence to Premiere Pro, DaVinci Resolve, or Final Cut Pro.

Example workflow

A fashion brand is producing a 60-second film with five shots, all featuring the same model in the same styled wardrobe across different urban locations. The team uploads three kinds of reference: a portrait of the model, a mood-board image for the treatment, and one location image per scene. The portrait and the mood board are connected to every video node; each location reference goes only to its own shot. The rooftop shot runs on Vidu, the subway shot on Kling O3, and the seated product moment on Seedance 2. They review the first take of each shot against the references before generating the rest, then sequence the five cuts and export to DaVinci Resolve for color.

Common use cases

Hold a visual treatment across a campaign

Use one style reference across every shot so the cuts read as a single piece rather than a set of unrelated clips.

Keep a location or set continuous

Feed the same scene reference into shots that are meant to happen in one place, so the environment does not change between cuts.

Animate a product without losing its detail

Supply the approved product photograph as the reference so labels, materials, and packaging carry into the moving shot.

Compare how models read the same references

Send one reference set to several video models for a difficult shot and pick the take that follows it most closely.

Turn an editorial photo set into motion

Use finished stills as references and generate video that keeps their framing and treatment.

Recommended model stack

Tips and common mistakes

Tips

  • Use the largest version of each reference you have. Models downscale, but detail that was never in the source cannot be added back.
  • Keep one reference per role. A composite that mixes subject, treatment, and location weakens all three.
  • Run a difficult shot through two or three models on the same reference set. Which model handles references best varies by shot type.
  • Keep the reference fixed for the length of the project. Regenerating it mid-way introduces drift that compounds shot by shot.
  • Save the reference set with the canvas so the next campaign starts from it rather than from scratch.

Common mistakes

  • Compositing subject, style, and scene into one reference image. The model reads it as a single input and loses the specificity of each.
  • Re-uploading references into every video node instead of connecting them from one canvas node, which makes it hard to tell which shot used which version.
  • Choosing a model that does not support references for a shot that depends on one.
  • Writing a prompt that re-describes the reference. The reference does that job; the prompt should cover the action and the camera.
  • Locking the subject but not the treatment, which still leaves the cuts looking like separate pieces.

Frequently asked questions

How many reference images can I use for one shot?

It depends on the model. Some accept a single reference, others take several: Vidu Q2 Subject Reference works with one to seven images, Bach 1.0 Reference combines up to nine, and Seedance 2 reference mode accepts a mix of image, video, and audio references. The model picker shows what the selected model takes. The right number is the smallest set that fixes what matters, not the maximum the model allows.

What should a reference image look like?

Uncluttered, well lit, and dominated by the one thing it is meant to fix. A style reference should be carried by its treatment rather than by a competing subject; a scene reference should show the space rather than the people in it. Uploads have to be at least 300 pixels on each side, and using the largest version you have gives the model more to work with.

Which image formats can I upload as a reference?

JPG, PNG, and WebP upload directly, and HEIC or HEIF photos are converted to JPEG in the browser first. GIF, animated WebP, and SVG are not accepted. Both sides have to be at least 300 pixels. PNG and WebP files have to be 30MB or smaller and are rejected above that; larger JPEGs are re-encoded in the browser instead, so the size ceiling is not a hard limit for them.

Can I mix references from different sources?

Yes. The subject can come from a portrait shoot, the treatment from a mood board, and the location from a scout photo. Keep each reference focused on its own role rather than compositing them, and connect only the ones a given shot needs.

What do I get when I download the result?

The generated video downloads in the format the model produced. Astorie streams the original file rather than re-encoding it, and the sequence export writes frame-rate-clean files for Premiere Pro, DaVinci Resolve, and Final Cut Pro.

Are my reference images private?

References and generated video are stored against your account. How they are processed, how long they are kept, and how to delete them are set out in the Astorie privacy policy, which is the authoritative source; this page does not restate its terms.

Can I use the generated video commercially?

Astorie does not claim ownership of what you generate. Commercial use also depends on the rights in the reference images you uploaded and on the terms of the model that produced the video, and some models require verified assets and moderation review before a real human face can be used as a reference. Check the Astorie terms of service before publishing.

Is there a free allowance?

Astorie has a free plan that does not require a credit card and comes with a monthly credit allowance. Some accounts are asked to claim that allowance once, on the credits page, before it can be spent. Video costs more per generation than images and varies by model, duration, and resolution, so check the Astorie pricing page for current rates.

How is this different from character consistency?

This page is about using reference images as an input to video generation, whatever the reference happens to be: subject, treatment, or location. ai-character-consistency covers the narrower goal of keeping one person identical across images and video, and consistent-character-video covers the multi-shot video delivery of that. Start here for how references work, go there for identity specifically.

Related how-to guides

Related models and tools

Related features

AI Character Consistency Across Images and Video

Keep a subject consistent across image and video generations on Astorie using reference workflows.

AI Style Transfer — Apply the Look of One Image to Another

Upload a source image, choose a style or describe your own, and generate a styled version on Astorie. Supported formats, downloads, and free allowance explained.

Multi-Shot AI Video — Build Connected Scenes, Not Isolated Clips

Plan, generate, and sequence multi-shot AI video on Astorie — keep characters, style, and motion consistent across shots.

AI Image to Video — Animate Stills Into Production-Ready Shots

Turn still images into production-ready video shots on Astorie's canvas — every major i2v model, side-by-side fanout, NLE-export ready.

AI Influencer Video Generator — Repeatable Character Pipeline

Design, generate, and scale AI influencer videos on Astorie — character library, voice cloning, lip-synced video, all on one canvas.

AI Product Video Generator — From Product Image to Ad Video

Create product ads and demos from product images on Astorie's canvas — chain product photo to multi-shot video across Seedance, Runway Gen-4, and GPT Image.

AI Talking Head Video — Spokesperson, Course, and Narration

Produce spokesperson, course, and narration videos on Astorie's canvas — Kling Avatar, OmniHuman, ElevenLabs, Fish Audio, locked identity end to end.

Video to Video AI — Restyle, Edit, Transform Source Footage

Restyle, transform, and edit source video on Astorie's canvas — Runway Aleph, Kling O3, Wan chained into multi-shot pipelines.

AI Video Generator — Multi-Model AI Video Production on Astorie

Multi-model AI video generation with text, image, reference, and editing workflows on Astorie's canvas.

Text to Video AI — Generate Video From Prompts on Astorie

Generate video from prompts and chain outputs into scenes on Astorie's multi-model canvas.

AI Explainer Video — Educational and B2B Demo Videos

Generate explainer videos, B2B demos, and educational content on Astorie's canvas.

AI Video Generator for Social Media

Build TikTok, Reels & YouTube content at scale with reusable Recipes and 50+ models — Sora 2, Kling 3.0 4K, Veo 3.1, Runway Gen4 — on one canvas.

Related docs

Related reading

Comparisons

Build it on the canvas

Open Astorie and wire this workflow up in minutes. Free to start — no card required.