How to Animate an Image Locally with Wan 2.2 on a 12GB GPU
Short answer: yes, a 12GB NVIDIA card can animate images with Wan 2.2. On my RTX 5070 (12GB) I render at 1536×864 with up to 55 frames, which takes about 70 minutes and gives roughly 4 seconds of video. The single biggest factor is the quality of your starting image. Everything below comes from my own renders in Loomance, not from a spec sheet.
What you need
| Part | Recommended |
|---|---|
| GPU | NVIDIA, 12GB VRAM (8GB can work at lower resolution) |
| System RAM | 32GB |
| Free disk space | about 40GB |
| OS | Windows 10/11 64-bit or Linux |
The first launch of Loomance downloads roughly 38GB of ComfyUI and model files, so the first start is slow. After that, rendering happens on your own GPU.
My real settings and render times
- GPU: RTX 5070, 12GB VRAM.
- Resolution: 1536×864, the maximum my card can handle. I push it this high because the paid cloud tools run on very powerful GPUs and I did not want a quality gap next to them.
- Length: I set 55 frames at most. The finished clip comes out at about 110 frames, which is roughly 4 seconds at around 24 fps.
- Time: about 70 minutes per render. It is slow. Plan to leave it running.
- Speed: if you want faster action or slow motion, do it in your video editor after rendering instead of fighting the model.
If you are testing an idea or you just want quick results, start at a lower resolution. It renders much faster, and you can raise it once you like the result. Upscaling the finished clip afterwards is also an option.
1. The starting image decides most of the result
This is the biggest lesson from my renders: image quality matters more than anything else. The model animates what is already there.
- Closer is better. A close-up or medium shot of a character gives more detail to work with, and the motion looks better.
- Show the action in the image. If the character is already doing the thing, the result is more convincing.
- Avoid many characters doing different tasks at once. The model can handle it, but you risk motion artifacts or faces that melt. One or two characters with a clear action is much safer.
Making the first frame with ChatGPT
ChatGPT works well for creating the starting image because it uses strong image generation tools. A few things I learned from using it a lot:
- Keep an eye on quality. Sometimes output gets worse for a while, or something changes in the model. Faces and body proportions are the first things to drift.
- For a consistent character, generate them close-up on a plain white background. Then give that image back and tell ChatGPT in the prompt not to change the character, only the pose.
- It does not always obey. When it stops listening, opening a new chat often helps.
You can also use your own photos or hand-drawn art. A good source image is the whole point.
2. Pick the right mode
- Stable Scene: keeps your character and framing steady. This is what I use when I want a character to stay exactly as drawn.
- Wild Mode: my favorite. Take a hand-drawn illustration and Wild Mode can re-stage it while keeping the original shapes. Sometimes you are disappointed and need a few more tries; sometimes you get a "wow" environment or a transformed character you would never have thought of. It is a gamble, and a fun one.
- First → Last Frame: you give a start and an end image and write a short prompt for the action between them. This is what makes longer videos possible (next section).
3. Build a longer video from several renders
One render gives about 4 seconds, but you can chain them without visible cuts:
- Render a first clip.
- Use Loomance's frame extraction to split it into individual frames.
- Take the last frame of clip one and the first frame of clip two (same characters, same action).
- Choose First → Last Frame, write a short prompt for the action you want in between, and render the transition.
Put the three renders together and you get one continuous video with no jumps. At high resolution with about 4 seconds each, that is around 12 seconds of high-quality video.
How it compares to a standard ComfyUI workflow
I tried the same prompts in a standard ComfyUI setup and in Loomance. Loomance is built on ComfyUI and Wan 2.2, so the difference is in the workflow. In my tests:
- The standard workflow had more motion artifacts and a "dead" environment. In the same prompt, leaves moved in Loomance but stayed still in plain ComfyUI.
- With several characters in the scene, Loomance tracked movement better and produced fewer artifacts.
- Loomance kept characters closer to the original image. Even hand-drawn art held the same style from start to finish.
These are my results on my setup, so your mileage may vary. If you like building workflows yourself, plain ComfyUI is a good place to experiment; see Loomance vs ComfyUI. I also recorded the comparison side by side. You can see it on the Loomance homepage and on our YouTube channel.
Try it on your own GPU
Loomance is free to use, with no render caps and no credits. Free users see a short video ad now and then when starting a render; the optional Founding Supporter plan removes ads. Your images and prompts are never uploaded. An internet connection is needed to start the app.
FAQ
Can I run Wan 2.2 image-to-video on a 12GB GPU?
Yes. On a 12GB RTX 5070 I render at 1536×864 with up to 55 frames, which takes about 70 minutes and gives roughly 4 seconds of video. Lower resolutions render faster and are good for testing.
How do I make a video longer than one clip?
Extract the frames of a finished render, use its last frame as the first frame of the next clip with First → Last Frame, and join the clips. Three 4-second renders give about 12 seconds of continuous video.
Why does the starting image matter so much?
The model animates what is already in the image. Sharp, detailed, close-up images with one clear subject give the cleanest motion and keep characters consistent.
Are my images uploaded anywhere?
No. Rendering runs on your own GPU, and your images and prompts stay on your computer.