Hands are the longest-running failure in generative imaging and the most recognisable. They have improved substantially and they are not solved, particularly in video where a hand has to stay correct across every frame rather than in one.
The practical answer is not a model recommendation. It is composition.
TL;DR
- Improved, not solved, and video is harder than stills because errors have to hold across frames
- Seedance 2.5 at 157 credits is the strongest on physical behaviour, which includes hands interacting with objects
- Veo 3.1 Standard at 448 is the quality ceiling generally
- The reliable fix is compositional: frame them out, occupy them, or keep them still
- For stills, editing a hand at 7 credits is cheaper than regenerating the image
Why hands are hard
Extreme articulation. A hand has more independent joints than almost anything else that appears routinely in images, and each configuration is plausible on its own.
Self-occlusion. Fingers hide other fingers constantly, so the model has to infer what is behind what, and it frequently infers an extra one.
Enormous variation in training data. Hands appear at every angle, in every configuration, at every scale. Unlike faces, which are usually photographed from a limited set of angles, there is no canonical hand view.
No strong prior on count. A face has two eyes and the model knows it. Finger count is weakly constrained because so many training images show partial hands.
In video, it has to stay correct. A hand that is right in frame one and wrong in frame forty is worse than one that is wrong throughout, because the change draws the eye.
Which models do better
Canon 8frame prices, a credit is $0.01 at pack rate:
| Model | Clip or image | Credits | Note |
|---|---|---|---|
| Seedance 2.5 | 5s at 720p | 157 | Best on physical interaction, including hands on objects |
| Veo 3.1 Standard | 8s with audio | 448 | Quality ceiling generally |
| Kling v3 Standard | 5s with audio | 86 | Reasonable, fails on complex motion |
| Wan 3.0 at 480p | 5s | 24 | Use for blocking, expect problems |
| Nano Banana Pro | image | 21 | Stronger on stills |
| Imagen 4 Ultra | image | 12 | Highest fidelity still tier |
| Flux Edit | image edit | 7 | Fix a hand in an existing image |
Honest assessment: the difference between models here is smaller than the difference between a shot that shows hands doing something complex and one that does not. Paying 448 instead of 24 improves your odds; it does not remove the problem.
The compositional fixes, which actually work
Frame them out. The most reliable solution. Compose above the wrists, or crop so hands are outside the frame. Most shots do not need them.
Occupy them. A hand holding a simple object with a clear shape (a cup, a phone, a box) performs far better than an open hand, because the object constrains the configuration and hides fingers legitimately.
Keep them still. Static hands hold up much better than hands in motion. If the hand must be visible, let it rest.
Partial visibility. A hand entering frame, or partially behind something, gives the model less to get wrong.
Distance. Hands at a distance are a few pixels and the error is invisible. Hands in close-up are the hardest shot in generative video.
Avoid specific gestures. Counting on fingers, a specific sign, playing an instrument, typing. These fail reliably and are exactly the shots people try.
For stills, edit rather than regenerate
If an image is right except for a hand, fixing it is cheaper than starting over: Flux Edit at 7 credits or Kontext Max at 13, with low strength and a targeted instruction. Regenerating at 6 credits gives you a new image with a new composition and possibly a worse hand.
For video there is no equivalent. A clip with a bad hand is a clip you regenerate or cut around.
Where this actually matters
UGC-style content. A person holding a product is the archetypal social ad shot and it is a hands shot. This is the single most common commercial case where the problem bites.
The practical answer: photograph the real hand holding the real product, then animate from that still. Image-to-video from a real photograph keeps the hand that was photographed, and it solves the product accuracy problem at the same time.
Craft and process content. Anything showing hands making something. Film it.
FAQ
Which AI model does hands best? Seedance 2.5 at 157 credits for physical interaction. The difference between models is smaller than the difference a good composition makes.
Are AI hands fixed in 2026? Improved substantially, not solved, and harder in video than in stills.
How do I avoid the problem? Frame them out, give them a simple object to hold, keep them still, or keep them distant.
Can I fix a bad hand in a generated image? Yes, Flux Edit at 7 credits with low strength. For video, regenerate or cut around it.
Compose around hands, or photograph the real one. The 8frame canvas is free and unlimited, and generation is paid from $19/month.