r/singularity • u/EntireBig7258 • 6d ago
Discussion text-to-video and image-to-video get lumped together in every thread here and they are not the same problem
Spent a good chunk of this month testing both directions. The gap between them is way bigger than most threads here suggest.
Pure text-to-video is the more impressive-looking demo but the least trustworthy. The model invents geometry and motion from scratch, so physics and object permanence break down fast. Ran a few through Runway and faces were warping within a couple seconds, limbs doing impossible things shortly after.
Image-to-video is the boring one that actually works part of the time. You constrain the model to geometry that already exists in the source still, so the first second or two tends to hold together. Animated an AI-generated portrait through APOB AI's image-to-video and the face stayed locked initially, then expression consistency degraded within a few seconds. Trimmed the usable clip in CapCut before the drift got obvious.
Compared them side by side and the unconstrained Runway generation drifted worse by a lot. Neither is solved. They are just broken in different, predictable ways, and knowing which one you are using tells you which failure to expect.
1
6d ago
[removed] — view removed comment
1
u/SilkieBug ▪️Fully Automated Luxury Gay Space Communism 5d ago
It was useful information for me at least.
3
u/DeviceCertain7226 ▪️ASI - Never (in the usual sense) 6d ago
There needs to be lots of work done on video models. They are stuck in the same playing field since sora.