AI Video Generation - General Discussion, Tips, Tricks, Frustrations and Showcases

If the LoRA you're after is of a specific looking (or just specific) person, just my opinion, but I would opt for working on the LoRA with an image model. That way I could use it to place that person in any environment I wanted via prompts, and then send the images through a video model like Wan or LTX. I think it is harder and more hardware intense to try to train a video LoRA for a specific person.

Am I understanding that to be your goal? Train a video LoRA to produce a specific subject?
 
Yes, you're correct, my goal is to train a character LoRA to use with, say, wan 2.2 14B. I'm not familiar with any image models, but I'll look into that. I didn't even think about using an image gen AI, then using the resulting pictures with a video gen AI, but yeah, you're definitely right that training a LoRA for an image model should be less resource intensive. I'll report back once I have something to show!

(Edit: I completely misunderstood your previous post. I thought that you meant that training a LoRA on videos was not possible with my rig, but training one on images was, but now I get it. You didn't mean training a LoRA on videos or images for a video gen AI, but training one on images for an image gen AI.)
 
Last edited:
  • Like
Reactions: Casshern2
I think starting with an image LoRA is definitely a good way to get familiar with the whole process before jumping into video LoRA training. Once you have a character LoRA working well, you can use it to generate a consistent set of images and then experiment with feeding those into a video model like Wan 2.2.
 
  • Like
Reactions: Casshern2
I finally found an upscale workflow that didn't have OOM crashes. And it was a very unassuming video from an equally unassuming YT channel. Although, not sure this qualifies as a true upscale because the node that does the size increase is a simple resize node. But it works.

This was also a good test of a Flux 2 Klein workflow to place different subjects in the same space, so all these lovely ladies were creates separately then placed on this street.

Tokyo Street Walking 1
 
I think I'm done. I have made two LoRAs, neither worked, then I tweaked it, and now it would take forever to train one, so I put that aside for now to get wan 2.2 14B running. I eventually got it to run, chose one of the screenshots that I wanted to train a LoRA with, but the videos were not at all what I prompted. The first one was abysmal; I tweaked the workflow and the second one was better, but still not what I asked for. I linked the AI-upscaled photo and the two videos, you can view/play them in your browser. The prompt was "A young woman with blonde hair looks to her side and notices a penis." For the second one I changed that to "an erect penis", but that didn't help. This will be up for 3 days:


Also, while the videos were generating I caught a whiff of a burnt smell. The fans were not working at max speed, the GPU temperature was around 65 Celsius, so I'm not sure what it was (it could have been simply dust according to ChatGPT), but I'm not sure if I want to risk it being something serious. Unless this is normal?

So yeah, I think I'll just wait until they make a retard-proof software for people like me unless somebody here has advice on how to do better and whether or not I should be concerned about my hardware.