Keywords β Title β Shorts script β Voicebox audio β MiniMax H3 lip-synced videos (9:16, one per paragraph)
Enter your keywords; the LLM proposes 10 Shorts titles. Edit any title (βοΈ) or click one to use it.
Built from the keywords and your chosen title. Click Edit script to change the text in a popup before continuing.
Each paragraph is trimmed to its first sentence before synthesis (change in config.php β audio.first_sentence_only). The TTS model is unloaded from VRAM automatically once all clips are generated (voicebox.unload_after_audio).
Upload the person's photo β the same face appears in every clip. Rendering runs as a background job β you can close this tab; finished clips appear one by one as they complete (poll: api/video_status.php).
One paragraph per block (separate blocks with a blank line). First block = hook, last = CTA. Only the first sentence of each paragraph is narrated.