Seamless 30-second Kung Fu Panda from two MiniMax H3 clips: ComfyUI-H3-Motion-Context chains clips via latent slicing and audio continuation

Developer @eternityspring demonstrates (29-second video): MiniMax H3 generates at most 15 seconds per call — how do you join clips for a short drama without visible seams? He recommends the open-source ComfyUI-H3-Motion-Context pack, showing a 30-second Kung Fu Panda clip: "If I didn't say it, who could tell this 30s Kung Fu Panda was generated in two segments?" The common approach — feeding the last frame back as an image-to-video start — lacks motion information, so the model guesses direction and speed, causing stutter or sudden action changes at the seam. Motion Context instead pins the previous clip's last 22 frames (0.92s) at the head of the next clip, frozen at every sampling step — the model sees real motion, not a guess. Picture and sound are sliced directly from the previous clip's latent (no VAE decode-re-encode round trip), preserving the numbers bit for bit and eliminating the color drift and softening that accumulate over long chains; the pinned audio window ends at the join and reaches backwards, so the soundtrack continues rather than restarts — instantly audible on anything with a beat. After generation, a Trim node removes the pinned head and the two segments concatenate directly, no transitions needed. The key prompting trick: open the next segment by describing how the previous one ended — same character, same framing, same action held for a second or two before new plot. Frame-by-frame inspection shows pose, position and expression all matching across the seam. Core nodes: Motion Context / Trim / Save-Load Latent (carrying H3's paired video+audio latent across runs) / Seam Probe (quantifying whether a join is a real continuation or a convincing imitation) / Chain loop node (Approve auto-advances indices, segments controls how many clips run) — without modifying ComfyUI itself, refusing to run if layout checks fail.




