Overview

Logically coherent cross-shot generation. Given the context video, the input prompt, and a starting frame, our model generates a logically coherent video. In the top example, it generates a video in which the detective points to the likely culprit, following the causal relation implied by the context video. In the bottom example, it generates a video in which an employee brings in the same toy shown in the context video, preserving visual consistency across shots.

Chained Cross-Shot Generation

Each sequence consists of three or four shots. At each generation step, LogiShot uses the available video context, together with the starting frame and input prompt, to generate the next shot.

Comparison with Existing Baselines

For each case, all methods receive the same context video, input prompt, and starting frame. Their outputs are shown using the same presentation format.

Varying the Context Video

For each case, the input prompt, starting frame, and generation settings remain fixed, while only the context video varies.