Kling: the line that turned image-to-video into a storyboard tool
kling
Kuaishou's Kling was among the first Chinese video models to reach a level where it can take commercial work. Its strength is not raw text-to-video but image-to-video plus camera control: lock composition and character in the image stage, then let the video model interpolate from that fixed keyframe, with first/last-frame constraints and camera-move instructions producing predictable short clips. It represents how video generation actually enters a production pipeline today - not one prompt to a finished piece, but a three-stage workflow of storyboard, keyframe and interpolation. We grade its capability claims C (vendor-stated), because the public comparisons are largely cherry-picked samples.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- 单次生成时长档位
- Vendor Claim · 2025-06
- MATURITY
- Product
- research → demo → product → production
Our take<p>Kling earns its place on this ladder not by scoring highest but by <strong>settling on the usable posture for video generation</strong>. Pure text-to-video is close to unusable in commercial delivery, since neither composition nor character can be fixed in advance. Image-to-video with first/last-frame constraints lets the creator land on a keyframe they approve of first, then turn it into motion. That division of labour converts "uncontrollable" into "controllable upstream, bounded downstream" - the only workflow shape that currently ships reliably.</p><p>Confidence is honestly marked <strong>C (vendor claim)</strong>. Only two things are verifiable: the product is publicly available and has iterated across duration tiers and camera controls. Statements like "industry-leading physical realism" or "X-second clips without collapse" we have <strong>not tested side by side on a shared prompt set</strong>, and public evaluation samples are cherry-picked, so they are not used as conclusions here.</p><p>One thing worth watching: Kling's camera control (push, pull, pan, orbit, follow) is a rare productised attempt in this domain to treat <strong>the camera as a first-class parameter</strong>. If that line holds, video generation moves from lottery draws to scheduling - and scheduling is exactly the precondition for agent-driven creative tooling.</p>
The problem it solves: uncontrollability
Pure text-to-video has a fatal issue in production: you cannot know in advance what you will get. Composition, camera position, character appearance and lighting are all decided by the model, prompts exert weak influence, and a single deliverable becomes dozens of lottery draws. Kling's productised path pushes that uncertainty upstream - fix the keyframe with an image model (or an artist's own drawing) first, then let the video model generate motion on that settled frame. Composition, subject and colour become controllable; only "how it moves" remains open.
That looks like a workflow tweak but changes the cost structure of the whole stage. Re-rolls drop from dozens to a handful because there are fewer things to lose at once - you are no longer betting on composition and motion simultaneously. It is why the genuinely productive tools in this domain all converge on image-to-video.
Three tiers of usability
| Use | Usability today | Notes |
|---|---|---|
| Short product/object reveals (rotation, push-in, material reflections) | High | Simple subject, small motion - the easiest thing to ship reliably |
| Atmosphere shots (landscapes, streets, weather, light changes) | High | No hands or faces, so few failure modes |
| Simple single-person medium shots (walking, turning, looking up) | Medium-high | Needs first/last-frame constraints; breaks more often as duration grows |
| Multi-person interaction and hand close-ups | Low | Extra fingers and clipping are frequent failures |
| Narrative with lip-synced dialogue | Low | Requires a dedicated lip-sync pipeline; never one-shot it |
Why camera control matters
Kling exposes push, pull, pan, truck, orbit and follow as selectable parameters. The significance goes beyond prettier clips: it lets creators specify work in the grammar of cinematography instead of natural-language adjectives. "Slow push-in to a facial close-up" is a well-defined shot instruction; "cinematic stunning footage" is not. Parameterised shot language makes results reproducible and lets the cutting rhythm across shots be planned.
One step further, this is precisely the interface an agent needs to take over a creative pipeline: an agent that can read a storyboard does not need prompting skill, it needs the ability to fill in shot parameters.
Boundaries and failure modes
- Duration ceiling: single-pass generation stops at the 5-10 second tier. Longer content is relayed through first/last frames and stitched, with slight style and identity jumps at the seams that need manual frame selection.
- Physics is texture statistics: fluids, cloth and rigid-body collisions look right, but anything requiring conservation (volume poured, bounce height, object count) gives it away.
- Text and logos: signage and packaging copy in frame is almost always wrong - composite it in post, or avoid it at the keyframe stage.
- Audio is a separate pipeline: output is silent; voice, score and effects are done elsewhere, and lip sync only works inside dedicated features.
- Cost order of magnitude: priced per second, a single 5-10 second clip costs one to two orders of magnitude more than an image. With three to ten re-rolls normal, budget at that multiple.
Verification backlog (what we do next)
To move this from C upward we need: a side-by-side comparison of Kling against two or three other video models on shared prompts and shared keyframes, measuring temporal consistency and camera-instruction adherence; a quantification of identity drift under first/last-frame constraints (face similarity between first and last frame); and recorded per-clip cost and queue time at each duration tier. Until then, read the numbers here as vendor-stated.