Google's I/O 2026 production account documents the pipeline behind TPU Training Day. Director Laurie Rowan and Nexus Studios began with cardboard, markers, puppetry and simple 3D animation to establish performance, composition and camera motion. Nano Banana then generated styled first frames, a custom Google AI Studio tool compared those frames at scale, and Gemini Omni with other experimental models merged the base animation and visual treatment into sequences.
Each layer solves a different problem
| Layer | What it preserves | What needs review |
|---|---|---|
| Puppetry and simple 3D | Performance, framing, camera | Rhythm and readable motion |
| Styled first frames | Material, color, character appearance | Pixel match and character continuity |
| Sequence merge | Base motion and styled result | Jitter, deformation and visual breathing |
The sequence suggests a useful short-drama principle: make non-negotiable performance visible in a base before asking generation to supply style and finish. When pauses, eyelines, weight shifts and camera positions already exist, failure is easier to locate. When a prompt carries performance, framing, material and continuity at once, a team can only guess which layer failed in the final image. A hybrid pipeline is not automatically cheaper, but it can assign revision more clearly.
Google explicitly says the pipeline was designed to preserve small human imperfections in puppet performance. That is the maker's account of its own film, not an independent shot test. It still clarifies that consistency does not mean erasing every difference. A character asset system should separate locked structure, expressive variation and performance noise worth retaining. A quality sheet that tracks facial drift but not rhythm may produce clean continuity with little life.
