Google DeepMind premiered Dear Upstairs Neighbors at the 2026 Sundance Film Festival, a 6-minute animated short directed by Pixar story artist Connie He. The film follows Ada, a sleep-deprived woman whose noisy neighbors trigger increasingly unhinged hallucinations, rendered through fine-tuned versions of Veo and Imagen. that maintain artistic control while scaling hand-crafted animation.
A 45-person crew developed entirely new AI capabilities for this project. The result challenges assumptions about AI's role in creative work by positioning generative tools as a stylization layer rather than a replacement for human artistry.
The Core Challenge: Unique Styles That Off-the-Shelf AI Couldn't Deliver
The expressionistic visual styles director Connie He envisioned were central to the storytelling but extremely difficult to achieve in traditional animation. Her storyboards called for a series of hallucinations that shift through multiple painterly styles as Ada's night progresses, from muted bedroom tones to neon expressionism.
The team quickly discovered their vision was too specific for existing tools. Production designer Yingzong Xin (character designer on Turning Red, Soul; director of Nini) created concept art with extruded proportions and angular shape language that required the AI to learn deep artistic concepts, not just surface-level style transfer.
We didn't type this film into existence. We crafted it with a team of 45 people and brought it to life with this new technology.
Google's researchers built tools allowing artists to fine-tune custom Veo and Imagen models on their artwork, teaching the models new visual concepts from just a few example images. What the AI learned surprised the team: not just superficial details like color and texture, but principles like two-point perspective and how to maintain character silhouettes that follow 2D animation rules even as forms rotate in 3D space.
Video-to-Video: Show, Don't Type
Text prompting alone couldn't control the rhythm of Ada's sleepy fingers typing, the comedic timing of her facial expressions, or the exact framing of a camera reveal. Using text-to-video with the fine-tuned Veo model produced scenes that looked like Ada, but their movement was random, uncontrolled, and often bizarre.
The solution was video-to-video workflows inspired by how animators naturally communicate: by drawing or acting out scenes. Animators created rough animation in their preferred tools, which AI then transformed into fully stylized video with an adjustable balance between firm control and improvisation.
This approach kept motion and timing in human hands while offloading the labor-intensive stylization process to AI.
Multiple Pipelines, Same Philosophy
Different animators used different tools, all feeding into the same AI transformation process:
Maya to Veo - Animator Ben Knight created rough 3D animation, and researcher Andy Coenen used fine-tuned Veo models to transform it into the final painterly look.



