This episode of Denoised starts with a simple test that lands somewhere bigger: feeding a handful of cell-phone photos to OpenAI's Codex and asking it to reconstruct a real building in Blender. From there we get into the AI Odyssey trailer everyone argued about, Foundry's SmartRoto for Nuke, Seedance 2.5's pricey teaser, and America's first genuinely competitive open-weights model. It also serves as our SIGGRAPH primer, since the show is landing in LA and we will both be there.
Quick Take
One quick test set the theme for the whole episode: a coding agent took cell-phone snapshots of a theater and rebuilt it as editable Blender geometry in roughly 10 minutes. That reframes a bigger argument we keep having on the show about where AI fits into a 3D pipeline. Is the future fully synthetic generation, or is it 3D doing the precise structural work while AI handles the final look? The AI Odyssey trailer, Seedance 2.5, and a new wave of open-weights models all push on the same question from different angles.
What We're Watching: SIGGRAPH Lands in LA
SIGGRAPH runs in LA, which makes it easy to reach for people in media and entertainment. Addy's framing: it is NAB for the research crowd, where university labs and companies show work that can be five or ten years out. Netflix already teased some of the papers it plans to present, and the density of AI researchers makes the show a recruiting event as much as a technical one.
Two satellite events are worth flagging. The ASWF Open Source Days sessions run on Sunday and dig into the open standards, like USD and OpenTimelineIO, that this episode keeps circling back to.
There is also a free AI Workflows Summit on Tuesday evening that does not require a SIGGRAPH pass, with talks on static Gaussian splats from NVIDIA, real-time volumetric workflows, and a ComfyUI agentic tool called ComfyCode.ai that runs locally and hooks into production-grade pipeline standards. We are moderating the closing panel there.
What We Tested: Codex Rebuilt the Culver Theater in Blender
Here is the test that opened the episode. We pointed OpenAI's Codex, its coding agent and answer to Claude Code, at Blender and asked it to install the tooling itself.
The agent found the third-party Blender MCP server, set it up, and connected to Blender without hand-holding.
Then came the real ask: rebuild the Culver Theater and its surrounding block using a few phone photos shot at AI on the Lot. These were not ideal reference images, mostly one-sided, with no clean angle on the corner. The Art Deco detailing (curved forms, the sphere held up on top, intricate trim) is exactly the geometry a human modeler would dread.
The result, after one revision pass asking for a more rounded corner door, was a strong starting point rather than a finished asset.
What it got: the overall massing, the lamps, the rounded corners, and a believable interpretation of the Art Deco floor stamping and side-wall movie posters that were barely visible in the source photos.
What it missed: the theater marquee sign was the most obvious error, and it skipped several of the door windows.
Why it still matters: the geometry came in cleanly separated in the outliner, ready to modify. It even generated a camera move and, on request, relit the scene as a moody night exterior, though some emissive values landed on trim that should not glow.
The whole build took around 10 minutes once the server spun up, and the agent felt notably more responsive driving the computer than screen-reading control loops we have used before.
Where 3D Becomes the Pilot and AI Becomes the Renderer
The test crystallized a prediction we have been building toward. For our own work, photoreal accuracy is not the point; a gray-box model that gets you 70% of the way there gives you real geometry to run camera moves and feed into video-to-video passes, with generative cleanup handling the rest.
Addy's larger read: a year or two ago the assumption was that generative AI would pull production from 3D back to flat 2D, because why build in 3D when you can reach final pixel directly. That was wrong. You still need 3D for control, for the specific camera move, for the sphere held up by four supports. The workflow taking shape is 3D as the pilot doing the precise structural work, and AI as the renderer filling the photorealism gap. As GPUs push from billions of polygons toward trillions and inference keeps getting faster, the pitch is an AI 3D operator that handles modeling, shading, and rigging under the hood while a director just talks to the scene.
What We Explored: A New Video Format Built From a Prompt
If an agent can install its own tools and script complex software, it can also invent new plumbing. Alex Barashkov used Codex to build Aval, an open-source format for state-driven interactive video on the web, with small file sizes, low CPU overhead, alpha transparency, and a web-native runtime.
That is the part worth sitting with. The barrier to inventing a new file format or delivery standard used to be enormous; USD took Pixar the better part of a decade. The risk is a Wild West where every studio vibe-codes its own incompatible pipeline, which is why open standards matter more, not less, as building custom tooling gets cheap.
What We Debated: SmartRoto Brings AI Into a Pro Roto Pipeline
Foundry announced SmartRoto for Nuke, an AI-assisted rotoscoping add-on the company says can be up to 4x faster while keeping artists in control. Similar masking and tracking already lives in Resolve and After Effects, but Foundry has historically been careful about how it frames AI for high-end professionals.
That care is notable given Foundry now owns Griptape and is leaning into agentic pipelines. Griptape's studio push, which we covered when Foundry brought its agents into Nuke, Blender, and Maya, emphasizes OCIO color support and local GPU inference. That local angle addresses a real studio anxiety about pushing unreleased IP to the cloud.
We landed on a live question worth testing: AI generations carry inherent noise, and the open debate is whether that hurts rotoscoping or, more likely, camera tracking, which has to solve a 2D image back into a 3D camera move.
What We Questioned: Does an AI-Generated Odyssey Prove Anything?
An AI-generated feature trailer built for a few thousand dollars got framed as going head-to-head with Christopher Nolan's $250 million adaptation. The reality check: the clip runs under two minutes, so it is a proof of concept, not a competing feature.
Some individual shots look genuinely good, well color-graded and striking on their own. The breakdowns show up in the connective tissue. AI still cannot nail real lensing and glass, so perspective and distance shots fall apart. Camera geography and placement drift, and fully synthetic performances stay in the uncanny valley.
Addy's recurring argument is the useful takeaway: the real unlock is hybrid, combining AI with physical performance and a real cinema camera and lens for the reference plate, then letting AI fill in the world around it. Solve for the expensive parts of production and cut costs in half, and that is the win. Betting on fully synthetic everything trades away quality that the technology cannot yet recover.
What We Watched: Seedance 2.5's Teaser and the $5-a-Shot Problem
On the generation side, ByteDance released a teaser made with Seedance 2.5, a soccer-themed spot following a kid dribbling a ball through London. The ball-to-foot contact, normally a hard problem for AI, holds up well enough that we may end up eating our earlier skepticism, which is the whole point of tracking this stuff in the open.
Two caveats stand out. The model reportedly supports up to 30 reference images, which raises an open directing problem: whether it can place all those references correctly in space from a wide shot plus close-ups. And the cost is steep. A maxed-out 4K shot of a few seconds runs roughly $4 to $5, which makes real iteration expensive. At those prices, shooting something practically can be cheaper than blasting a model at it, a tradeoff that already pushed some software firms to rehire humans over pricey tokens.
What We Explored: America Finally Has a Competitive Open-Weights Model
In the open-source lane, Mira Murati's Thinking Machines released its first model, Inkling, an open-weights, multimodal model built to be fine-tuned. It is not at the level of the top closed frontier models, but it is a meaningful leap as an American open model in a category that has been dominated by Chinese labs like DeepSeek.
A few standout traits from the discussion:
Built to be fine-tuned. Addy flagged a feature he had not seen elsewhere: the model can fine-tune itself on your data, so an agent running a specialized use case can adapt without the current markdown-and-context-window workaround.
A large context window and a free playground to test it, with third-party providers likely to add hosting support.
A momentum play. Open-weighting the model is a way for a newcomer to build a user flywheel rather than trying to out-scale ChatGPT or Claude head-on, much like early Stable Diffusion did with image models.
That ties into a bigger watch list: Ilya Sutskever's and Yann LeCun's new ventures. LeCun's bet, per Addy, is that LLM scaling is hitting diminishing returns and the JEPA architecture points toward a more fluid model where text, image, and video inputs blend rather than getting stitched together. Meanwhile, Kimi 3 is teased at a reported 2 to 3 trillion parameters as an open-weights release, echoing the scale of Alibaba's 2.4-trillion-parameter Qwen preview. The uncomfortable note we ended on: the biggest open model may again ship from China rather than the US.
Bottom Line: 3D Structure, AI Finish
The through-line across every story is the same split between structure and finish.
Codex plus Blender MCP shows AI can now handle the tedious software work of building and lighting a 3D scene, turning phone photos into editable geometry that feeds a hybrid pipeline.
The AI Odyssey and Seedance 2.5 show pure generation nails individual shots but still breaks on lensing, continuity, and cost, which is why real cameras and physical performance stay in the loop.
Inkling and the open-weights wave show the models underneath are getting cheaper to adapt and self-host, with fine-tuning and local inference mattering more as API costs climb.
The pattern to watch is not whether AI replaces 3D or live action, but how quickly the agent layer makes both easier to control.
Links from This Episode
Tools & Platforms:
OpenAI Codex, the coding agent used to run the Blender test
Blender MCP, the server that let Codex drive Blender
Seedance 2.5, ByteDance's teased video model
Tools & Releases:
Foundry SmartRoto for Nuke, VP Land coverage
Foundry Griptape agents in Nuke, Blender, and Maya, VP Land coverage
Thinking Machines Inkling, VP Land coverage
Aval interactive video format, VP Land coverage
News & Analysis:
AI-generated Odyssey trailer, The Hollywood Reporter
Alibaba's 2.4-trillion-parameter Qwen preview, VP Land coverage
Events:
SIGGRAPH, official site
ASWF Open Source Days, VP Land coverage
AI Workflows Summit, free SIGGRAPH satellite event





