INTERVIEWS

Luma's New Agents Let You Talk Instead of Writing Prompts, Says COO Caroline Ingeborn

Luma's New Agents Let You Talk Instead of Writing Prompts, Says COO Caroline Ingeborn

Caroline Ingeborn thinks the AI industry sold creatives the wrong idea when it pushed them to become prompt engineers. The chief operating officer of Luma argues that skilled filmmakers should talk to an agent the way they would talk to a person, and let the software write the prompt.

In this episode of Inside the AI Studio, we talked with Caroline Ingeborn, chief operating officer at Luma, about why the field built AI models one modality at a time and why she thinks prompt engineering was the worst thing ever pitched to creatives.

  • Single-modality AI was "the wrong bet." Ingeborn says language-only, image-only, and video-only models hit a ceiling because creative work spans every format at once.
  • Uni-1 is trained jointly on text and image. Luma's unified model currently outputs images and is built for editing tasks like layering, with speed and cost treated as first-order concerns.
  • You talk to Luma Agents, not write prompts. The agent turns plain conversation into a prompt, and creators can click to see what it generated.
  • Luma Agents route to outside models. The system picks the best model for a job and quite often doesn't use Luma's own, naming models like Seedance and Ray as examples.
  • Jon Erwin shot a one-hour pilot in roughly a week. The Innovative Dreams production used Amazon's Stage 15, an LED volume, and Sir Ben Kingsley, working nonlinearly with actors and AI creators side by side.

Single-modality models hit a ceiling, so Luma trained Uni-1 on text and image together

Ingeborn traces Uni-1 to a call the field made early on, when labs decided to build intelligence one modality at a time: language models, image models, video models, 3D models. She does not think that was wasted effort, but she does think it was misdirected. "That wasn't a bad bet, but it was the wrong bet," she says, because that kind of intelligence "hits a ceiling."

The limit shows up fastest in creative work, which she says "is not image or video or language." Her fix is a harder engineering path: models jointly trained across modalities so they reason the way a person does.

Ingeborn's shorthand for why unified training matters is the way a mind works when it is not locked into one format:

I don't know what happens in my head when I dream. I don't know if I'm dreaming in text or video or image or 3D or whatever it is, but I know that I've never dreamt in a flow of text or a bunch of images. It's all meshed together and that's what the human brain is like.

Uni-1 is the first model on that unified architecture, trained on both image and text so it understands each. For now it produces image outputs.

Ingeborn frames its strengths as editing and layering, a look she calls "beautiful, it's cinematic," and efficiency in both time and cost, which matters when an agent is generating many images from a script rather than one at a time. Unified video models, she says, are "the next line of thinking," built on a different architecture than Luma's current video generation.

You talk to Luma Agents like a person, and the software writes the prompt

Luma Agents is an end-to-end agentic workflow that runs in a model-agnostic canvas. It is also where Ingeborn's prompt-engineering critique lands hardest. "You don't write prompts," she says, calling prompt engineering "one of the first and worst value props to any creator."

The alternative she describes is conversational. A creator can "talk to the agent exactly as you would with any other person," and the agent translates that intent into a prompt for the models. If you want to inspect the result, you can click through to it. "They look like a little bit of crazy work, but it works," she says.

Ingeborn's core argument is that the people who already know craft are the ones the tools reward, not a new class of prompt specialists:

If you're a really good VFX supervisor, if you're a good writer, if you're a good costume designer, if you're a great editor, you are going to be the people that can get the best results out of these products.

The agents also decide which model to use for which task, and Ingeborn says they "quite often don't use Luma." Experienced AI creators once had to know that "I need Seedance for this or Ray for this," she says, and she does not think most creators should have to carry that information.

The canvas is multimodal across sound, video, image, and text, and it lets work move nonlinearly instead of one step at a time. Google Flow took a similar step with an agent for planning and tool-building.

That skilled-operator idea matches what one AI-native studio reported after millions of prompts, where craft judgment mattered more than prompt trivia.

Jon Erwin shot a one-hour pilot on Amazon's Stage 15 in about a week

Luma's partnership with filmmaker Jon Erwin, under a company called Innovative Dreams, started with Erwin trying an unreleased build of Luma Agents. Ingeborn says he called her right after and said "this is amazing, this is so much fun."

The production milestone came about two weeks before Luma Agents launched. Ingeborn recalls Erwin telling her, "I think I can make a 1-hour show now, and I think I can make it in a week or 10 days," work he had expected to take six months. He then shot it on Stage 15 at Amazon Studios, put roughly 100 people to work, and cast Sir Ben Kingsley.

Ingeborn describes how the shoot came together on that compressed timeline:

He got Sir Ben Kingsley to participate and he got it made. But he got it made by working in a nonlinear way and being able to have actors and AI creators side by side.

Erwin worked with an LED volume, capturing background plates in camera so the footage was usable without a heavy post-production VFX pass. One brand-focused LED volume playbook is built around the same in-camera approach.

According to Ingeborn, Erwin's read on the tradeoff was that it is faster and cheaper, and, in his words, "more fun for me as a filmmaker." She also frames it as a location shift. Erwin spent years producing movies in Europe and can now work in Hollywood instead.

Ingeborn expects hybrid filmmaking and agents that reach across tools

The same logic that makes Luma Agents pick outside models is pushing Luma toward connecting with other platforms. Asked about MCP and interoperability with systems like OpenAI's, Ingeborn said agents will need to work with other agents, and that she would have liked Luma to support that already. Security is the gating factor, she says, especially for enterprise customers, where linking agents opens the door to consequential actions.

On virtual production, she is direct. Asked whether AI will be an unlock for the technique, she said yes, and noted animation is already there. For 3D environments and camera continuity, she does not think Luma should build everything itself, pointing instead to an ecosystem that connects the pieces. Skywork's open real-time world model targets exactly that kind of interactive scene generation.

Her forecast for the coming stretch of AI media is hybrid filmmaking, with live performance and AI generation on the same production:

I think the excitement of it being able to have real-life performance together with AI, greenlighting stories that otherwise wouldn't have been told is huge. It's positive, and I think it's going to move a lot faster than any of us can foresee.

For working crews, her bet is that agents absorb the mundane parts of a job and let skilled people move faster, and that the craft they already have is what makes the tools pay off.

The full conversation with Ingeborn runs in this episode of Inside the AI Studio, shot at AI on the Lot.

KEEP READING