Google has unveiled Gemma 3, an AI model that processes text, images, and video simultaneously while running efficiently on smaller hardware like single GPUs. The model combines an enormous 128,000-token context window, with multimodal capabilities that could transform post-production workflows by analyzing and understanding visual content alongside text instructions.
Gemma 3 represents a significant advance in making powerful AI accessible to film production teams without requiring massive computing resources.
The model uses innovative "local-to-global attention layers" that dramatically reduce memory requirements, making it possible to run sophisticated AI on standard production hardware
Its SigLIP vision encoder enables the model to analyze video content, identify objects, and even read text within images



