Google DeepMind introduces Gemma 4, a multimodal AI model with a new architecture

Google DeepMind has released Gemma 4, a multimodal AI model designed to operate efficiently on consumer hardware. Unlike traditional AI systems that rely on separate neural networks for different modalities, Gemma 4 utilizes a unified architecture that processes images and audio directly within the model's internal representation. By slicing input data into small patches or 40-millisecond chunks and projecting them directly into the main transformer, the system eliminates the need for dedicated vision or audio encoders. This approach forces the model to learn to interpret visual, auditory, and textual information simultaneously, significantly reducing the number of specialist parameters required. The model is notably smaller than many large-scale open AI models, yet it maintains high performance in reasoning and tool-calling benchmarks. The release includes a technical report detailing the architecture, which aims to improve efficiency and reasoning capabilities across various applications. The model is available for download and has been integrated into platforms like Ollama, enabling local execution on standard laptops.

Google DeepMind has released Gemma 4, a multimodal AI model designed to operate efficiently on consumer hardware. Unlike traditional AI systems that rely on separate neural networks for different modalities, Gemma 4 utilizes a unified architecture that processes images and audio directly within the model's internal representation. By slicing input data into small patches or 40-millisecond chunks and projecting them directly into the main transformer, the system eliminates the need for dedicated vision or audio encoders. This approach forces the model to learn to interpret visual, auditory, and textual information simultaneously, significantly reducing the number of specialist parameters required. The model is notably smaller than many large-scale open AI models, yet it maintains high performance in reasoning and tool-calling benchmarks. The release includes a technical report detailing the architecture, which aims to improve efficiency and reasoning capabilities across various applications. The model is available for download and has been integrated into platforms like Ollama, enabling local execution on standard laptops.

Gemma 4 uses a unified architecture that processes multiple modalities without requiring dedicated encoders. The model architecture significantly reduces the number of specialist parameters by integrating perception and thinking.

Gemma 4 is designed to run efficiently on consumer-grade hardware, including standard laptops. The system processes visual and auditory data by projecting patches directly into the model's internal representation.

Gemma 4 demonstrates improved performance in agentic reasoning and tool-calling benchmarks compared to previous iterations. The technical report for Gemma 4 provides insights into the architecture that can assist other AI systems in improving efficiency.

Chapter guide

Worth noting

  • The video is sponsored by Lambda, a provider of GPU computing services for AI research and development.

Watch the original video ↗