Generate and edit images, including transparent images and edits guided by reference images.
A multimodal reasoning model with a compressed key-value cache, published as open weights.
An open-weight text model focused on complex coding and long-horizon tasks.
A multimodal model for coding and agent tasks, with 18B active parameters from a 320B total model.
An experimental open-weight multimodal model with sparse attention and 262K native context.
Animate a character from a driving video, with text control over the camera viewpoint.
A dense vision-language model for coding, reasoning, and agent tasks, with image and video understanding.
An open-weight multimodal model for long-running coding, reasoning, and knowledge work.
A compact multimodal model that understands text, images, audio, and video for local assistants.
The smallest Gemma 4 variant, built for local assistants with text, image, and audio input.
A smaller on-device Gemma model for reasoning, coding, and text, image, and audio understanding.
A compact vision-language model for reasoning, coding, agents, and image understanding.
A distilled reasoning model built on Qwen2.5, for exploring reasoning on your own infrastructure.
A 12-billion-parameter image model designed for generation in one to four steps.
A compact multimodal model for text and image understanding; a building block for private assistants.
Code generation, repair, and reasoning in a compact open-weight coding model.
An open-weight language model with switchable thinking and non-thinking modes.
An open text-to-video model for experimenting with local video generation.