日本語 ← Back to home
Technology

Google Open-Sources ML Drift for Faster On-Device GPU AI

ML Drift is Google's open-source GPU inference engine for edge AI, supporting multiple graphics APIs and accelerating workloads in products including Photos and YouTube Shorts.

Article ID: TC-0006 Published: Updated: 2026-10-10

On October 8, 2026, Google's AI Edge team released ML Drift, an open-source GPU compute engine for running AI models on devices. It is licensed under Apache 2.0 and also powers GPU acceleration in LiteRT.

On-device AI can benefit from local GPU processing, but developers face a fragmented landscape of hardware, drivers and low-level APIs. Delivering both portability and high performance is difficult.

WHY ON-DEVICE AI MATTERS: Photo editing, speech processing and live video effects are sensitive to delay. Cloud processing offers flexibility but introduces network dependence and questions about data transfer. Local execution can improve responsiveness, provided the workload fits within device power and memory constraints.

WHAT A GPU INFERENCE ENGINE DOES: AI models rely heavily on tensor and matrix operations. GPUs are well suited to parallel computation, but software must map model operations efficiently onto hardware. ML Drift is an execution engine for running models, not an AI model itself.

ABSTRACTING GPU APIS: Android, Apple platforms, desktops and browsers expose different compute interfaces. ML Drift targets backends including OpenGL ES, OpenCL, Metal and WebGPU. The goal is to reduce platform-specific implementation work, not to promise identical speed on every supported device.

TENSOR VIRTUALIZATION: A tensor is a multidimensional numerical array. ML Drift separates its logical representation from its physical placement in GPU memory. That makes it easier to choose layouts and execution strategies appropriate for different hardware while preserving the intended computation.

OPTIMIZING GENERATIVE MODELS: Language-model inference has different phases. Prefill processes an input prompt, while decoding generates tokens sequentially. Prefill can exploit large parallel workloads, whereas decoding may be constrained by repeated memory access and latency. ML Drift incorporates optimization strategies that recognize these differences.

SUPPORT FOR FIVE-DIMENSIONAL TENSORS: Images already have multiple axes, and video or advanced model architectures may need dimensions for time, batches and other structures. Support for five-dimensional tensors expands the range of workloads that can be represented, though actual memory demands depend on the model.

READING THE YOUTUBE SHORTS RESULT: Google reports up to a 40% reduction in average frame latency for certain Shorts effects. Lower latency can help visual effects track a camera feed more responsively. The result is tied to specific effects and measurements and should not be generalized to all video processing.

READING THE GOOGLE PHOTOS RESULT: Some Photos operations reportedly improved by up to two seconds compared with the older GPU delegate. Shorter waits can make iterative editing feel smoother, but different operations and devices may produce different results.

HOW IT FITS WITH LITERT: LiteRT is Google's runtime for on-device machine learning, and ML Drift provides GPU acceleration within that ecosystem. Developers need to evaluate model compatibility and supported devices when migrating. Google says new features will not be added to the legacy TensorFlow Lite GPU delegate.

PRIVACY AND BATTERY TRADEOFFS: Local inference can allow applications to process images and audio without uploading them. Yet other application components may still send data to remote services. Intensive GPU usage can also affect battery life and temperature, so speed alone is not a sufficient performance measure.

WHAT DEVELOPERS SHOULD TEST: Real-device evaluation should cover startup time, latency, memory use, power consumption and numerical correctness. Tail latency and performance after extended use can matter more than a single average, particularly for interactive video features.

THE BROADER OUTLOOK: ML Drift aims to make local AI workloads easier to deploy across diverse GPUs. Together with smaller models and better hardware, such runtimes could move more capabilities onto devices. Actual benefits will depend on implementation, compatibility and measured performance.

ML Drift is designed to abstract differences among OpenGL ES, OpenCL, Metal and WebGPU. Its tensor virtualization approach separates a tensor's logical representation from its physical GPU allocation, reducing the need to maintain separate shaders for each backend.

The engine targets both traditional machine-learning workloads and generative models. Google describes optimizations that adapt to the prefill and token-generation stages of language-model inference, alongside support for five-dimensional tensors.

Google reports up to a 40% reduction in average frame latency for certain YouTube Shorts effects and up to two seconds saved in selected Google Photos processing compared with the legacy GPU delegate. These are workload-specific examples, not universal performance guarantees.

ML Drift can also be used as a standalone library. Google says the older TensorFlow Lite GPU delegate will no longer receive new feature updates, making this a significant transition for developers building local AI experiences.

Source

Google Developers Blog (2026年10月8日) ↗