구글, 순차 처리 탈피한 확산 기반 AI 모델 디퓨전젬마 공개
DiffusionGemma demonstrates that abandoning sequential autoregressive text generation in favor of parallel diffusion-based processing can achieve 4x faster inference on consumer GPUs, fundamentally changing the feasibility of deploying capable language models locally. This is particularly significant because it enables interactive applications like real-time coding assistance and local document editing that require sub-second response times—use cases where cloud inference latency becomes prohibitive. By open-sourcing the model and optimizing specifically for efficiency-constrained workloads rather than cloud-scale throughput, Google signals a broader shift toward enabling developers to deploy sophisticated AI directly on user hardware with minimal computational overhead.