Google offers Gemini distillation as a service for custom models
Build smaller, cheaper models from Gemini with 80% cost reduction...
Deep Dive
Google appears to be offering AI model distillation as a service.
Key Points
- Distillation reduces Gemini model size by 70% and inference cost by 80%
- Retains 95% of parent model accuracy on key benchmarks like MMLU and HumanEval
- Integrated with Vertex AI for easy fine-tuning and deployment
Why It Matters
Enterprises can now deploy custom, cost-effective AI agents without sacrificing performance on core tasks.