Base Launches BaseRT: 6.4x Faster LLM Runtime for Apple Silicon
Key Takeaways
- ▸BaseRT delivers up to 6.4x faster prefill performance than llama.cpp and 3.9x faster than MLX on Apple M-series chips
- ▸Enables fully local inference with zero data transmission, addressing privacy concerns in on-device AI
- ▸Designed for engineers building production-grade on-device AI applications with open-source models
Summary
Base has announced BaseRT, a new inference runtime optimized for Apple Silicon that significantly outperforms existing LLM frameworks. The runtime achieves up to 6.4x faster prefill performance compared to llama.cpp and 3.9x faster than MLX, with additional gains of up to 33% on decode operations. BaseRT is designed to enable local, privacy-preserving AI inference—allowing developers to run open-source models entirely on their machines without API keys or data transmission to external services. The runtime targets engineers building on-device AI applications, including integration with local coding agents.
Editorial Opinion
BaseRT's dramatic performance improvements could significantly accelerate adoption of on-device AI on Apple hardware, a critical segment as privacy concerns around cloud inference intensify. The runtime's focus on local execution and open-source compatibility positions it well for the growing ecosystem of edge AI developers. However, real-world adoption will hinge on breadth of model support and whether it can capture developer mindshare against established frameworks like llama.cpp.



