BotBeat
...
← Back

> ▌

Base44Base44
PRODUCT LAUNCHBase442026-07-20

Base Launches BaseRT: 6.4x Faster LLM Runtime for Apple Silicon

Key Takeaways

  • ▸BaseRT delivers up to 6.4x faster prefill performance than llama.cpp and 3.9x faster than MLX on Apple M-series chips
  • ▸Enables fully local inference with zero data transmission, addressing privacy concerns in on-device AI
  • ▸Designed for engineers building production-grade on-device AI applications with open-source models
Source:
Hacker Newshttps://www.basecompute.co/getbasert↗

Summary

Base has announced BaseRT, a new inference runtime optimized for Apple Silicon that significantly outperforms existing LLM frameworks. The runtime achieves up to 6.4x faster prefill performance compared to llama.cpp and 3.9x faster than MLX, with additional gains of up to 33% on decode operations. BaseRT is designed to enable local, privacy-preserving AI inference—allowing developers to run open-source models entirely on their machines without API keys or data transmission to external services. The runtime targets engineers building on-device AI applications, including integration with local coding agents.

Editorial Opinion

BaseRT's dramatic performance improvements could significantly accelerate adoption of on-device AI on Apple hardware, a critical segment as privacy concerns around cloud inference intensify. The runtime's focus on local execution and open-source compatibility positions it well for the growing ecosystem of edge AI developers. However, real-world adoption will hinge on breadth of model support and whether it can capture developer mindshare against established frameworks like llama.cpp.

Large Language Models (LLMs)Generative AIMLOps & InfrastructureAI Hardware

More from Base44

Base44Base44
PRODUCT LAUNCH

Base44 Launches Custom AI Model as Startups Seek Defensibility Against Frontier Models

2026-07-05

Comments

Suggested

Google / AlphabetGoogle / Alphabet
PARTNERSHIP

Ukraine Digitizes Government Licensing Process with Google's Gemma AI

2026-07-20
NVIDIANVIDIA
RESEARCH

NVIDIA Research Achieves Near Speed-of-Light Latency in GPU Collective Communication

2026-07-20
StokeStoke
PRODUCT LAUNCH

Stoke Introduces Hard Budget Caps for Runaway AI Agents

2026-07-20
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us