oMLX
Mac LLM server that cuts agent wait times from 90s to 5s
评分 74 ▲ 71 ◆ 2 免费 PH 精选
概述
oMLX 将 Mac 变成本地 LLM 推理服务器,通过分层 KV 缓存将 AI 编码助手响应时间从 90 秒压缩至 5 秒。
创新点
RAM+SSD 分层 KV 缓存设计使缓存跨重启存活,大幅降低冷启动延迟;原生 Swift 实现而非 Electron,性能更优;同时支持文本、视觉、OCR、Embedding、Reranker 多类模型;兼容 OpenAI 和 Anthropic 双标准 API
目标用户
使用 Claude Code、Cursor 等 AI 编码助手的 Mac 开发者,以及希望在本地运行 LLM 的技术用户
竞争格局
与 Ollama、LM Studio、llama.cpp server 直接竞争,差异化在于持久化 KV 缓存和面向编码助手场景的极致低延迟优化,但目前仅支持 Mac 平台限制了受众范围
标签
开源工具开发者工具本地推理LLM服务器AI加速Mac应用开发者
相关产品
GPT-6 Astra 94
OpenAI's most capable model for end-to-end work
Claude Fable 5 91
Anthropic's most capable model ever — free until June 22
AlphaGenome Atlas 88
Google's AI map of every possible human DNA mutation
WeatherNext 3 88
Our most advanced and accurate global weather AI model
Gemini Robotics 2 88
Google's AI brain for the next generation of robots
GPT-5.6 88
A new standard for intelligence and efficiency