My Journey Running AI Locally: oMLX saves the Day
After my experience with vLLM on NVIDIA hardware, I was painfully aware how behind Apple Silicon was in terms of software support. My 11 tokens per second on MacOS versus the 30 per second I saw on...