Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

MacBook Pro has plenty of compute for local LLMs to be usable. I'm getting up to ~150 tokens/s with Deepseek-v4-Flash on a MacBook M5 Max. It's quite capable for coding assistant usage.

In general LLMs are bottlenecked by memory bandwidth rather than raw compute power.



But it's quantized right? It's not the same almost-free DeepSeek you get from the API.

And once the context gets large, it slows down.

"up to" 150, is doing a lot of work there.


Yes, it's quantized (4 bit). Sure, it's... not quite as good as what's on offer via API. And sure, "up to" does a lot of work (I don't have an average/median for you but it feels fast to me).

But it's usable, fully local, fully private, and has no subscriptions and no operating costs other than electricity.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: