MacBook Pro has plenty of compute for local LLMs to be usable. I'm getting up to ~150 tokens/s with Deepseek-v4-Flash on a MacBook M5 Max. It's quite capable for coding assistant usage.
In general LLMs are bottlenecked by memory bandwidth rather than raw compute power.
Yes, it's quantized (4 bit). Sure, it's... not quite as good as what's on offer via API. And sure, "up to" does a lot of work (I don't have an average/median for you but it feels fast to me).
But it's usable, fully local, fully private, and has no subscriptions and no operating costs other than electricity.
In general LLMs are bottlenecked by memory bandwidth rather than raw compute power.