This mainly shows that you need to watch out when it comes to unified architectures. The sticker bandwidth might not be what you can get for GPU-only workloads. Fair point. Duly noted.
But my overarching point still stands: LLM inference needs memory bandwidth, and 200GB/s is not very much (especially for the higher ram variants).
If the M1 Max is actually 90GBs that just means it's a poor choice for LLM inference.
But my overarching point still stands: LLM inference needs memory bandwidth, and 200GB/s is not very much (especially for the higher ram variants).
If the M1 Max is actually 90GBs that just means it's a poor choice for LLM inference.