For comparison, you can now get < 35-watt Intel Ivy Bridge processors, with 2 or 4 physical cores. Would not be surprised to see them outperform on a performance / watt or performance / $ basis, especially if you're running OpenCL accelerated computations on the GPU.
the wattage measured here is not just the processor, its also the entire boards.
Ive done some benchmarks testing Intel Atom processors Atom330's and D525's against Tegra3 processors. On average the Terga3 processors outperformed the Atoms 4x-8x in Flop/s per Watt.
Admittedly that is not the only important metric, but gives an initial performance comparison between ARM and x86 performance
That being said, the Atoms themselves are nor a good benchmark in performance / Watt. I'd rather be interested in a comparison vs. current gen. xeon or interlagos systems. Sounds silly, but ARM has been making progress and I could see them being used in the future as companions to GPUs in computing clusters. With current GPGPU computing models like OpenACC it does not really make sense to put 16 race horses (Interlagos) besides an ant colony (Fermi GPU), except if you head for high flexibility.
Well, I think Intel wants Atom to compete with ARM for that market, so it may be relevant.
There's also Nvidia's Project Denver, which will probably come out in 2014. It's based on the 64 bit ARMv8 architecture, it's a custom CPU made in collaboration with ARM, and I think they want to pair it with their next-gen GPU architecture Maxwell. It's intended for servers and supercomputers.
I've heard about that Nvidia project. I think they are on the right track. The only thing that's missing is enough programmers (und thus software) for this model. As an example (and I'm saying this as a layman in terms of databases) I think that DBMS might be able to profit a lot from the GPGPU based model. For high read traffic databases you could scale a system with n GPUs, based on how much "storage" the database needs (the storage being the GPU ram, continuously mirrored to harddisks when writes occur).
I think the consumption of the rest of the system is fairly negligible, especially considering how inaccurate the TDP figures are to begin with. Eg this[1] system with a 95W TDP CPU draws 92 watts from the wall. And that is with off-the-shelf parts. Swap in a low-power CPU and overall optimize for power consumption and I wouldn't be surprised if you'd get below 50W.
You're giving Intel graphics as a pro for servers using OpenCL? Either way, I don't know about this set-up, but next year a set-up with Cortex A15 and Mali T658 will have shared cache between CPU and GPU, which should be a lot more efficient than anything Intel or even AMD has today regarding GPU compute. From what I understand cache "latency" between CPU and GPU is a pretty big problem in the desktop space, and this ARM concept should make it a lot better.
> Cortex A15 and Mali T658 will have shared cache between CPU and GPU, which should be a lot more efficient than anything Intel or even AMD has today regarding GPU compute.
Intel's had cache (last level cache, LLC in the diagram below) since January, 2011:
Ok, but then maybe you'd want a small and cheap cluster gather knowledge on cluster deployment and management. I am sure interested in the type of stuff. I guess you will hit scale problem sooner on an underpowered motherboard. What do you think ?
Oh, and as already said by other, wattage is for the whole board. That said I am curious about how they will actually power this cluster.
http://ark.intel.com/products/65703/Intel-Core-i5-3470T-Proc...
http://ark.intel.com/products/65735
http://ark.intel.com/products/65714/Intel-Core-i7-3517U-Proc...