3 ms·
It was an issue of fundamentally differing directions. In the mid-2000's, Sun decided to take SPARC towards designs with many small SMT cores. In the era of si
by sam1714 14d ago
It was an issue of fundamentally differing directions.
In the mid-2000's, Sun decided to take SPARC towards designs with many small SMT cores. In the era of single-core processors, the UltraSPARC T1 had 8 cores x 4 threads per core. This was at the same time Intel released the Pentium 4 with hyper-threading, so it was an industry trend.
This of course works great for very specific applications, particularly considering efficiency, but is awful for others. Scientific computation was especially bad because the T1 had only one FPU for 8 cores.
Fujitsu's SPARC64 didn't go in this direction, and stayed with a conventional design (2 way SMT at most). Sun realized this and started to also sell the Fujitsu SPARC64 for customers who couldn't use the thread level parallelism, an arrangement that lasted until the end.
The idea of lots of slow cores is still a thing today: Intel's Sierra Forest Xeon is 144 E-cores.
- whizzter 14d ago1 FPU for 8 cores is wild, but then again they were selling webservers. Still, iirc FPU's were silicon heavy back in those days and it'd be interesting to know how far ahead the foundries Sun and Fujitsu were, maybe it was simply a factor of being too far behind in the foundry race that left Sun with few options. Intel's E-Cores still are functionally complete for most parts though (excl Avx512?)? Todays limits seems to be memory bandwidth and power and I guess many of the 144 core customers are in it for virtualization and servers?
- sam1714 13d agoMy understanding is the lots of small cores are for applications where you have lots of simple, independent tasks that tend to be IO-bound: web and other front-end servers, databases, transaction processing. SMT is optimal here because you can execute other threads while waiting for IO. I haven't been closely following it, but the Intel E-cores are definitely less wide and have less execution units. The FPU/Vector hardware is 128-bits wide instead of 256. If you ask it to do 256-bit SIMD, the front-end breaks it into 2 operations.