Interestingly, M1 doesn't have that much secret sauce in it - it's extremely wide and benefits immensely from avoiding the x86 tax on the decoder, with a few tricks up its sleeve. Apple have shown it's possible, so - although starting now isn't ideal - if Samsung have the wherewithal to do it they absolutely could.
More difficult, however, would probably be selling it - I would imagine some HPC clusters would love it, for example, but for consumer products Apple can charge through the roof because their customers are used to paying pretty high prices for (sometimes better, sometimes worse) hardware. Apple's vertical integration also means they basically don't have to bother building an acceptable ecosystem around their new chipset - there's basically no documentation or vTune-style performance tools. It's also partly a question of priorities, but AMD still lag behind Intel even after decades in the game.
The documentation is depressingly more detailed that what Apple bless you with. I always get the feeling that Apple create for the same reason a pretentious chef cooks, you're merely proof of my greatness.
Agreed, nothing magic, no tricks, just solid engineering. Larger caches, lower latency caches, wider issue, larger re-order buffer, lower memory latency (30ns without TLB misses), etc. The impressive part is that much engineering resulted in a low power chip that apple can afford to put into $700 desktops and $1000 laptops and competes with the per core performance (and wins on perf/watt) against Intel and AMD.
Yes, but keep in mind if you spec those laptops out to have 16Gb and decent storage they're suddenly double the price. If M1 is anything it almost definitely isn't cheap
What’s interesting is that even when emulating x86 they outperform AMD/Intel. That means the decoder improvements of doing a bunch of decodes in parallel (which x86 can’t do due to being CISC) aren’t the whole story. It could be that Apple just translates x86 code into fixed width to work around this or it could be other architectural improvements.
When emulating x86 Apple translates the x86 code into ARM code, so it is executed later at almost native speeds.
Only for programs that depend heavily on just-in-time compilation for Java, JavaScript etc. Apple must fall-back to interpretation of the x86 instructions, being then much slower.
Also the outperforming of AMD/Intel has been somewhat exaggerated. Apple M1 is faster in single-thread than any old Intel or AMD, but it is slower than any new Zen 3 CPU and it will be slower than the top models of Intel Tiger Lake H and Intel Rocket Lake. In multi-thread M1 is easily beaten by many processors.
On Christmas I have upgraded the CPU in one of my computers with a Ryzen 9 5900X, so I could verify that in the single-thread benchmarks that I could run on my computer, the values matched those published elsewhere for Zen 3 and they also exceeded the highest of the values published for Apple M1, by e.g. from 3% to 4% in Geekbench 5, until 24% in gmpbench.
I agree that the advantage has been exaggerated, but to be fair the M1 is a laptop class processor that is performing just slightly less than a gaming CPU. Your CPU cooler is probably close to half the size of a mac mini alone, and the CPU alone probably draws a lot more power than an entire mac mini.
When the M2 comes out with a big heatsink and fan, it is going to be extremely competitive. Although AMD might be on 5nm by then.
The TDP of the 5900 is 105w with 12 cores. So let's say 8.75w/core. A 5950X is slightly higher per core wattage.
The M1 has a maximum power consumption of 15.1w (13.8w + 1.3w) so for the high powered cores that is 3.45w/core and a miniscule amount for the low power cores. A mac mini tops out at around 20w during 8-core CPU benchmarks. While a 5950X PC will draw around 96w at the wall when idle.
I know a Ryzen desktop CPU probably won't necessarily draw its entire TDP but it is not in the same league as an low powered CPU such as the M1 or the latest Ryzen mobile CPUs in terms of power consumption.
Sorry - I meant 5990X. The 5950X is solidly within the "diminishing returns" frequency/voltage range.
The 5990X is at 4.37W/core at its maximum power consumption - 280W/64 cores. This is from full TDP, which yes it very rarely touches, and when it does it's crunching way more numbers than most benchmarks use.
This is completely irrelevant. The multicore TDP of a Ryzen can be as low as 5W per core and as high as 20W per core if you boost it a lot. It's purely a matter of what you are setting the frequency to.
The M1 core is not that far away. A single core can boost up to 15W or more. The reason why laptops catch up to desktops in single core benchmarks is because their single core power budget is almost exactly the same but every time people act as if there is still a 4x energy efficiency gain left to be exploited when there isn't.
Idle power is a matter of how integrated your system is. The M1 is highly integrated so it will consume less power just like any other integrated SoC. When people buy desktops they want to get as far away from an integrated system as possible.
Intel / AMD have a 4-wide decoder, of which can be a 6-wide decoder if executing out of the uOp cache.
Apple just went with an 8-wide decoder, surprising a bunch of people. There's not much difference between 4-wide and 8-wide, aside from Apple deciding that such a wide single-core unit was worthwhile.
That depends on how many execution units it has to play with, although I guess the bottleneck at that point could end up being the length of a predictable flow.
The vast majority of instructions by a computer are executed in loops, if you perform the translation once it's an O(1) overhead in the ideal case.
Intel and AMD both perform decodes in parallel, but not as wide as M1, but having such a large and complex decoder costs silicon and power. Taking RISC-V as an example, a fixed width frontend is a project for a student, a modern x86 decoder is probably millions to write and verify.
I bet both Intel and AMD aren't exactly in love with x86 at the moment
> I bet both Intel and AMD aren't exactly in love with x86 at the moment
This has been the case for, eh, decades now? To me it's almost a joke how many times Intel has tried to replace ia32 with a modernized replacement and miserably failed to make any headway.
I suppose Intel doesn't want to try competing with the legion of architecturally similar RISC cores and would rather "go big" with ideas like Itanium/i860/i960/iAPX. It's funny to imagine, but maybe one day they'd release a RISC-V implementing processor. Can't imagine them licensing the rights to ARM, and going with Power or MIPS also seems out of character.
This segmented sum / prefix sum / Kogge Stone is taught in carry sum adder class at the undergrad level. Sure, it's non obvious that it applies to parallel decoding but id expect this sort of thing to be a student exercise at the masters level.
Is it simple? Variable length encoding doesn't have to be overly difficult, but x86 is a weird ISA. I'm not entirely familiar with what you cite so I can't really comment, but x86's semi-unbounded prefixes and sheer volume of extensions make things difficult.
A master's project might be to write a decoder, possibly to verify (prove - everything in a CPU has to be formally verified or generated from some other formally verified tool) it.
As proof I raise that many disassemblers, which run as software and are therefore much easier to write, still disagree and struggle with x86
Kogge-Stone proved that ANY associative function can be parallelized with a prefix-sum arrangement. Associative defined as in f(f(x, y), z) == f(x, f(y, z)... or more commonly (A+B) + C == A + (B+C), where + is any associative operator)
Now "add" (or +) is associative. But it also works for *, min, max, and even weird stuff like "Can the Bishop move here" or "Can the Rook move here". So the goal is to find an associative operator that you apply byte-per-byte.
------
Okay, a brief detour. It seems obvious to me that a finite-state machine can decode x86 byte-by-byte. FSM is the "obvious" sequential algorithm that determines whether the byte is the start-of-instruction, or the middle-of-instruction, as well as what instruction it is by the end.
Remember: we're just trying to make a FSM decode ONE instruction right now. That's pretty obvious how to do that. (Alternatively, imagine a RegEx that can parse an instruction from the bytestream when given the start-of-instruction. All Regular-expressions have a finite-state-machine representation).
> Since this composition operation is associative, we may compute the automaton state after every character in a string as follows:
> 1. Replace every character in the string with the array representation of its state-to-state function.
> 2. Perform a parallel-prefix operation. The combining function is the composition of arrays as described above. The net effect is that, after this step, every character c of the original string has been replaced by an array representing the state- to-state function for that prefix of the original string that ends at (and includes) c.
> 3. Use the initial automaton state (N in our example) to index into all these arrays. Now every character has been replaced by the state the automaton would have after that character.
In short: computing the "state" of a sequential FSM applied across its inputs can be EASILY performed in parallel through the prefix-sum model.
Any finite state machine can be converted into parallel prefix form through this mechanism.
"Work Efficient" Parallel Prefix arrangement is O(log(n)) depth and O(n) total elements, which means that parallel decoding to any width (ie: 8-way decoder, 16-way decoder, or even 1024-way decoder) is LINEAR in terms of power-consumption and O(log(n)) with respect to time.
I'll have to take your word for parts of this because I'm not familiar with this proof.
I'm not sure x86 decoding is oh-so-simple, because if you bring in memory the alignment is not guaranteed so you don't even know where the instruction starts let alone where it ends. An x86 instruction can theoretically have an unbounded number of prefixes, meaningful only after decoding something after that - maybe you can do it with a FSM but an enormous one.
All in all this doesn't sound like a master's project in the slightest, because remember I said design and verify not just build a never-used-again toy.
There have been probably hundreds of millions of dollars spent on this over the years, and they're basically stuck on the current width and multiple pipeline stages.
Just as I expected: finite-state machine used to decode. Now write a FSM-compiler (which is an undergrad-level project), to automatically parallelize the implementation.
The parallelization step is probably Master's level, but a very advanced undergrad student can probably accomplish it: since all the individual elements are undergrad projects (Kogge-stone applied to associative operators, finite-state machine compiler / regular expressions)
As a FSM, verification is simple. Just generate all x86 instructions (there are a finite number of them after all), and ensure your FSM properly goes through all of them.
That state machine looks ahead more than one symbol and has a lot of memory even if you consider it one big state.
You have to verify it decodes invalid instructions as a fault which means testing the entire search space.
Hypothetically you could use a bounded model check but you need to test roughly 10^36 combinations which is still thousands of years if you can do a 1000 billion billion a second. You need a formal proof of the operation too.
You can formally verify a finite state machine by simply testing all state-transitions.
You don't need to do an exhaustive check of the 2^8^15 all 15-byte combinations. You just check all state-transitions of the state machine.
Or to put it another way: you don't need to test all 2^8^15 byte combinations. You just need to check all invalid-instructions that have ONE invalid byte, to prove the attributes of the finite-state-machine.
Now more than ever they won't because they shuttered their CPU microarchitecture team. Mongoose is dead for good and they're using ARM's reference architectures, like Cortex X1 and A78.
> macOS competitor that doesn't require me to grow a neck beard or have a voice assistant built in.
What is genuinely so bad about Linux these days? Although I like playing with low-level stuff, I derive zero pleasure from OS-fettling and I just stuck Mint on my laptop - it's been basically perfect apart from a problem with the dual-boot setup which I think is an SSD problem waiting to explode.
Half the time I read these complaints, I either see people who actually already have neck beards in denial, or people who are expecting Linux or Windows to be identical to macOS.
Everytime I install Linux, there is always something that's frustrating. Like the other day, I couldn't watch NYTimes video in fullscreen mode without Chrome crashing. Ok, so, I installed Firefox. Fullscreen video goes to my vertical monitor. No way to change this behavior.
Desktop Linux is developed by people who love CLI. It's not built by people that try to address problems of common mundane people off the street.
I'd rather use Windows than Linux for daily driver. Atleast UI will not freak out like in Ubuntu.
Half the time I read people saying have you tried Linux Desktop, I either see people who are tasteless and love being nerdy with tmux environment, or people who are expecting others to be like themselves.
As Linus Torvalds would plainly put "Desktop Linux sucks. It is the worst piece of shit attempt at Desktop OS". :-)
Are you using the Nouveau graphics driver? Chrome reliably crashes my machine with Nouveau (though nothing else seems to). After I install NVidia's driver, it's rock solid.
Google chrome, linux, and Nvidia are the best bet on trifectas. Full screen video works, cool new webGL pages work well, games work well (steam has quite a few that work with proton), stable (can login for months), netflix, amazon prime, youtube, etc just works full screen or in a window.
Most of the glitchy stuff I've seen like you describe is either nvidia+nouveu driver, or AMD's GPUs.
Desktop Linux is developed by people who develop Linux.
Our operating system in an outlet for so many cultures — privacy minded, seeking freedom or free, people who hack and patch, etc, etc. And your first instinct is to take it away, make it another Windows/macOs for you who don't even run Linux.
I feel sorry for people who tried to help you, that's you who expect others to be like yourself.
Linux gives me the system I want — functional understood minimalist retro style [1]. And I am not alone [2], these are not tasteless, that's my home. Get off my lawn!
This is truth, but not the whole truth. Beside money, you need somebody with a right vision at helm.
Look at the direction modern Gnome goes. They have the money from Red Hat and Canonical. BTW systemd too was developed by salaried employees of Red Hat. Money makes things happen, but not always the nicest things.
I disagree. Even with money, there's too much fragmentation in the Linux ecosystem, and while choice is nice, the parts often don't play well with each other. What Linux needs is more big companies taking over and forcing their direction. Ubuntu and Fedora did a lot for desktop Linux, because they forced some controversial decisions upon the community, instead of being stuck in endless battle of supporting every legacy toolkit and supporting every fork that comes up whenever a controversial decision occurs.
I sold my Hackintosh this year to upgrade. Since I didn't want to deal with setting up a Hackintosh again, I decided to go Linux for my development work after finding WSL2 unsatisfactory.
I shopped around for Linux distributions and finally settled on Pop!_OS. I loved it. No fiddling needed and the user experience is the closest to MacOS I have gotten.