I think it's great but it's important to remember that we haven't seen Xiaomi's chips in actual devices under real world tests. We don't know if the speeds are sustainable, under what wattage, etc. Competition is still great and I look forward to learning more, of course.
I find it funny not because I support Apple. I find it funny because it feels like the leapfrogging happened in early superscalar CPU evolution. The era when the Moore's Law was working.
I enjoy it because of progress, not because of Apple.
It's been really weird reading laptop reviews over the last few months. I've seen a bunch of reviews where they have their usual bar graphs comparing a bunch of laptops, and the laptop that is at the bottom is an Apple. It's the really inexpensive Neo, of course, but even the M5 laptops are consistently beat on performance and battery life by Intel Windows laptops.
I assume the M6 will take the crown back and then a few months later Intel/AMD will release a new chip and take that crown back again. That's the state of the world we used to expect, but it's a state that has been missing ever since the release of the M1 in 2020 until Intel finally caught up again this year.
Reading Infinite Games, they tell an anecdote about a Microsoft exec on a flight telling an Apple exec that the Zune was a way better portable music player than the iPod. The Apple exec was just “yup, you’re probably right” and then soon after the iPhone dropped
Zune was probably a better player. It was too late, and arrived when MS didn't much care.
There are emails unearthed in various lawsuits where you can read Bill Gates screaming at his subordinates: "why the hell can't our partners like Sony and Creative create a similar device? Give them all, give them early access to everything, work with them". In the end MS felt compelled to make their own.
Even w/the pricing spike, inflation adjusted we are back to roughly the prices of a new Mac SE/30 for something that can beat a turing test w/o sweating.
I yield the floor to no one when it comes to pessimism, but that's incredible.
I'm expecting it to somewhat collapse. I don't know if it'll go back to pre bubble prices (here's to hoping), but I do expect a pretty sharp decline around 2030... probably not before then.
Basically everyone that makes memory is building new fabs, meanwhile I'm not sure how much longer AI datacenter demand for ram will last. I think the decrease in AI ram demand and the new fabs will likely coincide leading to a collapse in pricing.
That is, of course, assuming the memory manufacturers don't pull their favorite trick and collude.
Cycles tend to be 5-10 years long not 25. Without knowing anything else I would expect a fab you seriously start planning today will be at full capacity in about 5 years. Nobody serious likes delays - in particular the banks don't like loaning money that won't at least start paying off. They know it takes some time to design a building - but factories typically are standard buildings so once you know about the size you can get it done fast - I expect 1 year to have the building done is the worst case (and it can be done in 3 months possibly if your project management is good - after interest this is cheaper than the 1 year). It takes time to build and install the specialized machines that go inside - this is the largest problem, but you typically order them first and then plan the building around the needed space and when they will arrive. Then you need 6 months to setup the inside of the building. From there it is just ramp up time.
The above is a standard project management problem. We do this for lots of industry all the time. There is every reason to think you can get a new factory running in 5 years.
Note that I said 1 factory above. Some of the special machines we don't have the ability to make them fast enough to do 2 (I don't know the real number!) new factories in 5 years. Existing factories are using most of the special machine capacity to replace machines that wore out on the way - this can be corrected as well, but it adds another year and the expenses are much larger. Realistically though 1 new factory is likely enough.
What does that have to do with the Turing Test? The TT has clear rules: There are judges that have a dialogue with anonymized AI/humans. The humans cannot cheat and impersonate a machine, they have to act normally. The AI obviously should try to sound human.
No AI would pass this test with experienced judges.
The test doesn't say that the judge has to be experienced. But I also don't care if some random gullible person can't tell the difference. Nothing passes the Turing Test for me yet.
Of course, and I'm sure OP wouldn't disagree with you, it was clearly a joke for emphasis. Some people round here need to clean and calibrate their humour detectors more often.
Yeah...but in context, "gullible" seem a bit pejorative. Humans are also hopelessly incapable of sensing radioactivity, methanol in their alcoholic drinks, carbon monoxide, and a great many other things that our ancestors just didn't encounter much.
Though we're pretty good at sizing up a person's emotional balance/maturity and competence at familiar tasks. So maybe have an old blacksmith watch the AI/robot interact with horse owners for a while, then shoe their horses, and see how well it does.
Nowadays, I prefer AI CS agents to humans. I just had a chat with an AI yesterday, it understood me perfectly even when I made mistakes, I was impressed.
In contrast, humans tend to paste me the same barely-relevant macro over and over, no matter how much time I spend explaining my issue.
Customer service is very different. Crappiest audio quality possible & scripted answers all the way down.
Almost like humans are forced to behave like machines.
The Turing test is more complex than what gets suggested.
And the "popularized" version is faulty also since it uses an ideal, abstract human judge (like the "spheroidal economic agent").
But if you want to add declinations to the said popularized image of the Turing test, you may add Maxim Lott's IQ tests at trackingai.org . Between the end of 2024 and the beginning of 2025 LLMs reached an equivalent IQ of 100, for example.
Yes, which is exactly and entirely the point of the whole paper. Somehow missed still -- despite how important AI has become, shockingly few people actually read the short, layperson-accessible paper that started the whole field.
It depends who takes the test. I am not yet, to my knowledge, fooled by AI.
I've tried [1] and I almost 100% detect which is the AI. I really want to convince myself I have failed, does anyone know of a better site/resource for this?
I know it might be moving goalposts but I would consider AI to have passed in a well and truly undisputed manner when [2] is resolved.
But in a more practical sense, if AI can impersonate humans so well today then why are state of the art frontier models so obviously AI when they create PRs, commit messages, documentation, etc. Are the companies deliberately making them unnatural?
we might need to bring back the Voight-Kampff test. anthropic at the very least is introducing a water making system to Claude which might make them more identifiable to humans as well as much easier to detect for machines.
Not sure where you're quoting from but if it's the metaculus question comments, many of them are from 2023. The consensus is it will resolve in 2029. I believe it will not resolve before 2035.
Sure, when you look a little wider, since 2000 we have seen the following major improvements:
- Extreme poverty has dropped from 30% to under 10% globally.
- Child mortality rates have dropped in half
- Internet access has exploded from 10% to 70%
- Solar energy costs have dropped 90%
- Cancer death rates have declined by 30%
All of these massive improvements in less than 30 years.
While there certainly are issues to solve, and if you simply follow journalism you may think the world is worse off, but for many, their lives have been significantly improved.
In part, sure. From drug discovery to crop yields to education, compute and AI have material benefits. It's kind of shocking that anyone could doubt this.
This being HN, I hasten to add they also have massive downsides, we're all doomed, nobody programs the right way anymore, those poor people just think compute & AI are improving their lives, etc, etc.
AI in the sense of LLMs is too new to really make an impact yet. But it is compute driven. Modern science and engineering would be impossible as we do it now without high-end compute.
Just like the computer revolution, it will make a small number of people richer and put everyone else at their mercy. In real terms the average person is far poorer than in the 20th century. Used to be able to buy a home and support a family on a single income. Now it can take two just to survive.
Why in the world are you choosing to live that way?
I'm writing the best music of my life, realizing games and art projects I never had time for, and writing higher quality software in addition to dramatically more of it. Who has time for pablum?
How exactly are you using these tools that you have that experience?
Apple Studio with maxed out M5 Ultra, 256GB RAM and 16TB storage is 18,299$. The 512GB RAM version apparently is coming in October, considering that the difference between 96GB and 256GB is priced at 4000$, the 512GB upgrade must be eye watering.
So, on the mini the RAM upgrade runs at 25$ per GB on all tiers, the same as the Studio therefore the upgrade to 512 will probably cost 6400$.
The fully maxed out Apple Studio then will be 24699$. It's 17199$ if you don't upgrade the storage(1TB).
So is downpayment on a house. I would buy the house and just pay for tokens as needed. The house will get more valuable and that wealth would buy a lot of tokens in the future - which will probably get cheaper.
EDIT: or buy AAPL. If I had bought Apple stock instead of buying a Mac LC II in 1992, then I would have about $2 million in Apple stock.
Or more simply – $25K (+ tax) put in a savings account will earn about enough interest to pay for a $100/month AI subscription indefinitely. And at the end of it you still have the $25K.
Not 'more simply', there are basically zero savings accounts that are going to net you a 5%+ interest rate to give you that $100 a month. And that $25k becomes less valuable over time. $25k is now only worth $19k because inflation.
I tripled my money on my RAM purchase of three years ago. So, yes, for short-term appreciation that's hard to beat. But I don't think it's something that will continue.
Apple hasn't been selling just ram for a long time, they sell vram. Try getting 512 gb of HBM on current Nvidia cards - it's gonna cost way more than $ 24k. And here you get the same amount of memory for weights right in a quiet unit under your desk
Put together a similar build with a couple of rtx 6000 Ada cards and Apple's price tag suddenly looks pretty damn reasonable
i remember when SGI boxes were $50k and then literally worthless just a couple bears later. i remember my university had a pile of them for free outside the deans office.
Ten years ago a bought an expensive MBP because I do a lot of stats in R, Python etc that benefited from it. But the next Mac I’ll buy will be a much lower-end model, because it’s just easier these days to do that work in notebooks in the cloud.
0% APR for 12-24 months from Apple Financial Services for a device that can approach or even exceed what we were paying for new cars just a few years ago. Apple is definitely making bank off these financing offers, and with very little risk as unlike a car these Mac Studios don’t lose 20% of their value when you drive them off the lot.
Rumors say that Apple will only release M6 base variant and skips M6 Pro, M6 Max and M6 Ultra variants to concentrate all efforts to create a good AI capable M7:
"According to reports from Bloomberg, Apple will be skipping its M6 Pro, M6 Max, and M6 Ultra chips to accelerate development of the M7 chip. That means the only chip to be released from the M6 family will be the base M6.
The reason for this break with tradition: AI. Apple had been planning major neural-processing upgrades for the M7 family and ultimately decided those improvements were important enough to justify accelerating the next generation rather than completing the M6 lineup." https://9to5mac.com/2026/08/08/apple-m7-chip-heres-why-it-ma...
I'd skip M5 and M6 chips for LLM work and wait for a year for M7.
I live right next to a micro center and remember when they started offering that deal... still so pissed at myself for not buying one. I ended up just buying a raspberry Pi for what I was doing, but seeing as where the prices are now, I messed that up a bit. Also my worst sin was not buying 64GB of DDR5 when I was doing my computer upgrades back in August last year.
I returned an M4 Mac mini, 64GB, unopened... because I thought it was excessive for my needs then. I swear it'll be one of the things flashing before my eyes when this all ends.
The 512GB Ultra is amazing, sure. But who is it for exactly? VC funded big spender founders? In that case why would they need local AI? The ultra rich enthusiast? But there can't be too many of those. So who actually buys these?
I think there's some mental inertia around what a computer is and what it's worth. This thing can build custom software for you, mostly autonomously. It can monitor things happening on the internet that are relevant to you, in a holistic and flexible way. We have one crawling the web for local events we'll like, and it judges them based on what it knows about us, and it tells us about the best matches every weekend, which has yielded some awesome outings we wouldn't have known about. It reads the literature on a subject in seconds and uses it as context to help in decision support. It's not the same value proposition as a computer 3 years ago, where most people are mentally anchored on what a computer should cost. Having it at home means that you can use it as a personal agent that always puts your interests first, regardless of what ad model the commercial providers decide to put in, and you can stash in it your medical data, what you buy, what you make, your worries, hopes, and dreams, without worrying about that being used as training data, or worse, something to exploit you commercially. I think it'll become considered totally reasonable to consider spending the cost of a small car on a computer, for many families.
Also, a lot of companies are looking at how to run capable models locally to cut some of their (massive) cloud AI bills. An easy answer is worth a lot to them.
I like this train of thought. The inverse is saying that the cost of this computer is the value we give away to AI companies by doing compute on their servers with our data. And to take it another way, is the value to you, the cost of a small used car?
> And to take it another way, is the value to you, the cost of a small used car?
For me personally, not quite that valuable yet, but I think it's getting there quickly. Deepseek V4 Flash massively increased the value of local AI to me, to the point where it's displaced most of my Claude Code usage, its upcoming vision enabled version should bump it further, and it's only going to get better from there.
It's a lot faster, but a lot of it is also feeling free to discuss things I wouldn't be comfortable sending to Claude, with the idea that that info is now theirs in perpetuity. I got my genome fully sequenced recently (it's cheap now!), and I get a battery of blood tests every year. Wouldn't do processing on any of that with Claude, but local AI? Totally great.
And if I was running a company with a large cloud AI bill, I'd probably buy a wheelbarrow full of these macs. Cheaper, but also a more solid/predictable base to build on.
Other than the local AI crowd which is much recent it is professionals using Final Cut Pro for video editing, Logic Pro as a DAW and music production, Video transcoding, Photoshop and other tasks for high performance computing that don't need Laptops but want above 128GB of ram and prefer a Mac. Then there is the obvious group of developers that are making Apps for all of their products. Also, these are great for the workplace. AI is much more recent thing that Apple products were used for.
If Apple didn't sold these things they wouldn't make them but, also the level of marketing that Apple is talking about for AI is basically the new group they need to capture because the ones I just listed are already buying Macs and or easily to motivate with the other obvious CPU / GPU performance upgrades for code compilation, faster memory and video transcoding.
If I can get Sol level capabilities on a $20k machine, then it is well worth it for my employer to buy me that machine for work as a workstation. When you start paying in tokens vs subscription costs due to enterprise agreements, you really start to see how much cash utilizing frontier models at the frontier costs (and I'm efficiently using luna and other models where possible!)
this comment has always existed behind every apple release, most especially anything vaguely pro-ish.
to answer your question : looking at the aftermarket availability of Apple's prior best and brightest : practically no one buys them.
"people here buy them" , well, 'here' is one of the most affluent groups of people in the world.
They're available as movie and television set pieces (undoubtedly disappearing into the home of someone close to the staff post-production), and for administrative/boss types that can slip the cost into a ledger somewhere that few will ever see.
It has been a hobby of mine every few years to check out the apple site and see how big I can option a machine. My record was when I was in high school years ago and was able to option some pro studio-ish apple desktop thing to like 61,000 usd out the door.
Movie set pieces, as a motivation for Apple making these high-end configs available? That makes no sense.
For one thing, you can’t tell from a movie what the specs are. A $999 Mac Studio looks exactly the same as a $20,000 one.
For another, Apple updates the industrial design on their products so rarely, a 6-year-old Mac, iMac or MacBook also looks nearly indistinguishable from a brand-new one.
The competition is a custom multi-GPU NVIDIA RTX pro desktop, which go for much much more. $20k is cheap for 512GB addressable memory. The old mac pro could easily be configured to cost that much.
Local AI is the future, and a lot of people want the first mover advantage or to toy around with it. I know a guy with a small rack of Nvidia Spark machines that he uses for that purpose; it's as much as a decent used car.
But is it? If it’s cheap enough latency doesn’t matter. Unlike say cloud gaming, where latency does matter. I’ll take my games local and my text bots cloud
Damn. I just bought a maxed out MacBook Pro M5 Max 128GB 8TB, still waiting for it to be delivered. I could get 256GB RAM M5 Ultra 1TB for roughly the same price, and it's double the memory bandwidth. Which one would you recommend? I do plan to run local LLMs.
Leasing now an option, only $50/month (cheaper than inference subscription?), so even cash-poor can go the amortized-investment route.
I've often felt there is tremendous value locked up in underutilized old computers. It would be interesting to see Apple in 3 years offering compute as a service using lease returns (or more likely, partnering with someone else to operate it (perhaps exclusively in secondary markets like China or India, to address political demands for local siting or jobs). Apple is in the best position to work around or even gap-fix older software/hardware limitations in a controlled environment, and now they can do so without cannibalizing new hardware sales.
I'm not sure who that line is supposed to impress. Gamers focusing on graphically demanding AAA games would laugh at this. People who don't game much probably won't know whether this is good or not.
The page says 170G/s memory bandwidth for the NPU and 1.2T/s for the GPU. Why the discrepancy if it's all "unified memory"? The former is nothing to write home about as far as AI compute is. The latter is really nice.
Which one is it you can run local models on? I suppose the NPU only.
Apple still has the best hardware so I moved to it for the last few years, but the closed software ecosystem is terrible for taking advantage of it.
I wasn't able to debug network errors (restartin my Mac worked), Metal was missing low level disassembly / debugging tools (there is some hard to use UI), but the worst thing was the inflexible windowing system.
Even getting all the window handles on all screens/desktops with their titles and programs is impossible.
I just decided that I move to Omarchy 4 (basically Hyperland + QuickShell) + NVIDIA GPU, and I already was able to customize it more than my Mac in years.
I will miss Apple's hardware for sure, but not MacOS and the missing hardware documentation
I had a chance to try Omarchy past few days and it's just very different vs macOS. I honestly never had a problem with the windowing system ever since I built my own customization scripts (i.e. Hammerspoon). I can see the appeal for someone who wants ultimate customization though but macOS still wins overwhelmingly when it comes to polish, ecosystem, user experience, and apps (nothing comes close).
It's interesting because I just haven't felt the polish.
For example when using PyTorch I wanted to try to speed up my NN kernel by 2x by just using half precision and haven't noticed any speedup at all. Also I was missing the easy to use GNU tools that had to be mixed with Apple's tools.
I loved using Arc browser as well, and I'm missing it, but I guess I will do without it somehow (Chrome's vertical tabs are just not the same).
My main program missing from going back to Linux was ChatGPT Desktop, but now it's there.
I just checked out Hammerspoon, I'm happy for you that you wrote it, and looks great, but it has the same problem that I had: for security reasons Apple stopped allowing the window APIs to get all important information on other workspaces. You can only do it with Accessibility API. I was trying to fight with it but have up.
It’s not just Omarchy, there’s really not much out there in the desktop Linux sphere for those who are mostly happy with how macOS works out of the box. Everything is either in a similar vein to the Omarchy setup (hyper-minimal tiling WM), Windows-like (KDE, Cinnamon, most other DEs), or a chimera with a grab bag of design bits from every desktop and mobile platform (GNOME, Pantheon, COSMIC).
It’s a bit depressing because it means that if I ever feel forced to switch my daily driver, it won’t come without a dump truck load of friction, frustration, and lost productivity, which I’ve validated by using the various Linux desktops on secondary machines.
I ordered an ASUS Zephyrus G16 with 5090 NVIDIA card + 1.9kg (quite an overkill, and I know that I will have to limit power output), but hasn't arrived yet.
But what's fun is that I love QML+QuickShell with its hot reloading, Hyprland with its Lua support.
With AI nowdays it's just so easy to do deep UI changes that wasn't possible a year ago.
I bought a 128GB M4 Max Mac Studio a while back, and for a while I thought like I had done really well to buy it when I did.
The problem I'm having now is that no models are targeting RAM of that size. Everything is either much smaller, targeting laptops, or much larger, targeting hardware well out of reach of enthusiasts.
Please, AI people, start making models targeting 128GB machines again. The last interesting one was Qwen 3.5 122B.
I use both the base M4 Mac mini and an M4 MacBook Air for Final Cut Pro, Photoshop, and Fusion, and while they're not as almost-always-perfectly-smooth as my Mac Studio, they're about 100x better than the experience I used to have on my old Intel MacBook Pros in the 2010s.
I edit 4K ProRes and H.265 footage, sometimes with multicam (up to 4 streams) and color adjustments, titles, etc. It's only after stacking 3-5 effects before things can stutter, really.
Or if you try doing something CPU-intense in the background _while_ running some heavy creative software. I just don't do that.
I have Mac M1 Max and I'm quite happy with it. But these advances make me think that maybe I should upgrade to Mac M6 (or something) when it is released.
Sick. Particularly stoked for the 10gb network card ($100 option) when using the Mac Mini as a server. Just wish the memory + NVMe prices could come back down to pre ai-goldrush prices. As $2999 for the M5 Pro with 64GB RAM feels painfully over-priced.
RAM and SSD in apple gear has always been way over-priced. There was a short blessed period in March where the M5 Max macbook pro was out, but the general 30% price hike had not yet happened. In this period, given the insane inflated RAM prices, the price apple was charging for the M5 Max with 128 GB RAM was actually _reasonable_.
My example is I'm looking for a new media creation machine to replace my homebuilt PC from 2015. Since that old machine can't run Windows 11 and also because Apple storage is so expensive, my idea is to turn the PC into a Linux storage machine with a 10G NIC. Then I should just be able to edit off of the storage instead of worrying about caching it locally.
The prices are ridiculous though. I may just keep rolling with my Windows 10 setup.
M5 Pro in a Mac Mini with 64GB RAM and 10Gbit Ethernet seems like the perfect Jellyfin server and Ollama test server. All for just over $3K (I specced with only 1TB local nvme).
This may depend on the size of your library. I tried installing Jellyfin on a Synology NAS, which runs Plex just fine, and it ran so poorly it was basically unusable. It “worked”, but it was painful.
not sure which cpu that Synology has, but I run jellyfin on ugreen nasync dxp8800 plus which has intel with QSV and it breaks no sweat in both transcoding (if needed) and serving over 10GBE.
From what I’ve noticed, Apple products have been getting worse in quality year after year. Sometimes they even ruin their own devices with updates... I guess it’s all because of marketing.
I’m mostly talking about their iPhones, where new updates sometimes make older models worse. I know a lot of people whose iPhones started lagging after iOS updates.
Maybe in a year or so or maybe never. It depends on how 14a turns out and if it is comparable to TSMC 2NM. They may also choose to utilize 18a-p / 14a for other chips and not the M-series.
Is Apple just going to announce everything silently from now on? No more getting excited for the big events to see what's new -- it just appears on the blog one day?
Also: "a staggering 1.2TB/s of unified memory bandwidth" -- yay, the GPU has reached the year 2020! (I'm a bit bitter that my M4 Max is near useless for local LLMs because of its low memory bandwidth.)
will be great fun if one M5 Ultra with 512GB memory at 1.2T bandwidth capable of doing 3x smallish local model inferencing each at Opus 4.5 level of intelligence.
The memory bandwidth and size seems to be there, but what is the tokens per sec on like a qwen model? And you can basically do 3x opus 4.5 on the $100 a month claude plan. Your payback will be near infinity years after electricity.
"M6 also introduces a Dual 16-core Neural Engine, providing up to 2x the peak compute over previous generations to make on-device AI workflows run even faster" .
> M5 Ultra features a massive amount of high-bandwidth unified memory, up to 512GB, and delivers a staggering 1.2TB/s of unified memory bandwidth that is 50 percent higher than M3 Ultra.
Apple never needed to participate in the AI race to zero. Because they were already at the finish line years ago building their own chips that can run large >100B parameter AI models locally.
As someone who works in AI now, I have found it pretty amazing that Apple basically didn't do much with AI software, and focused more on the hardware side. I think this is what the future of AI is going to look like, local models run on your mac for your workflow.
It's possible that they're working on their own LLM that's going to work very well on their chips, and possibly outperform anything out there when they do release it.
10 years ago 32GB ram laptops sounded too much. 8 was enough. These days even I would get that much ram since it’s soldered. 64GB is higher end.
In a few years we should see such high end hardware commonplace. Working with a local LLM to get work done is the ideal way to go which has mostly hardware limitation as of now that gets solved in due time.
> 10 years ago 32GB ram laptops sounded too much. 8 was enough. These days even I would get that much ram since it’s soldered. 64GB is higher end.
Ten years ago I got 64gb of ram in my laptop, same as I have now. I bought both for business and personal use. System ram capacity hasn't changed much in 10 years.
It makes me curious how old you were 10 years ago.
We were definitely outliers that long ago. I put 64 GB in a MacBook Pro back in 2019, and that was (a) overkill for everything I ever ran on that machine, and (b) stupidly expensive by 2019 standards (albeit almost affordable by 2026 standards)
>I have found it pretty amazing that Apple basically didn't do much with AI software
The iPhone 15 was almost entirely marketed based upon AI (I would say fraudulently so, advertising features they still haven't delivered), and a huge portion of the OS work was on local AI or AI integration.
And for that matter Apple has been dumping enormous sums into their own AI development. Their failure to have a lot to show for it doesn't void the fact that they tried really, really hard.
It's bizarre how often this "Apple sat on the sidelines and let the AI people fight...so smart!" narrative appears on HN. Apple hasn't gone down the path of spending hundreds of billions on nvidia GPU data centres, but they absolutely tried really hard to matter in AI.
It's worth noting that the 5090 (or the RTX Pro 6000 big brother with 92GB VRAM) will run rings around the Mac when it comes to compute.
My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.
In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.
There is no magic, if the data you compute as atomic chunk don't fit in cache then memory bandwidth R/W limit kicks in and architecture does not matter. On contrary - having multi gpu setup of same price and same memory size with even slower memories may give you effectively much higher bandwidth but at the cost of power consumption.
I was so blown away at all the discourse surrounding "Apple fumbling on models". They should never have been in the model game to begin with. Apple crushes hardware over the last decade and that's a huge advantage today. In the end, massive models have proven to be very strong, but small models have proven to be good enough (especially with the recent Qwen 2.8 27B drop) and that's where I imagine the future will lie for consumers.
I think this is a bit of a crazy statement. Everyone expects Apple to somehow build a category leading product every year. I'd expect something innovative every couple of years
* the iPhone
* the iPad
* apple watch
* airpods
* unified memory laptops and computers
Those are all products that either created a category or changed that industry.
They did participate early on with Apple Intelligence and failed miserably. Really good move to not double down and let the others explore the space first
Is there anything comparable that runs Linux, doesn't necessarily look as good, but is perhaps (a lot) cheaper/fixable? Or is this really pretty optimal?
I mean this is not nvidia based right? It's all custom? So we can use it under Asahi perhaps?
I want to get something for my company to run local models, wondering what would be a good option.
You can't run Linux directly on these. Asahi Linux supports up to M2 only.
Linux runs very well in a VM on macOS. There are many good options for this, some free and open source (QEMU, UTM, Lima, Colima), some proprietary (VMware Fusion, Parallels).
But Linux in a VM doesn't get access to the real GPU, so model performance is limited. Those running on the CPU perform well, and those needing the GPU don't.
However, macOS on M-series macs is excellent for local models. (Maybe not as excellent as a box full of the best nVidia GPUs, but still excellent).
So if you're getting Apple hardware, like Linux, and want to run all of it locally, a fine setup for a machine to run local models, with agentic characteristics:
- macOS running one of the many local model runners. I used to use Ollama and Whisper, and now use llama.cpp instead of Ollama. Others use LM Studio, oMLX, etc. Provide HTTP endpoints to access the models.
- Linux in a VM for overall control and orchestration, with standard VM settings, and bridged networking so it appears as its own machine on your network. Also, in here provide a robust shared file server for shared state. Use this VM as your desktop and primary access to the machine, if you like Linux.
- Linux in a VM to launch ephemeral, volatile containers, with the containers using a memory-only tmpfs overlay on top of a read-only Linux filesystem in a VM disk image, with tools in this filesystem. Alternatively, a writable Linux filesystem in a VM disk image, with disk buffering set to use macOS host buffering and discard fsync requests. These settings optimise for container disk performance for data that's only ephemeral which will be deleted soon or on system shutdown. (You can combined both VMs, but need to use two VM disks to get equivalent behaviour, and be careful about VM disk configuration of the two disks.)
- Containers spawned within that second Linux VM can be spawned very quickly and run quickly, so are ideal for LLM agents that need a quick sandbox. These sandboxes generally run faster than a macOS sandbox, despite being on the same machine with VM overhead, because Linux is faster at some things. Teach the LLMs to store files and memories they want to keep in the shared file server.
I love linux and would be using it if the ARM support was better. It's just not there and most distros that support ARM do it a little poorly. I just haven't seen anything even remotely comparable to Apple Silicon and unfortunately Linux is struggling very hard to support it.
I guess, what I mean is: Why are these tiny aluminum boxes so optimal?
I just want my butt ugly repairable beast machine to do the same trick. Why is my ram not unified? I have an iGPU in my server, but it can't access the 64 GB ram directly or something? It's on the CP right? Why did only Apple go for this architecture? So many questions...
The PC platforms have anemic memory bandwidth in comparison. Eg, Strix Halo is 256GB/s max. If money is a bigger limiter than performance it can be an option though. As can Nvidia DGX Spark machines. (Also limited to 128GB memory and comparatively low bandwidth, but higher compute than Strix Halo.)
AFAIK, apple does not release drivers open source, asahi is a reverse-engineering endeavour and does not support GPU.
For nvidia, there are both proprietary and open-source linux drivers. CUDA and inference works on linux with nvidia.
I would recommend checking out this video of Alex Ziskind to shop for a computer to run local LLMs: https://www.youtube.com/watch?v=mevUEQcumzU&t=224s.
TL;DR besides Apple he recommends, DGX Spark, Tenstorrent Wormhole N300, AMD Radeon 7900 and NVIDIA RTX 5090.
Apple has the mlx framework. Most/all major software for running models locally support it. Apple also has RDMA for interconnecting multiple machines across Thunderbolt connections.
Uhhh, they are? They’re a hardware co for sure, but to say they aren’t focusing on on-device models is absurd on its face. They’ve spent over 2 years on Siri AI which is (mostly) local.
I'm very very surprised that Xiaomi matches Apples speed even with the newest release, its not diminishing Xiaomis success.
I enjoy it because of progress, not because of Apple.
I assume the M6 will take the crown back and then a few months later Intel/AMD will release a new chip and take that crown back again. That's the state of the world we used to expect, but it's a state that has been missing ever since the release of the M1 in 2020 until Intel finally caught up again this year.
That, or I can’t read.
There are emails unearthed in various lawsuits where you can read Bill Gates screaming at his subordinates: "why the hell can't our partners like Sony and Creative create a similar device? Give them all, give them early access to everything, work with them". In the end MS felt compelled to make their own.
I yield the floor to no one when it comes to pessimism, but that's incredible.
Fab capacity is being bought online; there’s just lead time.
Noticeably greater intelligence is being achieved at the same number of parameters (see: Qwen3.8).
I think the future will be bright, it might be a matter of time. And for tinkers, a used Epyc + DDR4 server can be great fun and epic value.
Basically everyone that makes memory is building new fabs, meanwhile I'm not sure how much longer AI datacenter demand for ram will last. I think the decrease in AI ram demand and the new fabs will likely coincide leading to a collapse in pricing.
That is, of course, assuming the memory manufacturers don't pull their favorite trick and collude.
Could you provide more details about the Epyc + DDR4 server?
The above is a standard project management problem. We do this for lots of industry all the time. There is every reason to think you can get a new factory running in 5 years.
Note that I said 1 factory above. Some of the special machines we don't have the ability to make them fast enough to do 2 (I don't know the real number!) new factories in 5 years. Existing factories are using most of the special machine capacity to replace machines that wore out on the way - this can be corrected as well, but it adds another year and the expenses are much larger. Realistically though 1 new factory is likely enough.
https://en.wikipedia.org/wiki/ELIZA_effect
It turns out the limiting factor isn't how sophisticated algorithms are, it's how gullible humans are.
No AI would pass this test with experienced judges.
Though we're pretty good at sizing up a person's emotional balance/maturity and competence at familiar tasks. So maybe have an old blacksmith watch the AI/robot interact with horse owners for a while, then shoe their horses, and see how well it does.
https://commission.europa.eu/news-and-media/news/safer-and-m...
In contrast, humans tend to paste me the same barely-relevant macro over and over, no matter how much time I spend explaining my issue.
https://www.psychologytoday.com/ca/blog/the-digital-self/202...
And the "popularized" version is faulty also since it uses an ideal, abstract human judge (like the "spheroidal economic agent").
But if you want to add declinations to the said popularized image of the Turing test, you may add Maxim Lott's IQ tests at trackingai.org . Between the end of 2024 and the beginning of 2025 LLMs reached an equivalent IQ of 100, for example.
I think there are elements showing lowering of performance and expectation.
I kinda have to link it now, so uhh here's a random PDF: https://www.hec.edu/sites/default/files/documents/Computing%...
ELIZA beat the Turing test and then everyone forgot about it. Humans are just really terrible at recognising robots.
https://en.wikipedia.org/wiki/Turing_test
I've tried [1] and I almost 100% detect which is the AI. I really want to convince myself I have failed, does anyone know of a better site/resource for this?
I know it might be moving goalposts but I would consider AI to have passed in a well and truly undisputed manner when [2] is resolved.
But in a more practical sense, if AI can impersonate humans so well today then why are state of the art frontier models so obviously AI when they create PRs, commit messages, documentation, etc. Are the companies deliberately making them unnatural?
[1] https://turingtest.live/
[2] https://www.metaculus.com/questions/11861/date-when-ai-passe...
Not sure where you're quoting from but if it's the metaculus question comments, many of them are from 2023. The consensus is it will resolve in 2029. I believe it will not resolve before 2035.
Is there anything better now though?
All I see from AI, is an amplification of the enshittification of the internet.
And people being even more alone.
- Extreme poverty has dropped from 30% to under 10% globally. - Child mortality rates have dropped in half - Internet access has exploded from 10% to 70% - Solar energy costs have dropped 90% - Cancer death rates have declined by 30%
All of these massive improvements in less than 30 years.
While there certainly are issues to solve, and if you simply follow journalism you may think the world is worse off, but for many, their lives have been significantly improved.
Evidence:
* https://arxiv.org/abs/2402.09809
* https://phys.org/news/2016-12-mobile-money-access-percent-ke...
* https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3893351
This being HN, I hasten to add they also have massive downsides, we're all doomed, nobody programs the right way anymore, those poor people just think compute & AI are improving their lives, etc, etc.
That's one big plus.
These times are exciting and rough seas make good sailors. Find your path forward.
I'd rather go to the library and read a book.
I'm writing the best music of my life, realizing games and art projects I never had time for, and writing higher quality software in addition to dramatically more of it. Who has time for pablum?
How exactly are you using these tools that you have that experience?
"Find your path forward!" he shouted with glee, as he ran toward the cliff.
E.g how is the perf/$ vs Wildcat lake
So, on the mini the RAM upgrade runs at 25$ per GB on all tiers, the same as the Studio therefore the upgrade to 512 will probably cost 6400$.
The fully maxed out Apple Studio then will be 24699$. It's 17199$ if you don't upgrade the storage(1TB).
Nevertheless I itch to have one :)
EDIT: or buy AAPL. If I had bought Apple stock instead of buying a Mac LC II in 1992, then I would have about $2 million in Apple stock.
Unless you need privacy for your inference this instant, paying for credits can get 80 to 90 percent of people everything they need.
Of course if you do need that privacy, then forking the $25K over to Apple is a no brainer.
Put together a similar build with a couple of rtx 6000 Ada cards and Apple's price tag suddenly looks pretty damn reasonable
Interesting times, to say the least!
"According to reports from Bloomberg, Apple will be skipping its M6 Pro, M6 Max, and M6 Ultra chips to accelerate development of the M7 chip. That means the only chip to be released from the M6 family will be the base M6.
The reason for this break with tradition: AI. Apple had been planning major neural-processing upgrades for the M7 family and ultimately decided those improvements were important enough to justify accelerating the next generation rather than completing the M6 lineup." https://9to5mac.com/2026/08/08/apple-m7-chip-heres-why-it-ma...
I'd skip M5 and M6 chips for LLM work and wait for a year for M7.
Please say more? Is it because it is a one-time cost, unlike a recurring subscription of Claude/Codex?
Also, a lot of companies are looking at how to run capable models locally to cut some of their (massive) cloud AI bills. An easy answer is worth a lot to them.
For me personally, not quite that valuable yet, but I think it's getting there quickly. Deepseek V4 Flash massively increased the value of local AI to me, to the point where it's displaced most of my Claude Code usage, its upcoming vision enabled version should bump it further, and it's only going to get better from there.
It's a lot faster, but a lot of it is also feeling free to discuss things I wouldn't be comfortable sending to Claude, with the idea that that info is now theirs in perpetuity. I got my genome fully sequenced recently (it's cheap now!), and I get a battery of blood tests every year. Wouldn't do processing on any of that with Claude, but local AI? Totally great.
And if I was running a company with a large cloud AI bill, I'd probably buy a wheelbarrow full of these macs. Cheaper, but also a more solid/predictable base to build on.
If Apple didn't sold these things they wouldn't make them but, also the level of marketing that Apple is talking about for AI is basically the new group they need to capture because the ones I just listed are already buying Macs and or easily to motivate with the other obvious CPU / GPU performance upgrades for code compilation, faster memory and video transcoding.
to answer your question : looking at the aftermarket availability of Apple's prior best and brightest : practically no one buys them.
"people here buy them" , well, 'here' is one of the most affluent groups of people in the world.
They're available as movie and television set pieces (undoubtedly disappearing into the home of someone close to the staff post-production), and for administrative/boss types that can slip the cost into a ledger somewhere that few will ever see.
It has been a hobby of mine every few years to check out the apple site and see how big I can option a machine. My record was when I was in high school years ago and was able to option some pro studio-ish apple desktop thing to like 61,000 usd out the door.
For one thing, you can’t tell from a movie what the specs are. A $999 Mac Studio looks exactly the same as a $20,000 one.
For another, Apple updates the industrial design on their products so rarely, a 6-year-old Mac, iMac or MacBook also looks nearly indistinguishable from a brand-new one.
That may not be many people, but there certainly will be some people who want to do that, and are willing to pay big bucks to do so.
Its cheaper than Nvidia AI hardware.
It’s when self hosting and local hosting was the norm, and why it’s also starting to come back.
There will be workloads that can never touch a public cloud, and for it solutions like this are an option.
I've often felt there is tremendous value locked up in underutilized old computers. It would be interesting to see Apple in 3 years offering compute as a service using lease returns (or more likely, partnering with someone else to operate it (perhaps exclusively in secondary markets like China or India, to address political demands for local siting or jobs). Apple is in the best position to work around or even gap-fix older software/hardware limitations in a controlled environment, and now they can do so without cannibalizing new hardware sales.
not getting on that bandwagon but wasn't that not the most demanding game as its a just a nonstop cutscene.
Which one is it you can run local models on? I suppose the NPU only.
I wasn't able to debug network errors (restartin my Mac worked), Metal was missing low level disassembly / debugging tools (there is some hard to use UI), but the worst thing was the inflexible windowing system.
Even getting all the window handles on all screens/desktops with their titles and programs is impossible.
I just decided that I move to Omarchy 4 (basically Hyperland + QuickShell) + NVIDIA GPU, and I already was able to customize it more than my Mac in years.
I will miss Apple's hardware for sure, but not MacOS and the missing hardware documentation
For example when using PyTorch I wanted to try to speed up my NN kernel by 2x by just using half precision and haven't noticed any speedup at all. Also I was missing the easy to use GNU tools that had to be mixed with Apple's tools.
I loved using Arc browser as well, and I'm missing it, but I guess I will do without it somehow (Chrome's vertical tabs are just not the same).
My main program missing from going back to Linux was ChatGPT Desktop, but now it's there.
I just checked out Hammerspoon, I'm happy for you that you wrote it, and looks great, but it has the same problem that I had: for security reasons Apple stopped allowing the window APIs to get all important information on other workspaces. You can only do it with Accessibility API. I was trying to fight with it but have up.
It’s a bit depressing because it means that if I ever feel forced to switch my daily driver, it won’t come without a dump truck load of friction, frustration, and lost productivity, which I’ve validated by using the various Linux desktops on secondary machines.
I ordered an ASUS Zephyrus G16 with 5090 NVIDIA card + 1.9kg (quite an overkill, and I know that I will have to limit power output), but hasn't arrived yet.
But what's fun is that I love QML+QuickShell with its hot reloading, Hyprland with its Lua support.
With AI nowdays it's just so easy to do deep UI changes that wasn't possible a year ago.
In US its $4000 upgade so $25 for 1GB.
Also:
> 512GB memory option for M5 Ultra coming late October
The problem I'm having now is that no models are targeting RAM of that size. Everything is either much smaller, targeting laptops, or much larger, targeting hardware well out of reach of enthusiasts.
Please, AI people, start making models targeting 128GB machines again. The last interesting one was Qwen 3.5 122B.
Should bench better than Opus 4.7.
You might expect the M5 Ultra to produce 50 t/s from Qwen 3.8 27B with a good context length.
I plan on maximizing my residual student benefits, and taking advantage of education pricing.
I edit 4K ProRes and H.265 footage, sometimes with multicam (up to 4 streams) and color adjustments, titles, etc. It's only after stacking 3-5 effects before things can stutter, really.
Or if you try doing something CPU-intense in the background _while_ running some heavy creative software. I just don't do that.
That's wild!
The prices are ridiculous though. I may just keep rolling with my Windows 10 setup.
Just amazing engineering push, the competition got the message and we benefit.
This may depend on the size of your library. I tried installing Jellyfin on a Synology NAS, which runs Plex just fine, and it ran so poorly it was basically unusable. It “worked”, but it was painful.
Also: "a staggering 1.2TB/s of unified memory bandwidth" -- yay, the GPU has reached the year 2020! (I'm a bit bitter that my M4 Max is near useless for local LLMs because of its low memory bandwidth.)
I can't believe that Apple still comes with this bullshit like 32 GBs is a lot. It's a lot for video memory - vRAM, but not RAM.
Is this a joke?
>Additionally, M5 Ultra features a massive amount of high-bandwidth unified memory, up to 512GB
Now we're talking. But at what cost?
256 memory gets you to like 11k. So like 15-20k.
"M6 also introduces a Dual 16-core Neural Engine, providing up to 2x the peak compute over previous generations to make on-device AI workflows run even faster" .
A m4 mac mini is better than al of these per dollar, msrp adjusted.
Hopefully by the end of the decade China figures out manufacturing at scale and fixes this.
Apple never needed to participate in the AI race to zero. Because they were already at the finish line years ago building their own chips that can run large >100B parameter AI models locally.
It's possible that they're working on their own LLM that's going to work very well on their chips, and possibly outperform anything out there when they do release it.
10 years ago 32GB ram laptops sounded too much. 8 was enough. These days even I would get that much ram since it’s soldered. 64GB is higher end.
In a few years we should see such high end hardware commonplace. Working with a local LLM to get work done is the ideal way to go which has mostly hardware limitation as of now that gets solved in due time.
Ten years ago I got 64gb of ram in my laptop, same as I have now. I bought both for business and personal use. System ram capacity hasn't changed much in 10 years.
It makes me curious how old you were 10 years ago.
We were definitely outliers that long ago. I put 64 GB in a MacBook Pro back in 2019, and that was (a) overkill for everything I ever ran on that machine, and (b) stupidly expensive by 2019 standards (albeit almost affordable by 2026 standards)
The iPhone 15 was almost entirely marketed based upon AI (I would say fraudulently so, advertising features they still haven't delivered), and a huge portion of the OS work was on local AI or AI integration.
And for that matter Apple has been dumping enormous sums into their own AI development. Their failure to have a lot to show for it doesn't void the fact that they tried really, really hard.
It's bizarre how often this "Apple sat on the sidelines and let the AI people fight...so smart!" narrative appears on HN. Apple hasn't gone down the path of spending hundreds of billions on nvidia GPU data centres, but they absolutely tried really hard to matter in AI.
Yep, Siri AI; they’re doing it in public.
But you get a generic computer and much more RAM.
And you lose a couple of organs.
It would take Apple one or two engineers to make Linux life much easier on macs. But Linux is outside their walled garden so it's ignored.
My Mac Mini is strictly a headless server for llama.cpp.
I use a Linux workstation.
If I were limited to use Mac hardware , I would install Linux in VMware Fusion and work from there.
My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.
In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.
* the iPhone * the iPad * apple watch * airpods * unified memory laptops and computers
Those are all products that either created a category or changed that industry.
I mean this is not nvidia based right? It's all custom? So we can use it under Asahi perhaps?
I want to get something for my company to run local models, wondering what would be a good option.
Linux runs very well in a VM on macOS. There are many good options for this, some free and open source (QEMU, UTM, Lima, Colima), some proprietary (VMware Fusion, Parallels).
But Linux in a VM doesn't get access to the real GPU, so model performance is limited. Those running on the CPU perform well, and those needing the GPU don't.
However, macOS on M-series macs is excellent for local models. (Maybe not as excellent as a box full of the best nVidia GPUs, but still excellent).
So if you're getting Apple hardware, like Linux, and want to run all of it locally, a fine setup for a machine to run local models, with agentic characteristics:
- macOS running one of the many local model runners. I used to use Ollama and Whisper, and now use llama.cpp instead of Ollama. Others use LM Studio, oMLX, etc. Provide HTTP endpoints to access the models.
- Linux in a VM for overall control and orchestration, with standard VM settings, and bridged networking so it appears as its own machine on your network. Also, in here provide a robust shared file server for shared state. Use this VM as your desktop and primary access to the machine, if you like Linux.
- Linux in a VM to launch ephemeral, volatile containers, with the containers using a memory-only tmpfs overlay on top of a read-only Linux filesystem in a VM disk image, with tools in this filesystem. Alternatively, a writable Linux filesystem in a VM disk image, with disk buffering set to use macOS host buffering and discard fsync requests. These settings optimise for container disk performance for data that's only ephemeral which will be deleted soon or on system shutdown. (You can combined both VMs, but need to use two VM disks to get equivalent behaviour, and be careful about VM disk configuration of the two disks.)
- Containers spawned within that second Linux VM can be spawned very quickly and run quickly, so are ideal for LLM agents that need a quick sandbox. These sandboxes generally run faster than a macOS sandbox, despite being on the same machine with VM overhead, because Linux is faster at some things. Teach the LLMs to store files and memories they want to keep in the shared file server.
I just want my butt ugly repairable beast machine to do the same trick. Why is my ram not unified? I have an iGPU in my server, but it can't access the 64 GB ram directly or something? It's on the CP right? Why did only Apple go for this architecture? So many questions...
and 512GB is so 2025.
Give us 1TB version. Where is the competitive spirit?
To be fair, where is the competitor at this form factor?
Can we please kill the xcode. It is worst pile of garbage I have to use just to develop ios app.
It's the Hardware, Software and Services in combination. None would work without the other (to reach the scale apple is)
More specifically, they're a hardware dongle company first
FWIW, Ollama, LM Studio and Lemonade (and oMLX) also wrap Apple's MLX framework.