Rendered at 14:17:06 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
apimade 5 hours ago [-]
I guess the question is; is this the late-cycle cash-out (aka harvest pricing) akin to Sun Microsystems at the dot-com peak -- or is it a repeat of the crypto pricing hijinks we've already seen from Nvidia, where they're just exploiting the lack of supply?
That's a 90% uplift since the original pricing, in a market that is already showing signs of seizing.
With Apple offering leasing options, CXMT knee-capping Samsung, all of the big tech players on a run to outspend on CapEx by the end of the year..
We're at a point where local models are exceptionally capable, model-on-silicon dies like Taalas (recently acquired by AMD) may be cutting inference cost substantially for the 90% of work we do day-to-day (similarly Alibaba's T-Head division with open-model-forward inference chips being produced domestically in China).
If we shifted all of the design/planning to cloud models like Fable 5/Sol 5.6 Ultra, and day-to-day operational inference to these chips -- it's quite likely we'll squash usage to single-digit percentages of what we're currently using. But we should also expect the model providers to take a similar approach.
In any case -- I'm keen on demand destruction, both as a consumer and someone without skin in the game.
saidnooneever 3 hours ago [-]
i think maybe by pricing up they want to make a reliance for people on their services rather than developing running their own. if these devices are out of reach for many it means less innovation and then less competition. nvidia is now tightly coupled financially to business who offer services that owners of such devices might replicate without using services of their partners.
areoform 6 hours ago [-]
> RTX Pro 6000 Blackwell has 96GB of GDDR7 VRAM
A mac studio with 96GB unified memory costs, $5,299.00.
As others have stated, the Blackwell runs circles around the Mac performance-wise. I haven’t run the numbers for this specific comparison, but in some cases the nvidia solution is also more efficient (tokens/W). Some people are willing to pay that premium for the above, just like they’d pay more for a 4-door sedan from BMW as compared to a Toyota Corolla.
JackSlateur 5 hours ago [-]
LLMs can re-write and cross-translate software really well.
Is that was true, claude would be written in Rust;
androiddrew 2 hours ago [-]
Buy yourself 6 AMD r9700's an older 5965wx thread ripper with mob, 128GB of ddr4 and run deepseek v4 flash lossles. Then pocket the other $6000
lolive 6 hours ago [-]
Just when Cyberpunk gets almost bug-free, the price for a decent Gpu to run it smoothly is skyrocketing. #doooh
saidnooneever 3 hours ago [-]
these are not gaming gpu, but on the other hand those are also completely unaffordable for what they offer.
That's a 90% uplift since the original pricing, in a market that is already showing signs of seizing.
With Apple offering leasing options, CXMT knee-capping Samsung, all of the big tech players on a run to outspend on CapEx by the end of the year..
We're at a point where local models are exceptionally capable, model-on-silicon dies like Taalas (recently acquired by AMD) may be cutting inference cost substantially for the 90% of work we do day-to-day (similarly Alibaba's T-Head division with open-model-forward inference chips being produced domestically in China).
If we shifted all of the design/planning to cloud models like Fable 5/Sol 5.6 Ultra, and day-to-day operational inference to these chips -- it's quite likely we'll squash usage to single-digit percentages of what we're currently using. But we should also expect the model providers to take a similar approach.
In any case -- I'm keen on demand destruction, both as a consumer and someone without skin in the game.
https://www.apple.com/shop/buy-mac/mac-studio/m3-ultra-chip-...
LLMs can re-write and cross-translate software really well. Why does CUDA still have a $11k price premium?
They're both good value (or crazy expensive) depending on how you look at it.
It depends how you price the ability to run a particular model at all, vs run the model quickly and serve several parallel streams.
Maybe Apple should bring OS X Server back, with a rack like drawer.
CUDA ecosystem is more than running PyTorch.
the last mac pro was rackable, and they discontinued it anyway.