Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output.

If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks.

In essence, GLM 5.2 is Opus 4.8 its little brother, at a way, WAY cheaper price.

There has been really no training on Opus models going on, really, none i tell you! /sarcasm



> GLM 5.2 Max = Opus 4.8 Max in thinking behavior

This is insane! I can't wait until technology progresses to the point we can run these things on consumer hardware!


Are there any indications that this will be possible? Consumer hardware will continue getting better but I can't see 512GB RAM in a MacBook Pro any time soon. I'm hoping linear attention techniques plus MoE will make breakthroughs in size/compression and throughput.


> but I can't see 512GB RAM in a MacBook Pro any time soon

Could totally see this being a comment from a forum in like 1994 but swap out GB for MB and MacBook Pro to whatever the popular consumer pc was at the time


Yeah but the price of RAM wasn't increasing at that point.


Well, we're probably not going to be running frontier models anytime soon, but I think the general assumption is smaller models will continue to improve until they're sufficiently good frontier models aren't needed.

There's potentially also augmentation through tools, harnesses and RAG to help boost how well they work without tons of parameters.


There will be a 1024GB unified memory MacBook Pro.


Not at a price that your average consumer can afford for a long, long time


Certainly not any time soon, but I have faith it'll happen one day.


In the last ten years laptop memory footprints have, what, doubled at the low end? Smallest MacBook Pro in 2016 was 8GB, smallest is 16GB today? Max I think has gone up 8x meanwhile, 16 to 128?

I wonder if there's a bit of a chicken-and-egg issue where there wasn't much that demanded 10x the RAM, so there wasn't much pressure to develop more or increase production to support it at consumer prices.

There's wayyyyyyy more demand for memory generally now, so assuming it's not a demand bubble that pops rapidly, I'd expect the new normal to end up at a much higher baseline. 512GB would be 4x greater than today's max, so even with the relatively slow last 10 years development pace, give it five years max?


The problem is that the situation in the RAM market might just... not go away. It's locked in for the next couple of years unless the AI market goes pop. Which it might! But if it doesn't, there's no particular reason to think that the incentives for cornering the market like OpenAI have would go away.

We might see that new normal in five years or so. We will see a new normal sooner than that if there's a run on AI because of the sudden availability of DRR fab capacity, but also we'll probably see the level of local models freeze at whatever state they've got to at that point. But an equally likely outcome is that any new DDR capacity that comes online is just immediately absorbed by frontier AI, and consumer devices stay at "just good enough" for a decade.


The new Macbook Neo is 8GB. I think that if we are lucky, the huge RAM demand right now means new factory buildouts which eventually means more supply and prices go back down, and capacity begins to go up. This level of demand was just not anticipated by anyone.


you need 8 x 96GB Blackwell or equivalent

so around US$150k which is Small/Medium-Enterprise territory already, but who knows when it will hit "reasonable" home consumer territory

I think there's hope future generations of unified memory machines may get this sort of memory availability when new fabs open in then next couple of years and then ramp up production for a few years afterwards - that makes ~2030s credible at this point, but nobody can really predict the market that far ahead


> I think there's hope future generations of unified memory machines may get this sort of memory availability

I hope you're right. This is a very exciting idea. The weights are out there. The demand is astronomical. The manufacturers just need to make it happen.


there are cheaper ways to do it. not like, consumer-cheap, but I'm setting up a rig for 80% cheaper than that.

I'm a tad worried about triggering a run on the particular hardware I'm buying though so I'll leave it vague here, but hit me up on Discord if you're curious.


Hey, very intrigued about how it can be done for cheaper. Sent a friend request to sterlind on Discord, interested if you do a write up


But at what kind of speed? We're aiming at some speed that would negate the point of even using an off-site provider.


This is quite evident for personal AI but general intelligence with current scaling laws and how model keep getting better with more number of parameters, certainly the path does not converge. Personal AI is more deprived of context today than quality of token. Having a on-system knowledge base paired with Gemma works well to large extend.


With such ridiculously long thinking traces I'm surprised max outperforms high. After all, performance falls off a hill after a certain amount of context, and long thinking traces can fill that up really quickly.


looking at the score this is rather a gemini 3.5 flash competitor, yes, for cheaper, but distance to opus and fable is as big as their price diff.


distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.


It is quite likely that the intermediate tokens don’t have ‘semantic import’[0]

There are methods like Habitual Reasoning Distillation or Inverted Reasoning Traces [1] that can help.

While there are reasons to hide the intermediate tokens from a IP protection stand point, there is also a need to hide more effective and efficient generating that doesn’t fit the R1 claims of an aha moment that has been debunked, but is a consumer expectation.

While hidden intermediate tokens do increase the difficulty, it is not a from barrier in itself, especially as they are billed, given information about their length.

[0] https://arxiv.org/abs/2504.09762v4

[1] https://arxiv.org/abs/2603.07267


Chinese distillation attacks are about as unethical as Robin Hood stealing from the rich to give to the poor. The real unethical scraping was done by Anthropic to train Claude.

To be clear, if Anthropic was using totally licensed data, I'd be sympathetic to these claims. But if you're going to pirate the world's creativity you'd better be willing to gimme dat shit for free[0].

[0] As said by Hungry Santa.


>such outrageous copyright infringement

Sarcasm, considering the source of their own training data?


Considering they called the company "Misanthropic", sarcasm is a safe bet.


Somehow, I completely overlooked that.


Narrator: it was sarcasm, indeed.


IP for me, not thee.


For Claude models at least, you can tell to just manually think in the output and it works fine. I do it reguralrly because for creative writing and summarization, they seem to believe they don't need to think at all, and get way worse results.


this helps so much. i do it too. with some of the newer frontier models its unclear if you can even turn it off in the first party chat apps. havent compared api semantics yet.


FYI: model outputs are not protected by copyright.


The companies that did copyright infringement and unethically scrapped data think that copyright infringement and unethically scrapping data is wrong and needs to be stopped.

Though only in particular situations, like when it’s done to them and not when they do it. Cause they have the power and are morally right and know better than you. And if you question this at all, well you’re a threat to American values and a supporter of the Chinese and leading to the break down of Democracy.

This isn’t a type of reasoning argument or manipulation tactic used by the rich throughout history to trick the naive and gullible masses or anything like that. Trust me, I’m rich and I’m morally right. /sarcasm


It’s been amazing to see the arc of tech people going from “evil Disney, copyright is an abomination, information wants to be free” to “OMG copyright is inviolable and AI is taking money out of Plato’s descendants’ pockets!”


> taking money out of Plato’s descendants’ pockets

Yeah, remind me - is it Plato's descendants that people are concerned about here, or is it every single author who had any work in Anna's Archive, any work published online, any work published on github, etc?

I think that people are probably upset about the harm to living people who had their work stolen by Meta and other LLM companies - regardless of license, terms of use, or any other attempted protection.


Sure, that’s the motte / bailey. Easy to point to living, starving writers who suffer grevious harm, in defense of perpetual copyright. Disney and others use literally this exact argument year after year.

I’m not even disagreeing. I’m just saying the shift in attitude about copyright in the tech space has been sudden, dramatic, and really funny. Remember “you wouldn’t steal a car”? Today’s anti-AI tech contingent are enthusiastically embracing that false equivalence that we all laughed at 20 years ago.


Having a static, immovable belief system about something like copyright that is unaffected by seismic shifts in the real world also doesn't seem very logical.

If like, Disney did a 180 overnight and bought rights from Google to scan every writer's saved work in Docs with some flimsy legal argument then a person saying "wait doesn't copyright actually protect that" would make sense. Even if you were previously upset about them suing schools for using 80 year art.


Sure. So you’re saying MPAA was right and you’ve come around?

Creative works have always been accretive. There had never been a creative work made out of whole cloth, with no debt to any previous work.

The fact your opinions about creative works change based on who’s profiting does not change that.


Reasoning models can coaxed to reason like they do in dedicated reasoning blocks, outside of those blocks: in normal parts of the response.

But Anthropic at least has openly admitted they try to detect that and interfere


Supposedly there are “jailbreaks” that expose considerably more of the thinking traces.


Simple trick: Use an agentic tool like Pi or OpenCode that allows you to switch models. First do some chats with DeepSeek or GLM who shows full thinking traces, then switch to Claude or GPT and it's more likely to show full thinking traces.


I don’t understand why there isn’t public dataset for reasoning that can be improved by humans/llms like Wikipedia (ie with auto judging contributions etc).


There is already a lot of effort to collect agent traces including reasonings, e.g. see the recent discussion: https://old.reddit.com/r/LocalLLaMA/comments/1u795pb/donate_...

We've been developing DataClaw for this: https://github.com/peteromallet/dataclaw


Did I get it wrong or the first link has dataset with 30 entries only?


For reasoning a manually-curated dataset is too small; you need to be able to automatically generate vast volumes of synthetic reasoning data with provably correct answers. That's presumably why Claude and GPT are so good at using Lean (the theorem prover), because they get fed a bunch of synthetic, verifiably correct training data.


Wikipedia is a lot of data as well but we manage to do it, no?


You can trivially leak the CoT of any current model, it's not a problem.

>outrageous copyright infringement

>unethically scrapped data

Hahahahaha




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: