

8·
15 days agoContext limit is not really a problem on local models. Qwen3.6 can do up to ~256k tokens, it’s not that far from what things like cursor uses. grok in cursor have exactly 256k tokens for example. Also, you can use opencode with custom config, where you need to set trimming close to that number.
That being said, quality wise it still kinda shit and looses to paid models. I think you need something like GLM 5.2 to compete, which needs 228gb at lowest 1bit quant (not including context size, which can be up to 1m tokens, so you can round up requirements for RAM up to 256gb), so yeah, the only thing that limits you locally — the fact that “AI” companies bought all supply of focking ram and we can’t afford any.
What doesn’t at this point? All repos on github contaminated at this point. In couple of years border between code created by humans and code created by llm’s will be basically non existent.
The end is near. git clone until it’s too late!