Quick Introduction We have all been on forums, chats, reddit, discord, youtube, or somewhere and heard “Oh! Model XYZ is AMAZEBALLZ!zomgwtfbbq” then downloaded it (or more likely, some quantized form of it) and said “eww… This sucks!” This post is going to be a rather technical series of experiments to demonstrate the impact of implementation-specific hazards with inference. I will be using the term “reference implementation” to describe the lab that published and offers first-party hosting of ...
Having used these extensively, I can say with confidence that they feel exactly as dumb as they are.
Do you mind sharing what models you used and what your experience was? In my opinion the Qwen 3.6 models or maybe the 3.5 were the first local models that were actually useful, but I don’t have that much experience.
Qwen is really good. For coding, it’s the best local model I’ve found.