llm A field note by Vikrant Sharma
Your local LLM is smarter than you think
Running a model on your laptop does not mean it is broken. The context window might be.
I run Llama models locally on a 32GB M1 and they regularly give me answers that feel worse than ChatGPT. Not vague worse. Confidently wrong worse. Turns out the problem is not the weights.
The thread walks through a common trap. You load a 7B or 13B parameter model with 4-bit quantisation to fit in GPU memory. It works. You prompt it. The answer is shallow or off-topic. You assume the model is too small or the quantisation broke it.
The actual issue is the context window. Most local setups default to 2048 or 4096 tokens. That is fine for a chatbot that answers one question. It breaks when you paste in documentation, a long email thread, or a CSV sample and ask the model to reason about it. The model sees the first chunk, loses the rest, and guesses.
OpenAI and Anthropic models handle 128k to 200k token windows out of the box. Your local runner might be capping you at 8k unless you explicitly configure it higher. The model has the capacity. The serving layer is cutting it off.
I checked my own llama.cpp setup. Context length was set to 4096. I bumped it to 16384 and re-ran a prompt that previously failed. Same model, same weights, coherent answer this time.
The fix is boring. Read the flags for whatever runner you use. Set , ctx-size or the equivalent to match what the model was trained on. Most recent models support at least 32k. If you are on a Mac with unified memory, you can push higher without running out of VRAM.
Local models are not dumb. The defaults are.