Back to the notes

Seven LLMs running in your browser tab, no server involved

Someone built a lab where you can test tiny language models that run entirely client-side. No API calls, no backend, just WebGPU and a lot of patience.

Automatic fan controller for server racks (38933511880)
Dilshan Jayakody from Maharagama, Sri Lanka / Wikimedia Commons. Resized and converted to WebP. CC BY-SA 2.0

The MicroLLM Lab lets you run seven different small language models inside a browser tab. No server. No API key. Just WebGPU and your GPU doing the inference locally. The models are tiny by today’s standards. The largest is 1.5 billion parameters. Most are under 500 million. They load in seconds and respond fast enough that the lag does not kill the experience. This is the kind of thing that should have existed two years ago but did not because nobody wanted to ship a worse model when ChatGPT was right there. Now that the novelty of cloud LLMs has worn off, the privacy and cost arguments for local inference are starting to win. The interface is minimal. You pick a model, type a prompt, get a response. No settings beyond temperature. No streaming tokens for drama. It feels like using a calculator, not a chatbot. WebGPU is doing the heavy lifting here. It is the browser API that finally makes GPU compute feasible without plugins or desktop apps. Safari and Chrome both support it now. Firefox is slower but catching up. The performance gap between these models and GPT-4 is obvious. The answers are shorter, less coherent, more likely to wander. But for tasks like quick classification, keyword extraction, or generating a single sentence, they are fast enough and private enough to be useful. What this proves is that you can ship a useful LLM experience without touching a server. The models are small enough to cache. The inference is fast enough to feel interactive. The privacy story writes itself. The next step is packaging these models into browser extensions or offline-first apps where the server being unavailable is not a dealbreaker. A grammar checker that works on a plane. A sentiment analyser that never phones home. A search assistant that does not log your queries. Small models running locally will not replace GPT-4. They will replace the ten thousand API calls you make every week for tasks that do not need GPT-4.


Source: MicroLLM Lab – Try 7 tiny LLM’s in the browser

Back to all notes

Behind the notes

Vikrant
Sharma.

Artificial Intelligence Engineer intern at Voxon Photonics in Adelaide. Studying a Master of Information and Communications Technology at UniSC, with a focus on data, machine learning and security.

Meet the person behind the work