Someone built a C compiler with an LLM for $100 and it boots Linux
A solo dev used an LLM as a pair programmer to write a C compiler from scratch. It compiles itself and boots the Linux kernel. Total AI budget: $100.
From the notebook
50 notes on this topic.
A solo dev used an LLM as a pair programmer to write a C compiler from scratch. It compiles itself and boots the Linux kernel. Total AI budget: $100.
Someone built a lab where you can test tiny language models that run entirely client-side. No API calls, no backend, just WebGPU and a lot of patience.
A new paper shows the system prompt is not what makes models say 'as a language model'. It is the chat template wrapper that turns raw output into self-aware hedging.
Someone built a tool that generates fonts where every tokenizer chunk takes the same visual width. Suddenly prompt length is readable.
A live benchmark that ranks models by performance per dollar spent, not just raw scores. Turns out the best model depends on your budget.
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
A Reddit user discovered that giving a model explicit instructions to behave like JEV (a fictional expert persona) produces cleaner outputs, even when the model has no training data about JEV.
Pirate Face archives open-weight language models that companies try to delete. Turns out model takedowns happen more often than I thought.
Someone built a tool to make AMD GPUs run local LLMs without the usual driver nightmares. The benchmarks show Vulkan beating HIP by 20 percent.
The Matasano founder uses Claude to draft, then rewrites everything by hand. No copy-paste. The LLM is a research assistant that never ships.
Yayster is an LLM that runs locally in Emacs buffers. No API calls, no cloud, just a model sitting in your editor watching you code.
A developer used an LLM to translate 1993 Amiga assembly into modern game engine code. The surprising bit is how much of it actually worked.
Someone fed an LLM a codebase to remember, then asked questions. The model started catching bugs the author missed. Not by design, by accident.
Running a model on your laptop does not mean it is broken. The context window might be.
A quiz shows that even technical readers cannot spot which GPT-4 response has a cryptographic watermark baked in. The detection gap is real.
An LLM reverse-engineered a 2008 HP laser printer's protocol and generated a working macOS driver. No human debugging required.
Researchers built an LLM using exclusively elementary-level text. The results challenge assumptions about what makes a model fluent.
A new interface shows chat threads as directed acyclic graphs where you can rewrite nodes and re-run paths. Fixes the branching problem most chat UIs ignore.
A new tool checks if ChatGPT's GPU kernels are actually correct before you run them in production.
Needle2 runs on phones, smartwatches, and Raspberry Pis. The entire model is smaller than a single photo.
Why learning communities ban code assistants while enterprises mandate them.
Someone is flooding the National Vulnerability Database with fake SQLite vulnerabilities written by language models, and it is breaking actual security work.
Someone just released a minimal codebase for supervised fine-tuning, direct preference optimisation, and group relative policy optimisation on consumer hardware.
A security firm ran Claude through GlobaLeaks' codebase and found medium-severity bugs for seventy-six dollars per finding. The question is whether a human would have caught the same issues faster.
Manifest deprecated their LLM router after six months. The reason is not what I expected.
A researcher deployed a fake human verification page that only AI crawlers would fall for. The logs are filling up.
OpenAI open-sourced the internal security guidelines they used when building Codex. Turns out threat modelling an AI code generator is different from threat modelling a database.
Research engineer roles at the big labs filter for production ML skills first, paper count second.
A Debian general resolution proposes banning LLM-generated patches. The reasoning is direct: you cannot verify the training data's licensing.
New research shows children anthropomorphise LLMs at much higher rates than adults, which changes how we should think about guardrails.
Antares models are 1B to 8B parameters, fine-tuned on security tasks, and Apache 2.0 licensed. This is not another rebranded Llama wrapper.
A Chinese LLM patched critical security bugs in a codebase where OpenAI's models and Anthropic's Claude refused to engage. The refusal problem is real.
Gwern argues personalised LLM assistants could filter spam, draft replies, and catch your mistakes before you send them. The privacy trade-off is obvious.
Anthropic's Model Context Protocol promised to standardise how AI agents talk to tools. A new audit shows most implementations ship with auth disabled by default.
Iroh built a system that splits LLM inference across volunteer nodes. The networking stack handles dropouts mid-inference. Wild.
The moment you realise you are debugging prompt chains instead of writing code.
Dan Luu measured the same coding task twenty times with the same prompt. The variance in output quality was higher than the difference between model versions.
Sidenote lets you comment on a rendered blog post, then an LLM writes the actual markdown diff. No forking, no pull requests.
vLLM's Micro-Agent proves that three coordinated 8B models can outperform a single frontier model on complex reasoning tasks.
A GitHub repo called Bash4LLM+ does what Python libraries do in thousands of lines, using only shell builtins and curl.
A city government announced a locally trained language model. Turns out it was two existing models stitched together with the weights renamed.
A CLI tool that flattens your data science repo into one massive prompt. Smart filtering meets the 200K token era.
An enthusiast loaded a 1T-parameter model into 768GB of Intel Optane DIMMs and got 4 tokens per second on a single GPU. Slow, but it worked.
A developer catalogues the tell-tale signs of AI-generated code. The patterns are obvious once you see them.
New research shows agents generating backend code slowly drop requirements like authentication checks. The longer the generation, the worse the decay.
Models.dev is an open-source database that tracks pricing, context windows, and rate limits across every major LLM provider. No more tab-sprawl to compare GPT-4 versus Claude costs.
Steering vectors let you nudge a model's behaviour without retraining. They fell out of favour when newer models stopped responding to them. DeepSeek-V4-Flash brought them back.
A new LLM observability tool runs without PostgreSQL or Redis. That is not a feature list, that is an architecture decision.
Mythos discovered a vulnerability that was already documented in the data it was trained on. The industry is calling this autonomous discovery.
Mythos, an autonomous security agent, caught a buffer overflow in curl that human auditors missed. The tooling works.