Training transformers in seconds, not hours
Someone built a 22KiB transformer that trains in 13 seconds. The trick is not the model size, it is how they squeezed PyTorch out of the picture.
From the notebook
41 notes on this topic.
Someone built a 22KiB transformer that trains in 13 seconds. The trick is not the model size, it is how they squeezed PyTorch out of the picture.
Someone built a lab where you can test tiny language models that run entirely client-side. No API calls, no backend, just WebGPU and a lot of patience.
A new paper shows the system prompt is not what makes models say 'as a language model'. It is the chat template wrapper that turns raw output into self-aware hedging.
Project Suncatcher is Google's proposal to deploy GPU clusters in low Earth orbit. The pitch is simple: space is cold, solar power is free, and latency to ground stations is acceptable for batch jobs.
A live benchmark that ranks models by performance per dollar spent, not just raw scores. Turns out the best model depends on your budget.
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
A Reddit user discovered that giving a model explicit instructions to behave like JEV (a fictional expert persona) produces cleaner outputs, even when the model has no training data about JEV.
Pirate Face archives open-weight language models that companies try to delete. Turns out model takedowns happen more often than I thought.
A graphics engineer got neural texture maps working with ES optimisation. No gradients, no autodiff, just mutation and survival.
A new transformer variant ditches the attention loop entirely. Faster inference, cleaner scaling, same performance.
A live playground for experimenting with domain-specific ML languages, no install required.
Someone repurposed their home security setup to log every bird species that flies past. The microphone was already there.
The GPU monopoly now owns the model zoo. This changes who controls open-weight AI.
A startup ditched everything GPUs do well and made a chip that runs one architecture 20 times faster.
Running a model on your laptop does not mean it is broken. The context window might be.
An LLM reverse-engineered a 2008 HP laser printer's protocol and generated a working macOS driver. No human debugging required.
Three years after first commit, the project that made local LLMs possible ships a stable release.
Researchers built an LLM using exclusively elementary-level text. The results challenge assumptions about what makes a model fluent.
The company that sells license plate readers to 5000 police departments just announced privacy controls. Timing raises questions.
Researchers figured out how to steal the internal reasoning steps from proprietary models like GPT-4o and Claude through API timing attacks.
Needle2 runs on phones, smartwatches, and Raspberry Pis. The entire model is smaller than a single photo.
Why my intrusion detection capstone needed an Isolation Forest, a Random Forest and an autoencoder to reach 0.95 F1 on real attack traffic.
Someone just released a minimal codebase for supervised fine-tuning, direct preference optimisation, and group relative policy optimisation on consumer hardware.
Research engineer roles at the big labs filter for production ML skills first, paper count second.
Antares models are 1B to 8B parameters, fine-tuned on security tasks, and Apache 2.0 licensed. This is not another rebranded Llama wrapper.
A developer trained MNIST digit recognition using only SQL queries. No Python. No frameworks. Just recursive CTEs and window functions.
Google announced Gemini 2.5 Flash will be discontinued in February 2027. Developers who built on it are scrambling.
The moment you realise you are debugging prompt chains instead of writing code.
Physicists found that the standard mean-field approximation for neural networks breaks down when you look at correlations between neurons.
Deepset's Haystack framework caught my eye because it treats retrieval-augmented generation as a data pipeline problem, not a chatbot wrapper problem.
Researchers trained particles to form complex shapes without central control. Each particle runs the same neural network, learns local rules, and the swarm organises itself.
WebGPU makes training tiny neural networks that grow patterns possible in real-time, no server required.
A GitHub repo sketches how Europe could pool scattered compute across universities and research labs to train a GPT-4 class model without buying a new datacenter.
Researchers found that transformer attention mechanisms lack the executive control functions that let human brains manage working memory. The models can retrieve information, but they cannot suppress irrelevant context.
A CLI tool that flattens your data science repo into one massive prompt. Smart filtering meets the 200K token era.
The Norwegian University of Science and Technology is running LLM workloads on Chinese storage hardware. The performance numbers are interesting.
A bill requiring platforms to moderate content encouraging violence against Jewish communities turns moderation from a platform choice into a legal obligation.
Linus Torvalds says automated vulnerability scanners have turned the kernel security mailing list into noise. The tools work, the signal-to-noise ratio does not.
Steering vectors let you nudge a model's behaviour without retraining. They fell out of favour when newer models stopped responding to them. DeepSeek-V4-Flash brought them back.
A programmer rewrote PyTorch's transformer architecture using Rust and abstract algebra. The result is dense but fast.
Mythos discovered a vulnerability that was already documented in the data it was trained on. The industry is calling this autonomous discovery.