Back to the notes

Greg Kroah-Hartman on why LLM-generated patches terrify kernel maintainers

The Linux kernel maintainer explains why code review is about to get much harder when everyone brings AI-written patches.

01equipment.jpg
Alice Wondrak-Biel / Wikimedia Commons. Resized and converted to WebP. Public domain

Greg Kroah-Hartman, the Linux kernel stable branch maintainer, gave a talk about security in the LLM age that cuts through the optimism about AI coding assistants. His concern is not that LLMs write bad code. It is that they make code review exponentially harder. The kernel maintainers already deal with thousands of patches every release cycle. Each one needs human review because subtle bugs in kernel code can brick machines or open privilege escalation holes. Now imagine those patches arrive faster, in higher volume, and half of them look correct at first glance but fail in edge cases the LLM did not consider. Kroah-Hartman’s point is that review effort does not scale linearly with submission count. If submissions double but quality drops by twenty percent, review time quadruples. The bottleneck shifts from writing code to verifying it. That is already the case in most open source projects, and LLMs make it worse.

The detector problem

Some suggest using LLMs to detect LLM-generated patches. Kroah-Hartman is skeptical. Adversarial examples exist. A contributor who wants to slip something past review can iterate their prompt until the patch passes both human heuristics and automated checks. Defence does not scale the same way as offence in this game. What stuck with me is his framing of trust as the real currency. Kernel maintainers trust certain contributors because those people have a track record. LLMs have no track record. They have outputs. If I submit a hundred AI-assisted patches and ninety-nine are fine, the hundredth still gets the same review effort as if I wrote it by hand. Trust is not transferable to the model. This applies beyond kernel work. Any codebase where review is already the constraint will feel this. I have seen it in security detection engineering. A rule that looks correct in isolation but fails on live traffic is worse than no rule at all. LLMs are great at generating the former.


Source: Greg Kroah-Hartman – Security in the LLM Age [video]

Back to all notes

Behind the notes

Vikrant
Sharma.

Artificial Intelligence Engineer intern at Voxon Photonics in Adelaide. Studying a Master of Information and Communications Technology at UniSC, with a focus on data, machine learning and security.

Meet the person behind the work