Posts for: #AI

Putting a Code-Review LLM in the CI/CD Pipeline

Putting a Code-Review LLM in the CI/CD Pipeline

What actually happens when you wire a self-hosted model into CI to review every push: a story of checklists that reviewed nothing, reasoning that buried its own findings, the surprisingly scientific business of measuring whether your robot reviewer is any good – and the winner’s first three days on the job.

[Read more]

WalledClaude: My Claude Container, Improved

Back in May I wrote up my setup for running Claude Code inside a Docker container, so an AI agent can work on my code without also being able to work on my workstation. The nicest thing that can happen to a post like that has now happened: someone more diligent than me read it, took the idea seriously, and built it properly.

That someone is Justin, and the result is WalledClaude. (He also asked, very courteously, whether he could license his work derived from my blog post under an MIT licence: of course the answer was yes, and I’m delighted – that’s how this is all supposed to work.)

My original post was mostly about convenience with a security boundary attached: get the agent off my filesystem, keep file ownership sane, contain the blast radius. Justin kept the convenience and then did the hardening I hand-waved past. From the repo:

  • Every Linux kernel capability dropped (--cap-drop=ALL) and no-new-privileges set, so the container can’t escalate its way to anything interesting.
  • Resource limits on CPU and memory, so a runaway agent is an annoyance rather than an outage.
  • The big one: a Squid egress proxy in front of the container, so outbound network access is controlled at the domain level – by default only the Anthropic endpoints the CLI actually needs. My version quietly trusted the container’s network access; his assumes the agent might phone somewhere it shouldn’t, which is the correct assumption.
  • Sandboxed and regular Claude data kept apart in a separate ~/.walledclaude/ tree, so the jailed agent doesn’t share state with anything outside the jail.
  • A README with an actual threat model and a “things to consider” section, including honest statements about what it does not protect against. Documentation of the limitations is the part most projects skip, and it’s the part I’d recommend reading first.

He also spotted that docker run --user takes a UID/GID directly, making my pass-the-environment-variables-to-the-entrypoint dance unnecessary.

Drafted by Claude, edited by me. I write these things mostly for my own memory – but it turns out sometimes they get compiled into other people’s better software, which is the best outcome available.

Why I Run Local Models

Why I Run Local Models

Yet again, Claude is down. Yet again, my local models just keep working. A case study in why the complexity of self-hosted AI is worth it.

[Read more]

AI on-demand with Kubernetes (Part 2)

AI on-demand with Kubernetes (Part 2)

Part two of the quest to deliver AI applications on-demand from Kubernetes. In this one, we’ll deploy a typical AI application, and configure it to scale up (and down to zero) on demand.

[Read more]

Inference in the Cloud with Modal

Playing with contemporary machine-learning models can demand hardware with a pretty hefty pricetag. Modal lets you do it in the cloud with a much more reasonable pricing model than the big Cloud Compute providers.

[Read more]