TLDR⌗

I run a two-tier AI code-review process on my GitLab: a frontier model does the heavy review when a merge request opens, and a self-hosted model reviews every push for free, because the GPU is sitting in the rack anyway – and because a CI stage should not depend on a third party’s uptime. Getting the free tier to be worth anything took three redesigns and a day of genuine experimental science – and the winning configuration has now had its first few days in production, so this post gets to end with mileage instead of a promise. The short version of what I learned:

  • If you dumb a model down to a deterministic checklist, you get determinism and no reviews worth reading. Use a lint for lint-shaped work; let the model reason.
  • For review work with Qwen models, reasoning mode actively hurt. The no-thinking variant found real bugs; the thinking variant talked itself out of them, or timed out mid-rumination.
  • A single review pass is a lottery ticket. What is stable is which kinds of bug a given configuration can find; whether it finds a specific one on a given run is a coin flip. Run more passes and union the results – local GPU makes this free.
  • Build your test set from your own review history. You already have labelled defects: the ones your expensive reviewer caught last month.
  • The ops half matters as much as the model half: know what your proxy does when the backend dies, know your CI runner’s real timeout, and stream evidence into the job log because artifacts will not survive the failure modes you care about.
  • Live results, three days in: 123 findings across fifteen tickets, 48 fixed – including several genuine bugs – and the rest declined with written reasons. The disposition discipline around the reviewer turned out to matter as much as the reviewer: it converts false alarms into documentation and repeat findings into one-line replies.
  • The specifics, for the impatient: Qwen3.6-35B-A3B in Nvidia’s NVFP4 quant, on an Asus GX10, vLLM behind a LiteLLM proxy. The full run command is at the bottom of the post.

The Setup⌗

Some background for anyone new here: I am developing a reasonably complex solo project (a full financial platform – microservices, Kubernetes, the works) with AI assistance doing a lot of the legwork, and I run local models on my own hardware alongside the cloud ones, served by vLLM behind a LiteLLM proxy, on the GPU nodes in my own cluster. This is partly to learn how - AI is going to transform the way we develop software, and as engineers we have a choice to either understand how we can be in control of that, or get left behind - and also to build a genuine platform. My mission here is to learn how to produce a genuine, rock-solid well architected platform, with AI doing the donkeywork.

The model doing the work in this post deserves naming, because these are the details I always want from other people’s write-ups: Qwen3.6-35B-A3B, in Nvidia’s NVFP4 quant, served by vLLM on an Asus GX10 – one of the DGX-Spark-class desk boxes. The GX10 currently lives outside the Kubernetes cluster, which is why the appendix contains a hand-run docker run rather than a manifest: it’s an ARM box, the cluster is x86, and when I last looked at mixed-architecture Kubernetes the challenges were real enough that I filed it under fun for another day. On my network the model answers to the name zeu; the zeu-nothink variant that this post is largely about is not a second model but a LiteLLM alias for the same server, with the proxy injecting the kwargs that switch reasoning off. The full vLLM incantation is in the appendix.

The development process is heavily automated: an AI author works a ticket, pushes, opens the MR, and a separate AI reviewer – different session, model, and Git identity, deliberately walled off from the author – reviews it. That deep review runs on a frontier model and is worth its cost, but it only happens when an MR is ready. Which leaves a gap: all the pushes before that point, where a dumb mistake can quietly become a foundation for three days of work built on top of it.

Hence the light tier: a CI job on every push that hands the diff to a local model and asks for findings. The marginal cost is zero – the GPU idles otherwise, and the job runs in parallel with the test stage, so anything under about ten minutes of wallclock is literally free.

Free is a fine argument, but it is not the load-bearing one. Cloud frontier APIs have genuinely dismal uptime – understandable, frontier stuff is hard – and in interactive use you shrug, retry, and get on with your day. A CI pipeline is a different setting: nobody wants a red build because a third party is having a bad afternoon. This is the same reasoning that has us version-pin and vendor our dependencies rather than trust upstream to be present at build time, and a review stage is a build dependency like any other. On my own hardware it fails on my schedule, and fixing it is my job rather than a status page to refresh. Which leaves only one question: whether the output is worth reading.

Version One: The Checklist That Reviewed Nothing⌗

The first design did what everyone’s first design does: it tried to make the model reliable by making it mechanical. A fixed checklist of failure patterns from past reviews, strict output contracts, keep the temperature of ambition low. Deterministic, predictable, safe.

It was also useless. Over a few weeks of shadow-mode operation – roughly 36 runs – it produced zero findings from about 120 candidate items it claimed to have “verified”. When I finally sat down and audited it against reality, it had walked straight past two real defects on the exact commits where the deep reviewer later found them. Same diffs, same SHAs. The light tier looked at the murder scene and reported the curtains were nice.

The diagnosis, once stated, is obvious: I was using an LLM for what grep is good for, and had squeezed out everything an LLM is actually for. Half the checklist genuinely was mechanically expressible – “does this cited document section exist” – so that became an actual lint: a deterministic script, free, zero false negatives within its scope, no model involved. The judgement-shaped work went back to the model, with a rewritten prompt that asks it to understand what the change is trying to do, reason about it, and report opinions honestly labelled as such. The prompt says, in effect: reasoned opinions are welcome; vague hedging is not; silence on a defective diff is the worst outcome.

Measuring It Properly⌗

Here is the part I would actually recommend to anyone doing this: you already own a labelled test corpus. Your review history is one.

Every finding your expensive review process ever produced is a labelled defect at a known commit. Your version-control host remembers the exact diff of every historical review round. So the calibration harness replays the light reviewer against historical MR states – thirteen of them, in my case: rows with known bugs of specific classes (a dead guard that could never fire, a wrong count in documentation, silently-swallowed error paths, an ARIA violation), rows that were later fixed (a correct reviewer must not re-report the fixed defects), doc-only rows where the correct answer is silence, and one row whose only hosted finding was later refuted – the correct output there is also nothing.

Then the bake-off: I wanted to compare two CLI harnesses driving the model (Claude Code, which I use for almost everything, versus qwen-code, which is Qwen’s own agent and might reasonably be optimised for my Qwen models) and two model modes (thinking versus no-thinking). Four arms. The critical design decision – learned the hard way after an earlier attempt died at a timeout with one arm complete and the other absent – is to run all arms per corpus row, back to back, on the same checkout, so that however far the sweep gets before something kills it, every completed row is a complete comparison. And to stream every result into the job log as it happens, which turned out to be prophetic, because…

Everything Fails, Log Accordingly⌗

A partial list of things that went wrong during one day of calibration sweeps, none of which were the model’s fault:

  • My CI runner caps jobs at two hours. The calibration job requested three; the runner’s own maximum wins, silently. The 3h sweep died at 120 minutes exactly.
  • A timed-out job uploads no artifacts, when: always notwithstanding – the kill arrives before the upload stage. Every carefully-collected results file: gone. The job trace, streamed line by line as results happened, was the only evidence that survived. It was enough, because paranoia had put per-finding detail into it.
  • vLLM crashed mid-sweep, and the container wasn’t set to restart. The subtle part: the LiteLLM proxy in front of it kept answering requests – with errors, but answering – so anything doing a TCP-or-HTTP-level health check thought the backend was fine. Only an actual authenticated completion request tells you the model is alive. The calibration script now sends a tiny one per model before touching the corpus, and aborts the whole job in two seconds if any fail. (The vLLM container now has --restart unless-stopped, which it should have had from day one, and I got to red-test the pre-flight against a genuinely dead backend, which is the kind of test fixture you can’t buy. [That’s Claude taking the optimistic view of this - Ed.])
  • A model mid-run ran itself out of context, and the harness misread the resulting “context length exceeded” as “the prompt was too big”, shrank the diff, and retried in a loop for half an hour. If the model started and then hit a context error, no smaller prompt will save it – that’s a distinction worth encoding.

None of this is exotic. It is exactly the mundane operational sediment that accumulates around any real system, and it’s why “just point a model at the diff” demos don’t survive contact with a CI pipeline.

What the Bake-Off Actually Found⌗

The scoreline, over thirteen corpus rows:

Configuration Labelled defects found Notes
Claude Code + thinking model 1, plus two timeouts The incumbent configuration. Worst arm.
qwen-code + thinking model 1–2 Honest, slow, mostly blind
Claude Code + no-think model 4–5 Fastest; perfect on the control rows; blanks ~1 run in 3
qwen-code + no-think model 5–6 Never blanks on a defective diff; noisiest

Three findings stand out beyond the raw counts.

Reasoning hurt. The folklore about Qwen models overthinking – “but wait…"-ing themselves into a tizzy – showed up empirically and decisively. The thinking arms burned five to fifteen minutes per review producing peripheral nits, and the worst single behaviour of the whole exercise was a thinking-mode run spending its entire 900-second budget reading files and second-guessing itself, producing nothing. The no-think variants went straight to the code and found actual bugs, in a fifth of the time. I had expected the agent harness to matter most, on the theory that qwen-code might handle Qwen’s reasoning loops better; instead the answer was to remove the reasoning.

For a no-reasoning model, the harness is the thinking. I’d assumed harness choice would only affect mechanical things like tool-call reliability. Not so: a model with no private scratchpad acts directly on its conditioning, so the system prompt, tool vocabulary, and loop policy around it shape the outcome more, not less. The two harnesses produced consistently different – and complementary – blind spots: one reliably caught mechanical-verification bugs (dead guards, wrong counts, lint scripts that can’t fire), the other reliably caught behavioural ones (silently-swallowed error paths, notices rendered into invisible containers, races). Neither profile is “better”; they’re different reviewers.

Individual catches don’t repeat; territories do. Re-running the same configuration on the same diff gives you a different subset of findings each time – the star catch of one run gets shrugged past on the next. What’s stable per configuration is engagement level and which classes of bug it hunts. The practical consequence: the reliable product is the union of multiple passes, which on free local GPU is entirely affordable. Three passes of one arm plus one of the other covered essentially everything either could ever find; any single pass covered maybe a third of it. Relatedly, when two different harnesses independently reported the same defect – which happened twice – it was a near-certain real finding. Agreement between diverse samplers is a much stronger signal than confidence from one.

And one philosophical result: the classes of defect that no configuration ever caught were all verify-against-a-reference shaped – ARIA-spec validity, citations pointing at the wrong-but-existing section, “is every route to this state actually gated” enumerations. Noticing-shaped bugs get caught; reference-checking bugs need either explicit verification recipes in the prompt or, better, an actual lint. Which is pleasingly consistent with where this whole story started.

Switching It On⌗

The decision landed the same day as the sweep: qwen-code driving the no-think model, on every push. Findings now post to the branch’s ticket as advisory notes, each with a stable id keyed to the commit that raised it, and the author agent must disposition every one – fix it and say so, or decline it with written grounds. Nothing blocks; the tier is an advisor with a paper trail.

Two details of the switch are worth recording. The cross-harness second pass – the thing I said I’d “probably” do, because the two harnesses’ blind spots interlock so neatly – got deferred, deliberately. The review prompt is the biggest lever in the system, and it can only be tuned against one reviewer at a time; a prompt optimised for two reviewers is optimised for neither. One arm, tuned hard, first. And the retuned prompt takes the calibration’s blind-spot lesson seriously: every attention pattern in it now carries an explicit verification recipe. “Check the citations” becomes “open the cited section and compare what it actually says with what the code claims” – because the sweep showed that noticing-shaped defects get caught and verify-against-a-reference defects get walked past unless the procedure is spelled out.

The switch-over MR was itself reviewed by the thing it was switching on, which produced the first live data: seven review rounds as the branch’s commits landed, nineteen findings – nine fixed, all real; ten declined with reasons. The per-round fix count went 3, 1, 2, 0, 1, 1, 1: a decaying curve that looks exactly like a reviewer running out of true things to say, which is what you want to see.

Three Days of Mileage⌗

Then real development happened: a new microservice scaffolded and built out over fifteen tickets – user identity: token exchange, a persistence layer, a mock identity provider – plus ADRs and process work. Fifty-five review rounds produced 123 findings between them, and the author agents dispositioned every one: 48 fixed, 74 declined, one ticketed for later.

A hit rate around forty percent is well above what I was braced for, and the fixes are not padding. A sample from the week:

  • The mock IdP’s discovery document advertised scopes_supported: [openid, email] while the generated login config requested openid phone – so a scopes-respecting client would never ask for the phone claims the whole flow exists to deliver.
  • The local-deploy recipe claimed to provision both mock IdP instances’ signing keys but only provisioned one, so a cold local start of the new service was simply broken. Flagged blocking, and rightly; fixing it swept the class and caught a third instance of the same gap the reviewer hadn’t even seen.
  • A server-side token-file read failure was mapped to HTTP 401 – telling the caller their credentials were bad when the server’s own filesystem was. Now a 500, with tests pinning the distinction.
  • In the repository layer, create() inserted the user document first and appended the version-1 audit snapshot second, so a failure between the two left a visible user with no audit trail. Reordered companions-first: a failure now strands only unreferenced, PII-free orphans.
  • A staleness check gained its boundary tests: age exactly equal to TTL is fresh under the strict > that matches the Mongo reaper; one millisecond past is stale.

Underneath those, a steady background of documentation drift – stale comments, citations pointing at the section next door, doc claims the code had outgrown. Individually trivial; collectively the reason documentation rots.

The declined pile turns out to carry value too, in a way I hadn’t predicted. Because a decline must be reasoned, the cheapest honest way to decline a plausible-but-wrong finding is often to make the code state its own rationale. The reviewer flagged a write-then-snapshot ordering in the update path as a consistency bug; it is the deliberate design (the alternative fabricates audit history for writes that never landed), but the decline forced that rationale out of the author’s head and into the module documentation, where it should have been all along. Another finding misdiagnosed a deliberate ordering as a bug but exposed, in passing, that the field’s rustdoc promised “successfully issued” when the code records the attempt – the finding was wrong and the documentation fix it provoked was real. A false alarm that leaves the code better explained is a strange kind of failure.

The Reviewer Does Not Read Your Replies⌗

One mechanism from the switch week deserves its own confession. The light-review note on each ticket keeps a folded history of past rounds, and before each run the harness fetches those rounds with the author’s dispositions and feeds them into the prompt, with an explicit rule: a declined finding is not re-raised without new evidence named against the prior finding id. Anti-nag machinery, in short.

It went live at 07:06. At 07:32 a finding was declined with evidence. At 07:50 the next round ran – the job trace confirms eight prior findings and their dispositions went into the prompt – and out came the same finding, restated, not naming the prior id, with a fresh factual error added for garnish. The suppression rule held for eighteen minutes.

The worst case was better still: across three consecutive rounds on one branch the reviewer raised the same false claim about an ADR section – asserting the section “does not discuss resource ownership” when the relevant sentence is verbatim in it – and escalated it to blocking on the second raise, straight past an evidence-anchored decline and then past a dual citation added to the code specifically to disambiguate it. Two other findings made third raises the same week. The pattern underneath is consistent with the calibration result: when auditing citations the model matches section titles, not section content, and it treats the fed-forward dispositions as scenery. Making history available to a model and making it binding on the model are different problems, and I have solved one of them.

The machinery still pays for itself, just not where advertised. Stable finding ids mean a re-raise is dispositioned in one line – “third raise; prior decline stands” – so the nag tax rounds down to seconds. And every recurrence gets logged against the calibration ticket, where the prompt’s attention patterns are scored fix-or-retire: a pattern that only ever produces the same false positive gets rewritten or removed, with live traffic as the evidence. The calibration corpus from the bake-off keeps compounding for free.

More Sediment⌗

The operational list from the bake-off day kept growing after go-live. Four entries from the first three days:

  • There were two definitions of “ready”, and nothing connected them. The deep reviewer’s readiness watcher considered a branch ready when the pipeline was green and the MR open; the light tier considered its work done when every finding was answered. So the expensive review could start while light-review findings sat undispositioned. The fix makes one definition mechanical – clean means the newest round fully dispositioned, which is emphatically not the same as zero findings – and gates the readiness watcher on it. The same not-clean state that pauses the reviewer also summons the author, so the autonomous loop cannot deadlock; and every failure path in the gate fails open, because an advisory tier that can block reviews when it breaks has stopped being advisory.
  • The disposition ledger needed a grammar. Authors naturally wrote dispositions as Markdown tables; the parser expected inline lines; three perfectly readable ledgers parsed, silently, to nothing. There is now exactly one parser module shared by every reader – the prompt feed-forward, the note folding, the readiness gate – accepting both shapes, with a closed verdict vocabulary so bold prose cannot masquerade as a verdict. (While writing this post I scraped the week’s ledgers with a quick regex, got a count of 106, and found the missing seventeen were… in table-format ledgers my regex didn’t parse. Some lessons apparently need learning per-person.)
  • The outgoing harness went down swinging. On its final day, the claude-driven wrapper got stuck in a permission loop: the model repeatedly attempting a tool call outside its allow-list, being denied, and retrying until the job was killed by hand. The structural moral: an allow-list that denies at call time invites retry loops, while qwen-code’s whitelist means a disallowed tool is never offered at all. There is a difference between telling someone no and not showing them the button.

Where This Lands⌗

Three days is mileage, not a verdict. But the shape is holding: every push reviewed for a marginal cost of zero, roughly forty percent of findings worth acting on, the rest declined for a line each, and the only harness timeout I saw all week retried automatically, thanks to the bake-off’s operational scar tissue. The false-alarm tax is real, but the disposition discipline caps it, and a surprising fraction of the tax refunds itself as documentation. My suspicion from the bake-off write-up has survived contact with production: tolerating the chatter is part of why this thing never goes silent on a genuinely broken diff. A reviewer terrified of being wrong finds nothing.

If I had to compress the week into one piece of advice for anyone building the same thing: the model is maybe a third of the system. The rest is lifecycle – stable finding ids, one grammar for dispositions, one definition of clean, gates that fail open – plus a feedback loop that turns live false positives back into prompt calibration. None of it is glamorous, and all of it is the difference between a reviewer and a random-opinion generator wired to your CI.

Coda⌗

While this post was being finished, a Claude Code update changed its system prompt in a way that Qwen’s chat templates did not appreciate, and Claude-Code-driven access to half the local models on my proxy stopped working. It can most likely be fixed with some Jinja surgery on the templates, and I’ll get to it. But the timing is pointed: the light tier, freshly moved onto qwen-code, kept reviewing throughout and never noticed.

This is not really a complaint. Claude Code is built by a frontier vendor for that vendor’s models, it remains my daily driver for the cloud ones, and pointing it at third-party templates is off-label use that worked until it didn’t. That is precisely the argument, though: a harness from the model’s own family tracks the model’s assumptions as a matter of course, while for anyone else’s harness your model is a compatibility target nobody is testing. The permission-loop incident above made the same case from a different angle. I chose qwen-code for review quality; the reliability came bundled, and this week it got demonstrated for free.

Appendix: The vLLM Incantation⌗

Since this is the part of other people’s posts I always end up searching for, the full serving command – most flags here are the product of at least one thing not working without them:

docker run --restart unless-stopped -d --gpus all --name vllm-serve \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HF_TOKEN=<redacted>" \
  --env "CUTE_DSL_ARCH=sm_121a" \
  --env "VLLM_MARLIN_USE_ATOMIC_ADD=1" \
  -p 80:8000 \
  --ipc=host \
  vllm/vllm-openai:v0.27.1 \
  nvidia/Qwen3.6-35B-A3B-NVFP4 \
    --served-model-name zeu \
    --trust-remote-code \
    --gpu-memory-utilization 0.6 \
    --max-num-seqs 8 \
    --max-num-batched-tokens 8192 \
    --enable-chunked-prefill \
    --async-scheduling \
    --enable-prefix-caching \
    --load-format fastsafetensors \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_xml \
    --enable-auto-tool-choice \
    --override-generation-config '{"temperature": 0.6, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 1.0, "repetition_penalty": 1.0}' \
    --attention-backend flashinfer \
    --moe-backend marlin \
    --quantization modelopt_mixed \
    --default-chat-template-kwargs '{"preserve_thinking": true}' \
    --enable-log-requests

Two footnotes to that. The --restart unless-stopped on the first line is the flag whose absence you read about earlier; it is load-bearing now. And zeu-nothink appears nowhere above because it lives a layer up, in LiteLLM: an alias for the same server that injects the chat-template kwargs to disable reasoning – one vLLM process, two models as far as every client is concerned.

Oh, and worth mentioning, I’m fairly sure vLLM is ignoring that “presence_penalty” in the generation config - but I do also inject it into the individual requests using LiteLLM. It’s not a silver bullet, but it definitely tames Qwen’s tendency for looping.

This is what a review round looks like from the GX10’s side:

nvtop on the GX10 showing VLLM::EngineCore holding the GB10 GPU at 85% utilisation during a light-review round
VLLM::EngineCore mid-review: 85% GPU, 50 of 121 GiB, 49 degrees, and an alleged 19 W, which I choose to believe.

Disclosure, in the spirit of transparency about AI use: this post was drafted by Claude – the same assistant that ran the bake-off it describes – and edited by me; the production statistics were tallied from the tickets’ disposition ledgers via the GitLab API. The hero image comes from the ComfyUI deployment on the same cluster, it seeming only fair to keep the artwork local too. The irony of the cloud frontier model writing up the case for the local models that review its code is noted – and, on reflection, enjoyed: frontier intelligence helping get the best out of local tools looks a lot like the whole arrangement working properly.