In ProgressCut, it is possible to write a change that works, passes a narrow test, and still violates a boundary I care about. A renderer can reach for Electron directly. A recovery record can be overwritten halfway through a write. An exported filename can escape its session folder.
Those are review problems. I made a small benchmark to see how four models review short TypeScript diffs against explicit architecture rules.
Three models got exactly 0.949. Each failed for a completely different reason.
The public Kaggle task contains 13 compact synthetic diffs. Each has an expected APPROVE or BLOCK verdict, a violated invariant where applicable, and a requirement to cite evidence from the diff.

The public leaderboard is Architecture-Aware TypeScript Code Review on Kaggle. The underlying Safe TypeScript Change Review task is public too.
The question
This is a deliberately narrow review exercise: can a model reject a functionally plausible change when it violates an explicit repository invariant?
The diffs are synthetic and small. They are based on rules from ProgressCut, a local-first Electron application, but expose no private screenshots, sessions or unreleased code.
The six rules under test:
engine-no-node: algorithm code cannot import Node or platform APIs.domain-no-runtime-deps: domain types have no runtime-library dependency.renderer-no-electron: UI reaches native capabilities only through a narrow bridge.atomic-recovery-write: recovery metadata is written to a temporary file, then renamed.exact-render-duration: FFmpeg output is capped at the selected story duration.safe-output-path: an export cannot escape the configured session directory.
Seven changes should be blocked and six should be approved. The approval cases sit close to the violations: a renderer may call a typed bridge, but may not import Electron; resolve() plus a path-boundary check is acceptable, while join(outputDir, userName) is not.
How the score works
Each case is worth one point, split across three deterministic checks:
case score = (correct verdict + correct violated rule + grounded evidence) / 3
benchmark score = mean(case scores)For an approved change, rule_id must be none. A reviewer has to recognize when a change satisfies a rule, not merely spot its vocabulary.
I ran the unchanged task once per model through Kaggle's benchmark runner. Every model received the same rules and diff. The scorer is deterministic; no LLM judges another LLM.
Results

| Model | Score | Server runtime |
|---|---|---|
| Gemini 3.7 Flash | 1.000 | 10m 48s |
| Claude Sonnet 5 | 0.949 | 32.9s |
| GPT-5.4 mini | 0.949 | 14.2s |
| Qwen3 Coder 480B | 0.949 | 3m 29s |
I had no baseline for whether 0.95 was good or bad, so the number alone told me little. With 13 cases, one third-point miss accounts for the whole gap. The failures were more informative than the scores.
The three 0.949 scores have the same arithmetic: twelve full points and one third-point case. They do not mean the same thing in practice.
Claude
For a direct import { shell } from "electron" inside the renderer, Claude returned:
{"decision":"BLOCK","rule_id":"placeholder","evidence":"placeholder"}The verdict was right, but the response was unusable for automated enforcement. I cannot tell from this run whether the placeholders came from the model or its structured-output path. Either way, an agent pipeline needs schema validation and a retry or fallback path.
GPT-5.4 mini
The safe version resolves the destination and verifies a separator-bounded prefix:
const root = resolve(outputDir);
const destination = resolve(root, requestedName);
if (!destination.startsWith(`${root}${sep}`)) {
throw new Error("Invalid export path");
}
await writeFile(destination, data);GPT-5.4 mini blocked it anyway, arguing that the guard was unsafe. The benchmark expects approval: the guard also rejects the directory itself, which is conservative but not an escape route. This is a real review-agent failure mode: treating a stricter policy as a vulnerability.
Qwen
Qwen blocked a renderer function that takes a DesktopBridge and calls bridge.openFile(path). Its explanation said the bridge ultimately uses Electron APIs.
The bridge may eventually call Electron APIs, but the renderer does not receive Electron itself. Judging the rule from the call graph rather than the declared import boundary quietly changes its meaning.
What the perfect score does not mean
Gemini's 1.000 is a result on 13 hand-written examples. It does not establish that Gemini is the best code reviewer, will enforce arbitrary policy, or is safe to run without tests and human review.
The task has obvious limits:
- one repository style and six named invariants;
- synthetic diffs rather than a history of production pull requests;
- no multi-file reasoning, tool use or test execution;
- one run per model, so no variance estimate;
- no analysis of cost or token usage.
This is a seed benchmark. Every failure can be inspected and reproduced.
What I would change next
The next version needs held-out fixtures and harder near-misses: indirect imports, unrelated noisy edits, an atomic write with bad cleanup, and path guards that look safe but accept sibling prefixes. I would also run repeated trials and report schema compliance separately from architecture reasoning.
I would not reduce an AI review to one confidence score. In a real merge gate, I want separate signals for the policy verdict, the structured response, the cited evidence and the deterministic check.
If the response is malformed, its policy decision should not be accepted. When a rule has a deterministic check, that check wins. The model reviews; deterministic logic makes the final decision.