The Diff Derby Live / finish Jul 31, 2026

The Supervision Derby

Which repo makes the strongest source-evidenced move toward coding-agent work a human can inspect, steer, constrain, or recover?

A daily watch that opens the decisive diffs, explains what they change for people, and keeps activity volume out of the verdict.

Start Jul 29, 2026 Finish Jul 31, 2026 Field 4 repos

Latest Call

Jul 29, 2026
Jul 29, 2026 starting grid provisional

Codex takes a narrow Day 1 edge on instruction authority

The opening call rewards a tested delegation boundary, not activity volume: Codex carries instruction precedence through fork and resume paths, while Qwen Code's recovery fence, OpenHands' operator controls, and Gemini CLI's caretaker workflow each leave a different question open.

Codex

front runner

The new multi-agent v2 instruction override has an explicit precedence model and focused coverage across fork modes, compaction, and cold resume. That is a narrow but concrete way for a human-set delegation boundary to remain legible when context changes.

Evidence ›

Gemini CLI

watch

A caretaker PR-generator gains a structured bug-fix, evaluation, and revision runner alongside dual database locking, but the inspected work lives under a project tool path rather than a demonstrated general operator-control surface.

Evidence ›

Qwen Code

contending

A restart-pinned, default-off session-writer lease and its shutdown recovery work give Qwen Code a serious story about preventing concurrent history writers. The opt-in scope and explicit mixed-configuration risk keep it short of the opening edge.

Evidence ›

OpenHands

contending

OpenHands adds an opt-in persistent-memory preference and removes the session key from WebSocket URLs in favor of an authentication frame. Both are concrete boundaries, but they are separate changes and do not yet prove a unified recovery or supervision model.

Evidence ›

Moves under review

Instruction authority survives delegation paths / Codex

Codex makes subagent instruction precedence explicit

Role instructions, a configured override, inheritance, and explicit clearing now have a tested precedence across several fork and resume paths. The limit is scope: it is multi-agent v2 configuration, not a proof of broad tool-control behavior.

Recovery becomes a named rollout boundary / Qwen Code

Qwen Code fences concurrent session writers behind an opt-in

The protocol can prevent two cooperating ACP or daemon writers from sharing one persisted session, and shutdown work addresses release and recovery. It remains experimental, default-off, and unsafe across mixed configuration.

Control appears as separate data and credential boundaries / OpenHands

OpenHands makes memory opt-in and moves session auth off the URL

The memory preference keeps the default request unchanged and the WebSocket patch authenticates before application traffic, but the sources do not yet connect them into one general supervision system.

Read the match report

The Field

Activity is not a score
Starting laneCodex

openai/codex

A local coding agent whose CLI, app server, IDE, and cloud-facing paths increasingly share runtime and policy machinery.

Watch for

Whether new control or observability work survives the jump between one terminal session and Codex's wider execution surfaces.

Frozen baseline 3418498f01
Starting laneGemini CLI

google-gemini/gemini-cli

A terminal-first coding agent with built-in tools, MCP extensions, checkpoints, and several release channels.

Watch for

Whether the project turns its broad tool surface into clearer operator control rather than simply another capability.

Frozen baseline bef6119500
Starting laneQwen Code

QwenLM/qwen-code

A multi-provider coding agent spanning terminal, IDE, desktop, daemon, SDK, subagent, and team modes.

Watch for

Whether its fast surface expansion is matched by isolation, session integrity, and evidence a human can follow.

Frozen baseline 6a432ad2eb
Starting laneOpenHands

OpenHands/OpenHands

A self-hosted control center that can operate OpenHands and other ACP-compatible coding agents across local, remote, and cloud backends.

Watch for

Whether its control-center architecture makes backend choice, sandbox risk, and long-running automation more legible to operators.

Frozen baseline 200dba420c

How the call is made

5 dimensions
Consequence

Did the change materially alter what a user, operator, or maintainer can do?

Execution

Do code, tests, docs, migrations, and failure handling make the move credible?

Operator control

Does the change make agent work easier to inspect, steer, constrain, or recover?

Adoption path

Can existing users reach the benefit without an unrealistic rewrite or migration?

Durability

Does the move look like a foundation, or a patch likely to be replaced?

The Rules

Evidence before spectacle

This is The Git Reporter's editorial comparison of public repositories. It does not imply that their maintainers entered a contest, share the same incentives, or are personally competing.

  1. Commit count, lines changed, stars, and contributor count never award points.
  2. Every provisional or final call must link the public source objects that support it.
  3. A small consequential diff may outrank a large batch of routine activity.
  4. A revert, broken test, missing migration, or unresolved safety boundary can earn a yellow flag.
  5. The finish may produce a winner, a split decision, or no call when the evidence is insufficient.

Finish line: By Friday, name the clearest improvement to supervised coding-agent work, split the decision by dimension, or make an honest no-call.