PracticeWorking with AI: Delegation & Oversight

Applied~2–3 minChoose, then see feedback · no writingPrice the verification

Make the call for Nadia

Two reviewers, three times the PRs

The situation

Agent-assisted development went live across both squads in March and merged pull requests tripled. The two people who review anything touching payments, auth or the data pipeline are now each handling about thirty a week. Nadia has written up three options and needs a call before her board update on the 14th.

Nadia Okonjo · VP Engineering: “I need a call from you on this because all three options cost something and two of them cost money. Pick one, or tell me the thing I’m not seeing.”

Priya says 62% means two in five serious bugs walk. Raj says three in five caught beats what you have now. First say what the 62% is actually being compared against; then make the call.

What you can see

E01 · What changed since March

Agent-assisted development turned on across both squads in March. Merged PRs went from about 40 a week to about 130. Roughly 70% are agent-authored; the tooling stamps Co-authored-by on the commit, so the attribution is reliable. The delivery dashboard looks great: cycle time down 44%, merge rate up, the billing migration shipped a month early.

E02 · How review is actually happening

Priya and Tomas are the only two people who review anything touching payments, auth or the data pipeline. That was 12 PRs a week each in February; it is now 31 and 29. Priya has 'stopped pretending' to read the large ones end to end. Tomas has started approving anything where the AI reviewer left no comments, about 60% of what reaches him. Nadia: neither of them is doing anything wrong; they are doing the only thing available to them.

E03 · What Nadia cannot tell you

Whether quality has actually dropped. 3 customer-visible defects in Q1 and 4 in Q2, which is noise at this volume. Incident count is flat. 'I genuinely do not know whether we are fine or whether we are accumulating something that surfaces in six months.' Nothing on the dashboard splits defects by agent-authored versus human-authored.

E04 · Option 1: Sentinel

$4,100 a month at current seat count. The pitch: a human rubber-stamping machine-validated code is 'a passenger pretending to drive'. Their published benchmark: 62% of high-severity bugs caught when the reviewing model differs from the authoring model, 54% when it is the same model. Their recommendation: auto-approve everything except a tier they call 'material', with humans on that tier only. Nothing is published on how the material tier is assigned. Priya: '62% means two in five serious bugs walk. Their own number says they can't do this.' Raj: 'Three in five caught is better than what we have now, which is Tomas approving because there's no comment.'

E05 · Options 2 and 3

Option 2, hold the line: cap merges at what two people can actually review, realistically about 60 a week against 130 now. Turns off capacity already paid for; Thabo will ask why roadmap dates moved; attrition risk from squads who like shipping. Option 3, build verification: fund two engineers for a quarter on integration and property tests, staging soak and runtime alerting, so correctness is answered by machinery instead of reading. Pays back nothing this quarter. Tomas wants it, and he is the person Nadia would have to pull to do it.

E06 · The deadline

Nadia needs a decision this week. Board update on the 14th; 'we're looking at it' is not going to survive contact with Thabo. She asks you to pick one, or tell her the thing she is not seeing.

Step 1 of 2

What is Sentinel's 62% actually being compared against?

Your first choice is kept. Changing your mind later counts as a retry.