When programmers become AI operators: the cost of routine approval

From 1970s machines to AI coding agents: how approval becomes routine, why failures demand reconstruction, and how meaningful human oversight helps.

AI-generated conceptual illustration. A possible progression from careful review to habitual approval.

The agent proposes a change. The explanation sounds sensible. The checks pass. You approve it. Then you approve the next one.

Nothing dramatic happens. That is precisely why the habit can be difficult to notice. Little by little, approval may stop meaning “I understand this decision” and start meaning “the last few decisions worked.”

For programmers, this deserves attention. We build tools that other people operate. With AI agents, we increasingly operate tools that help build the next tools. We are moving into a supervisory role whose difficulties were familiar long before generative AI.

A familiar problem, fifty years later

IBM’s System/370 arrived in 1970, supporting business applications and centralized computing. Programming and operating a computer were already distinguishable responsibilities. IBM history

Imagine a payroll run on a mainframe of that period. The program executes, the output appears, and an operator checks completion. A normal finish does not, by itself, establish that the payroll rules or input records were correct. This is an illustrative example, not a documented incident.

There is a physical counterpart. Siemens dates its SINUMERIK System 7 computer numerical control to 1976. Siemens history

Picture a programmed machine cutting a part. Following the commanded path and producing the right part are different questions: the setup, material and instructions still matter. Again, this is an analogy, not a claim about a specific failure.

In 1983, Lisanne Bainbridge described how industrial automation could leave people responsible for abnormal situations while making their role more difficult. Ironies of Automation

The practical lesson I take from that history is uncomfortable: successful routine operation tells us little about how prepared someone is to handle an unfamiliar failure.

Three ways checking can become a ritual

There is no single established “normalization effect” that explains everything here. Three related concepts help.

Automation bias is accepting automated advice despite contrary evidence. Automation complacency is relaxing attention to a system assumed to be reliable. NASA research

Normalization of deviance describes departures from standards becoming accepted through repetition without obvious harm. NASA discussion

In a development workflow, those could look like three different moments: trusting a reassuring agent summary over a worrying diff; gradually scanning changes less carefully; then accepting that a required review is routinely skipped. These are possible applications of the concepts, not findings from a study of this particular workflow.

Not every quick approval is careless. Familiarity can support good judgment. The problem begins when the evidence needed to approve a change quietly disappears from the approval process.

A possible progression from careful review to habitual approval.
AI-generated conceptual illustration. Possible approval drift, not a measured or inevitable progression.
  1. REVIEW — I understand this change.
  2. REASSURANCE — The previous changes worked.
  3. ROUTINE — Approve without reconstructing why.

The programmer becomes an operator too

It would be inaccurate to call this the first time programmers have supervised automation. Build systems, deployment pipelines and code generators are familiar examples. Builders and operators have also overlapped for decades.

What feels distinctive about agents is how much implementation can be proposed between interventions. An agent can select an approach, edit several files, interpret an error and propose another change. The developer may spend more of the session reviewing decisions they did not make step by step.

My concern is that a fluent explanation can make that distance easy to underestimate. Reading a plausible account of a change and understanding its consequences are different achievements. Technical expertise remains valuable, but expertise needs access to the actual work and time to examine it.

When it breaks, the earlier decisions return

Consider a hypothetical agent refactoring an order-processing service. It changes identifier handling, modifies retries and updates the checks. The developer approves the series after reading summaries. Later, an intermittent duplicate-order problem appears.

Now finding the defective line may be only part of the job. Why was an identifier treated as unique? Did a retry preserve it? Did the updated checks still cover the original behavior? Which change introduced the assumption that connected them?

The developer has to reconstruct both the implementation and the reasoning visible in the proposals. If the agent changed a check to match faulty behavior, passing that check provides no independent reassurance.

I think of this as understanding postponed. Time saved during implementation can become time spent rebuilding context during recovery. That is an engineering concern, not a measured claim that agents always cost more time. Good records, small changes and independent checks can make reconstruction much easier.

Three connected questions reconstruct the duplicate-order example.
AI-generated conceptual illustration. Hypothetical example: recovering the assumptions behind an order-processing failure.
  1. IDENTIFIER — What did “unique” mean?
  2. RETRY — Was the same identifier preserved?
  3. CHECK — Did the check validate the original rule?

Human involvement needs to remain meaningful

My proposal is to design oversight around consequential decisions, instead of making every action another approval prompt.

  • Keep changes small enough to explain. Approve a bounded change with clear assumptions, rather than a growing bundle of unrelated edits.
  • Show evidence beside the proposal. Include the diff, what was checked, what remains uncertain and any change to the checks themselves.
  • Keep an independent reference. Existing requirements and checks should not all be rewritten by the same process that produces the implementation.
  • Preserve a recovery record. Retain observable actions, versions, proposals, results and approvals. An agent’s explanation is a claim to verify, not proof of its internal reasoning.
  • Make stopping practical. Set permission boundaries and usable checkpoints. Match autonomy and workload to the reviewer’s ability to follow the work.

For low-risk, reversible work, bounded automatic execution can be sensible. Decisions about permissions, data meaning or recovery deserve deliberate attention. Asking for approval constantly can itself make the button feel routine.

Research on cognitive forcing interventions suggests that requiring active consideration can reduce overreliance, although it adds effort and does not eliminate it. Buçinca and colleagues

Evidence, human judgment and bounded execution form a review loop.
AI-generated conceptual illustration. Proposed oversight workflow. Records support review; they do not guarantee correctness.
  1. PROPOSAL + EVIDENCE — Diff, assumptions, independent checks.
  2. HUMAN JUDGMENT — Accept, revise or stop.
  3. BOUNDED EXECUTION — Small scope, checkpoints, recovery record.

More capable systems need understandable decisions

Human-in-the-loop design cannot guarantee freedom from bias. It can give people a realistic opportunity to notice a problem and intervene. That opportunity depends on what they see, what they understand and whether they have authority to act.

My view is that good automation should preserve our ability to explain and recover from what it does. As agents become more capable, the oversight must become more intentional.

Before approving the next change, could you explain what you are accepting—and where you would start if it failed tomorrow?