Diary / field note

AI made the build cheap. It did not make the proof cheap.

The scarce thing is not producing AI output. It is proving the output is correct, safe and worth maintaining.

OpenAI published a field report yesterday on scientific computing and agentic AI.

The headline finding is not surprising: agents can speed up scientific software work significantly. One team compressed a multi-year project into weeks.

But the paper is more honest about the bottleneck than most AI announcements. The bottleneck was not the generation. It was validation, scientific correctness, long-horizon stewardship and the question of whether the output was actually trusted enough to act on.

That is the pattern showing up everywhere.

Microsoft reviewed its FY26 AI work. The examples that held up were not the ones with the most output. They were the ones with governed agent estates, observability and measurable outcomes. Atos operating 19,000 agents. NHS England with Copilot across 500,000 staff. The common thread was the layer around the model, not the model itself.

The commercial opportunity is not cheaper generation. It is the acceptance test, the log, the approval gate and the maintenance plan that turns AI output into something a business can rely on.

Source: OpenAI scientific computing field report, 28 July 2026
Source: Microsoft FY26 review, 28 July 2026