Insight / signal

Marketing agents need a ledger before they touch the budget

If an agent can change spend or kill creative, require an experiment ledger with evidence and rollback.

Marketing agents need a ledger before they touch the budget

The first useful marketing agent will not be the one that writes 100 ads.

It will be the one that remembers why 94 of them were rejected.

That sounds less exciting, which is usually how you know you are getting closer to the real work.

Most AI marketing demos still obsess over the visible bit: generate copy, make images, produce variants, turn one idea into twenty hooks. Fine. That was impressive for about five minutes. It is now table stakes.

The harder question is what happens once those variants meet a live market.

Which customer pain point was this ad testing?
Which source produced the idea?
Which creative family has already been overused?
How much did it spend?
What was the real success metric?
Who approved it?
Why was it killed?
What did we learn that should change the next batch?

If the system cannot answer those questions, it is not a marketing agent. It is a content generator wearing a hi-vis jacket.

Today’s useful signal came from a Startup Ideas episode where Cody Schneider described using Claude Code, MCPs and data connectors to run marketing loops. Strip away the founder-podcast gloss and there is a solid shape underneath it: data from paid media, analytics and business systems lands in a warehouse; the agent researches pain points and competitors; it generates copy and creative; it publishes through Meta’s Marketing API; it waits a couple of days; it kills weak variants; it records the campaign and ad relationships; then it repeats.

That is much closer to a real AI marketing system than “write me ten Facebook ads”.

The most important part is not Claude Code. It is not the image model. It is not even the Meta API.

It is the registry.

The registry is where the work compounds. Every hypothesis, ad ID, creative asset, source link, spend figure, result, decision and reviewer has somewhere to live. Without that, the agent has no memory. It can produce endless new material, but it cannot build judgement.

And judgement is the expensive bit.

There is a reason this matters now. Meta’s ad system has changed. Andromeda, Meta’s personalised ad retrieval engine, is designed to sort through a much richer supply of candidate ads and match them to people at the retrieval stage. Meta’s own engineering post talks about improved retrieval recall and ad-quality metrics on selected segments. That does not prove that “creative is targeting” in the simplistic LinkedIn way, but it does mean creative diversity matters more than it used to.

So yes, there is a real opportunity here. If the platform can evaluate more creative possibilities, businesses need better ways to produce and test useful creative possibilities.

But that is also where the naive version goes wrong.

More variants is not the same as better marketing. It can just mean more noise, faster spend, faster fatigue and a very tidy spreadsheet full of misleading early winners.

Killing an ad after two or three days might be sensible in one account and stupid in another. It depends on spend, impressions, conversion volume, margin, learning windows, lagging revenue, seasonality, audience size and whether the metric you are using is actually tied to profit. If the agent kills the wrong thing because it optimised for cheap clicks or early noise, the demo still looks clever. The business just gets poorer.

That is the bit we need to talk about before everyone starts calling their ad generator a media buyer.

A marketing agent needs an operating agreement.

Not a 40-page legal document. A practical agreement inside the system:

What can it read?
What can it draft?
What can it publish?
How much can it spend?
What must a human approve?
What evidence is required before a decision?
What counts as a win?
What stops it?

This is where the wider AI market is pointing anyway.

Google’s Interactions API is not just another model endpoint. It is built for stateful work: server-side interactions, background execution, typed steps, tool use, managed agents and retrievable histories on paid tiers. Google is effectively saying that agent work needs state, progress and recoverability.

Anthropic’s expanded Cognizant partnership says the same thing from the enterprise side. The pitch is not “Claude is clever, good luck”. It is Claude inside existing industries, platforms, engineering standards and client delivery frameworks. In one example, Cognizant says Claude Code runs inside Flowsource’s spec-driven development module and evaluates output before production. Again, the model is not the whole story. The operating layer is.

Even Anthropic’s Opus 5 launch leans into long-running work: coding, automation, computer use, professional workflows, tool changes mid-conversation and fallback routing. Whether you buy every benchmark claim or not, the direction is clear. AI is being sold less as an answer machine and more as a worker inside a managed process.

Marketing cannot pretend it is exempt from that.

If an agent is only drafting ideas, the risk is mostly taste. You reject the bad ones and move on.

If an agent is publishing into Meta, changing budgets, pausing ads or pulling conclusions from live performance data, the risk is commercial. Money moves. Accounts can get restricted. Brands can make claims they cannot defend. Teams can learn the wrong lesson because the experiment was badly designed.

That is why the first useful version should probably be boring.

Run the agent in shadow mode.

Let it ingest the account data. Let it propose the next creative batch. Let it explain which pain point each ad tests. Let it recommend what to pause or scale. Let it produce the experiment receipt.

Then a human media buyer, founder or marketing lead makes the actual change.

Do that for a few weeks. Compare the agent’s recommendations against the human decisions and the eventual outcome. Track where it was right, where it was early, where it missed context, where it overreacted, where the metric lied.

That is how you build trust. Not by giving the machine £500 a day and hoping it grows up quickly.

A proper experiment ledger does not need to be glamorous. It needs to be useful.

For each variant, I would want at least this:

  • hypothesis;
  • source insight or customer language;
  • copy and creative asset hash;
  • campaign, ad set and ad IDs;
  • objective and audience settings;
  • spend cap and live spend;
  • impressions, clicks, conversions and revenue where available;
  • margin or qualified-lead quality where possible;
  • decision window;
  • kill, keep, scale or revise decision;
  • human reviewer;
  • reason for the decision;
  • rollback or restore state.

That is not admin for the sake of admin. That is the difference between activity and learning.

The old agency model could hide a lot behind output. More ads, more decks, more calendar posts, more weekly reports. The AI-assisted version of that model is even worse because it can produce the pile faster.

The post-agency model has to be different.

It should sell the loop.

Map the workflow. Connect the right data. Define the permissions. Run controlled experiments. Keep receipts. Measure against the business constraint. Improve the system. Show what changed.

That is the work clients actually need, even if they currently ask for “an AI agent” because that is the phrase everyone is using this month.

For a business owner, the test is simple.

If someone pitches you an AI marketing agent, ask to see the ledger.

Not the prompt. Not the demo video. Not the list of tools in the stack.

Show me the experiment history. Show me what it tested, what it learned, what it spent, what it changed, what it got wrong and where a human had to intervene.

If they cannot show that, they are not selling an agent.

They are selling a faster way to make a mess.