Skip to content
2pizza.teamBlog

AI Media-Buying Agents: L1, L2, L3 Autonomy Explained (2026)

Ivan Bolonikhin
Founder, 2pizza.team

TL;DR: Autonomous media-buying agents that spend without human review are a lawsuit waiting to happen. Bounded autonomy in three tiers is the working pattern - L1 the agent recommends, L2 a human confirms, L3 the agent acts inside a strict KPI envelope with kill switches. Skip L3 for regulated verticals until you have three months of L1/L2 data proving the agent does not do stupid things.

The conversation about AI agents has bifurcated in 2026. On one side, autonomous agents that plan, act, and spend money on their own with minimal human oversight. On the other, tightly bounded agents that make recommendations a human confirms before any action. Both sides claim theirs is the future.

The pattern that ships in production, at least for media-buying and other spend-committing agent workloads, is a three-tier autonomy framework - L1, L2, L3 - with the agent operating at different tiers depending on the specific decision it is making. This post is the working version of that framework, with the caveats we have learned from production deployments.

Why not full autonomy

The pitch for autonomous media-buying agents sounds compelling. The agent watches your ad accounts, notices when a campaign is underperforming, reallocates budget, tests new creatives, and reports back. No human in the loop. Total operator freedom.

In production, the failure modes of full autonomy have been named repeatedly - the agent decides at 3am to double a campaign that looked promising but was actually being gamed by bot traffic. It reallocates budget away from campaigns underperforming today that would have recovered tomorrow because of a normal weekend dip. It generates new creative that inadvertently violates platform policies or brand guidelines. Each of these is a fixable problem individually. Together they are the reason no serious operator lets a media-buying agent run truly autonomous over their spend.

The compliance dimension amplifies this in regulated verticals. In iGaming, in finance, in healthcare adjacencies, an agent that spends money on advertising without human oversight is exposing the operator to regulatory action for anything the agent does that violates content rules, targeting rules, or attribution rules. The operator remains liable, even if the agent made the decision.

The L1/L2/L3 framework

The framework that works in production splits agent behaviour into three tiers, each with different human involvement and different scope of action.

L1: Recommendations only

The agent observes the state of the ad accounts, campaigns, and creative library. It produces recommendations - reallocate this budget, pause this ad group, test this new headline against the incumbent. It does nothing without human confirmation. The recommendations arrive as tickets, Slack messages, dashboard rows, or Telegram alerts. A human reads them and clicks approve or reject.

L1 is where every media-buying agent starts. It is also where many should stay. The value of L1 is not that the agent makes decisions - it is that the agent surfaces decisions the human would otherwise miss. Underperforming campaigns get noticed on day one instead of day seven. New creative variants get tested in a structured cadence instead of when someone remembers. The bottleneck moves from noticing to deciding, which is the direction you want it moving.

L2: Human-confirmed actions

L2 is L1 plus one-click execution. The agent recommends, the human confirms, the agent then executes. The human still sees every action before it happens. The difference from L1 is speed - the human does not need to leave the dashboard to act on the recommendation, and the agent handles the API calls to the ad platforms.

L2 works well for routine actions where the human is confident the agent has the context right. Pause a specific ad. Move budget between two ad groups. Duplicate a working creative and test a new variant. Each of these is a low-stakes decision where the human is willing to trust the agent's execution as long as the human is choosing what happens.

L3: Bounded autopilot

L3 is the tier where the agent acts on its own, but strictly inside a KPI envelope with hard kill switches. The envelope is defined per campaign, per action type, per day. If the agent wants to shift more than 20% of a campaign's daily budget, it drops to L1 and asks. If a KPI metric crosses a floor threshold, the agent stops all L3 actions and reverts to L1 pending human review. If total daily spend on autonomous actions exceeds a hard cap, everything stops.

L3 works for narrow, well-defined actions where the agent has been observed at L1 and L2 for weeks and has proven its recommendations were reliable. Pausing an ad group that has dropped below CTR floor for three consecutive days. Reallocating up to 10% of daily budget from one campaign to another within the same portfolio. Rejecting a new creative variant that failed the brand-guideline check.

The three-months-of-L1/L2-data rule is important. Before moving any action to L3, we require that the agent has been observed making L1 recommendations on that specific action type for at least three months, and the human confirmation rate is above 85%. If the human is rejecting more than 15% of the agent's recommendations on that action, the agent is not ready for autonomy on that action.

Kill switches and enforcement

L3 does not work without kill switches. There are three we require on every production L3 deployment. First, per-action-type spend caps enforced by a middleware layer - the agent cannot exceed the cap even if it tries. Second, KPI floor breaches trigger a revert to L1 across all L3 actions - not just the one that breached. Third, a manual kill switch that any human on the operator's team can trigger to freeze all L3 actions immediately.

The enforcement point matters. If the kill switches live inside the agent's own code, the agent can theoretically override them. We put them in a middleware layer between the agent and the ad platform APIs - the agent asks the middleware to execute an action, the middleware checks the envelope, and either passes the request or refuses it. The middleware has no autonomous behaviour of its own - it is a passive gatekeeper.

What the agent is actually good at

The workloads where media-buying agents at L1 or L2 have consistently added value in production - underperforming-campaign detection at the ad-group level, creative variant fatigue detection (when a creative that used to work stops working), audience segment drift detection, and continuous A/B test management (winner selection at statistical confidence, not gut feel).

Where agents at any tier have consistently underperformed against a human media buyer - new-market entry decisions (too much context outside the agent's window), creative concept generation (agents produce technically correct but boring creatives), and any action that involves reading between the lines of brand or compliance guidelines.

For iGaming operators specifically

iGaming media buying has an additional layer - compliance. Ad content rules vary by market. Targeting is restricted in some jurisdictions. Attribution and disclosure requirements affect what the agent can say in creative copy. Any autonomous agent operating in iGaming needs to check every proposed action against the current compliance rules for the specific market before executing.

We keep L3 disabled for anything that touches creative content in iGaming deployments. The agent can pause, reallocate, and A/B test at L3 within limits. It cannot generate new creative or modify existing creative copy without a human review pass. This is not a scaling limitation - it is the pattern that keeps the operator compliant.

What good looks like in the first six months

The realistic six-month arc for a media-buying agent deployment: months one and two, L1 only, human reviews every recommendation, agent is calibrated by rejection patterns. Month three, selected action types graduate to L2 - the human still confirms but with one click instead of manual execution. Months four and five, L3 opens up on the narrowest actions with the strictest envelopes. Month six, evaluation of what worked, what got rolled back, and where the envelope should widen or narrow.

The teams that get this wrong try to go straight to L3 in month one, hit a bad decision that costs money, and roll everything back to no autonomy. The teams that get it right stay patient at L1 for longer than they think necessary, build human trust in the agent's recommendations, and graduate individual actions to higher tiers based on evidence.

Deploying a media-buying agent and worried about the autonomy question? Book a scoping call. We help operators structure the L1/L2/L3 framework for their specific ad platforms and compliance requirements. See /igaming or /services/ml for the broader engineering picture.

Want us to look at your setup?

Free 30-min audit. We tell you what to automate first and what it would cost.

Book a free audit
iGaming
Per-player Scoring Architecture for Online Casino Operators (2026)
15 min read
Hiring
Best AI Automation Agencies for Small and Mid-Size Businesses in 2026
13 min read
GEO & AI Search
Generative Engine Optimization: What the Data Actually Supports in 2026
17 min read