Agency KPIs: The Metrics That Actually Run Delivery
Most agency KPI lists I get shown were assembled from whatever the project management tool reports by default. That is the wrong starting point, and it produces the same page of numbers everywhere: metrics that are true, tidy, and have never once changed a decision. I have spent twelve years running delivery operations for marketing agencies, across 79 completed engagements with a 4.9 out of 5 client rating, and what follows is the short list I actually use to run a delivery week.
Measurement is one section of five in my agency operations guide, and this piece is the deeper version of that section. I want to take each metric apart the way I would in a review: what it tells you, how it gets gamed, and the decision it is supposed to trigger. That third test is the one that matters. A metric nobody acts on is reporting, not a KPI.
What makes a KPI worth tracking?
Three things, and all three have to be present. A named person who reads it. A cadence short enough to act on. A decision that changes depending on what the number says. Miss any one of them and you have built a fact, not a key performance indicator.
The test I apply out loud in operations reviews is blunt: if this number doubled tomorrow, what would we do differently on Monday? If the honest answer is nothing, the metric comes off the list. Most scorecards lose half their rows to that question.
Which KPIs actually run agency delivery?
Five, in my experience, carry most of the weight. Utilisation, on-time delivery rate, rework rate, scope variance, and cycle time from brief to ship. Each one is useful, and each one is gameable without anyone intending to cheat.
Utilisation
Utilisation tells you whether the capacity you are paying for is pointed at work that earns. It is a capacity question and nothing more. It says nothing about whether the work was any good, which is why an agency can run hot on utilisation and lose money at the same time, if a meaningful share of those hours is work being done twice.
How it gets gamed: definition drift. Billable quietly widens to include internal work, admin time gets logged against whichever client task is open, and Friday-afternoon timesheets get reconstructed from memory. None of it is dishonest. It is what happens when a number is watched and never defined.
The decision it should trigger: hire, subcontract, or decline the next project. High utilisation with dates still slipping is a capacity problem and you should be buying capacity. Low utilisation with dates still slipping is a process problem, and hiring into it makes the problem larger and more expensive.
On-time delivery rate
Measured against the date the client was given, not the date the task was later moved to. That is where the metric usually dies. The practical fix is a separate field holding the first client-facing date, one the tool cannot overwrite when a due date is dragged.
Two quieter versions: declaring something delivered when a draft reaches internal review rather than the client, and counting a partial delivery as a delivery, which is how a month reports well while clients wait on the half that mattered.
The decision: where the constraint is. If slips cluster at one stage, that stage is your bottleneck and it needs either an owner, a resource, or a smaller queue. If slips are spread evenly across every stage, you are not slow, you are overcommitted, and the fix belongs at intake rather than in production.
Rework rate
Rework is the most useful number on this list and the one agencies are least willing to count. It measures the quality of the brief, not the quality of the work. Hours spent doing something a second time are almost always an intake or scoping failure surfacing three stages downstream, wearing the costume of a design problem.
How it gets gamed: it gets absorbed. Nobody wants to log hours into a category that reads like an accusation, so rework is booked as production and disappears. The second route is attribution. Everything ambiguous becomes the client changed their mind, which is sometimes true and is also the most comfortable available explanation.
Two things make it countable. Name the category neutrally, revisions beyond the agreed round, and forbid fault attribution at the point of logging. The decision it should trigger is never disciplinary. It is which stage’s exit condition you rewrite this month, because rework concentrated on one service line is a briefing template problem you can fix in an afternoon.
Scope variance
Scope variance is the distance between what you sold and what you delivered, measured in hours or deliverables. It is the metric that explains a profitable-looking agency with no money in the bank.
How it gets gamed: by kindness. Extra rounds and small additions get absorbed to keep a relationship warm, and because nobody logs them the variance reads as zero while the margin quietly leaves. A dashboard showing no scope creep in an agency that does client services is not a clean result, it is a measurement failure.
The decision: pricing or scoping, and you have to be honest about which. Variance scattered randomly across clients is a change-order habit you have not built. Variance that repeats on the same service line every time is not an account manager failing, it is that service priced or scoped wrong at the template level.
Cycle time from brief to ship
Cycle time is the only one of the five that measures the machine rather than the people in it. Start the clock when the request arrives, not when someone starts work, and split the total into working time and waiting time. That split is the whole value of the metric. In every agency I have worked inside, the waiting half is the larger one.
How it gets gamed: by excluding the inconvenient waits. Client approval time gets carved out as not our fault, which may be fair and still leaves you quoting turnaround times your clients experience as wrong. Measure the whole elapsed time, then report the two components separately.
The decision: which handoff to remove. Long waits before a stage mean a queue, and queues are solved by limiting how much work is allowed into that stage at once, not by asking people to hurry. It also gives you the one thing every agency guesses at, which is a defensible timeline to quote at intake.
What is the difference between a leading and a lagging KPI?
Four of the five above are lagging. On-time delivery, rework, scope variance and cycle time all report a verdict on decisions you already made. They are necessary and they arrive too late to save the quarter they describe.
Leading indicators are the ones you can still act on. Forward capacity load per person for the next fortnight. The count of projects sitting in a stage longer than that stage normally takes. Briefs entering production without a signed-off exit condition. None are impressive on a slide and all are actionable while the outcome is still open.
The rule I would hold any agency to: for every lagging KPI you care about, own at least one leading indicator that moves before it. Without that pairing, your reporting is a well-formatted account of things you can no longer influence.
What should I compare my numbers against?
Your own agency last quarter. That is the honest answer, and I want to be direct about why I am not giving you a table of industry figures: I do not have benchmark data on record, and inventing one would be worse than useless. Numbers like that get quoted in leadership meetings a week later as though somebody measured them.
Even where benchmarks exist, they rarely compare what they claim to. Two agencies reporting the same utilisation figure may be counting billable differently, running a different service mix, and working under different contract shapes. Importing an outside figure also has a predictable failure mode: it becomes a target within a quarter, and a target gets hit by redefining the metric long before it gets hit by improving the work.
Self-comparison only works if you do one unglamorous thing first: write the definitions down and freeze them for the baseline quarter. What counts as billable, what date on-time is measured against, what qualifies as rework. Change a definition mid-quarter and you have not improved, you have swapped metrics. Trend beats precision, but only when the thing trending stayed the same thing.
Which agency metrics look good and change nothing?
The recurring offenders, all of which I have seen presented as evidence of a healthy operation:
- Tasks completed. A measure of how finely you slice work, nothing more.
- Hours logged as a headline figure. Effort is an input. Without the rework split it cannot tell you whether the effort produced anything.
- Percent complete. An estimate wearing the clothes of a measurement, and it is the estimate of the person least motivated to say a project is late.
- Averages with no tail. An average turnaround hides the three late projects that are costing you renewals. Look at the worst decile before the mean.
- Satisfaction scores nobody routes. Collecting feedback that never reaches the stage that caused it is data entry, not measurement.
The pattern is the same in all of them. They describe activity, or they describe the tool, and no decision hangs on the answer.
How many KPIs should an agency actually run on?
Fewer than you are currently tracking. Five is enough for most agencies under fifty people, each with one named owner and a weekly review, defined in writing. Reviewed monthly, a delivery metric on a weekly delivery cycle is decoration. How you display them is a separate question, and I have written that up in the piece on agency dashboards and the reporting stack. I work inside ClickUp, Asana, Monday.com, Airtable and Notion and automate across them with n8n, and none of them will decide which five matter to you.
Where to start with your own numbers
If you want an outside read on which five belong on your list, that is the first half of the Agency Ops Audit. It is $1,500 fixed, takes two weeks, and ends in a written 90-day roadmap and a 60-minute readout call. If your measurement is already sound and the problem is somewhere else, I would rather tell you that in the readout than sell you a dashboard.