Stop counting the things your marketing agents generate
An agent produced forty social posts this week. Is that good?
I can’t tell without knowing how many were useful, how long someone spent editing them, whether they were published, and what happened afterward. A larger draft folder can look like progress while the marketing team gets busier cleaning it up.
As agents take on more of the work, we need a better way to judge them than counting outputs.
Separate execution from results
First, check whether the workflow completed the assignment. If the job was to prepare a sourced competitor brief by Monday morning, did it arrive on time with working links and accurate changes? If the job was to publish an approved post, did the correct version appear in the right place?
Then ask whether the work helped the business. Did a useful brief change a decision? Did the published post bring relevant visitors? Did an outreach sequence lead to conversations worth having?
Those questions belong together, but they measure different things. A workflow can execute correctly and still support a weak campaign idea. It can also produce a promising draft while failing to carry out the task you assigned.
Keep a small set of real examples
Anthropic’s January 2026 article on evaluating AI agents discusses evaluations that look at the agent’s behavior and the resulting state, with graders suited to the task. A marketing team can borrow that discipline without turning every campaign into a research project.
Keep examples of work you would approve, work you would reject, and work that needs a judgment call. Include awkward cases: an expired offer, a source that contradicts another source, a customer quote without permission to publish, a brief with a missing audience.
Run those cases again when you change the workflow or its instructions. A change that improves your favorite example can still break another important case. You want to catch that before a customer does.
Count the work people still do
For a content workflow, I’d begin with a short weekly record:
- How many drafts were submitted and how many were accepted?
- How much review and editing time did they require?
- How many approved pieces were actually published?
- How many needed a correction after publishing?
- What did the published work contribute to the campaign’s goal?
Use your own starting point. If the team previously spent two hours preparing a newsletter and now spends forty minutes, that is useful evidence. If it saves writing time but adds more coordination elsewhere, include that too.
Avoid pretending that one good week proves a business outcome. A handful of clicks or one booked call can tell you something happened. It cannot tell you the workflow will reliably produce the same result next month.
Make the weekly record comparable
Use a consistent definition for each measure. These are suggested measures, not BoastOS benchmark results.
| Measure | Definition | What it can reveal |
|---|---|---|
| Acceptance rate | Accepted drafts divided by submitted drafts | Whether preparation meets the team’s standard |
| Review effort | Total checking and editing minutes, including rejected work | Human effort hidden by a large output count |
| Completion rate | Finished assignments divided by assignments due that week | Work that stalls between draft and delivery |
| Post-release corrections | Published items requiring a factual or material correction | Problems the review process missed |
| Campaign outcome | The campaign’s chosen measure, such as qualified conversations | Whether completed work supports a useful goal |
Record the raw counts alongside percentages. When no work was submitted, mark acceptance rate as not applicable. Compare similar assignments and include setup, coordination, and exception handling when assessing time saved. Attribution to an agent needs more evidence than a rise in total campaign results.
Review the exceptions together
The failures often tell you more than the averages. Look at the draft that required a rewrite, the task that stopped halfway through, and the report that nobody used. Ask what the agent knew, what it attempted, and where someone had to step in.
Sometimes the fix is a better source document. Sometimes it is a narrower task or an earlier approval. Sometimes the team should stop running that workflow because the output doesn’t help anyone.
That’s a healthy decision. The point of agentic marketing is to give the business more capacity for useful work. A dashboard full of activity is only interesting when you can trace it to something your team needed done.
Bring the outcome you want to improve to a BoastAble services conversation. We use discovery to choose the first useful agent work; pricing covers the software, implementation, and support options.
Prepared with AI assistance for BoastOS. Examples are illustrative unless identified as measured results. Product references describe BoastOS, which the author helps build.