AI realism

Where the Humans Have to Be

I run six marketing functions on a quarter-million, a dozen agents and no direct reports, and I know exactly which judgments still need a person.

Company
A clinic management software company

There are two things I know about running a marketing organization on AI. Both were expensive to learn, and both run against what the people selling you the tools will say.

The first: humans sit at both ends of an AI system, never just the front. You need people to design the process before the machine touches it, and you need people doing quality control and feeding corrections back in once it runs. The human load shrinks over time as the system earns trust. It does not reach zero, and anyone promising you zero has not operated one of these systems against a live business.

The second is why. The reason you can never fully remove the human is structural, not a maturity problem that the next model release solves. These systems drift. Left alone, their output pulls toward a consensus middle, the average of everything, distinctive at first and beige by the fortieth iteration. And because every task is effectively fresh to the model, unexpected failures never stop arriving; they just stop arriving where you learned to look. Quality control around AI is therefore not a phase you complete. It is a permanent seat in the organization, and pretending otherwise is how companies get hurt.

Take those two things seriously and the org chart follows. An AI-first marketing organization is not the old organization with licenses bolted on. It runs on a different division of labor: machines carry production, people hold design and judgment, and quality control stops being a step and becomes a role. The work is deciding which judgments in your operation require a person, putting your best people on exactly those, and building everything else to run itself.

None of this comes from a conference stage. I run a marketing function built this way: inbound, advertising, content, social, email, a community platform and the revenue operations stack, on a quarter of a million dollars and a dozen agents, with nobody reporting to me. Machines carry the production across six functions. I hold the design and the judgment. And I know where that line between them sits because I have put it in the wrong place in production, twice, at real cost. That evidence is worth walking through, because the judgment is only worth something if you know what it cost.

What I inherited

I came in off the back of an exit, into a market I knew nothing about, having decided before my first day that I would build the function AI-first. HubSpot was the source of truth for the business, finance pulled from it, and years of different hands had left the usual sediment: live workflows, dormant workflows, automation nobody still at the company had written.

The structural problem took longer to see. One field, lead status, was doing five jobs at once: position relative to the demo, disqualification reason, terminal states, all of it. Lifecycle stage, the field built for exactly this, sat unused, because nearly everyone arrived through a demo request and was stamped a sales qualified lead on entry. There was no funnel above SQL, and nowhere to put a person who was interested but not yet buying. So thousands of them sat in a bucket called post-demo stored, untouched. The most valuable group in there was planning to open a clinic a year out, and with one salesperson carrying a hundred open deals, a prospect twelve months from opening was never going to be on his radar in month ten.

What I tried

The design was straightforward, and I had built its equivalent at three companies before: capture the planned opening date, calculate the distance to it, route the record on that number, so far-out prospects leave the pipeline for nurture and come back when the timing is real.

This was late summer 2025, before the connectors existed. I exported the instance, worked the design up in ChatGPT, took it to leadership, got agreement, and went to build. The calculations the design depended on turned out to sit behind a premium tier at close to a thousand dollars a month. ChatGPT had never mentioned the tier; it described the capability as available in the same confident register it used for the workflow configuration it also got wrong. I built around that. Then I turned the system on and enrolled every matching contact, reasoning that enrollment would move each one to its correct state. What I had not mapped was an existing demo-request cadence that changed lead status and lifecycle stage as a side effect. My workflows fired, that one fired back, and thousands of contacts landed in states they did not belong in, inside the system finance was reading from.

I rolled back to a day-old backup and paused everything.

The story I told myself first was that a model had lied to me. It had, and that is the least interesting fact in this piece.

I was arrogant. I had done this work before, so I stood it up without a sandbox, without mapping the existing automation, without testing a branch in isolation. The model gave me a plan I did not verify, and I gave it a live system I had not understood. I have a language technology background and had studied exactly where these models fail. Knowing did not stop me, which says something uncomfortable about knowing. Our CEO took it better than I did, these things happen, work in a sandbox, you now know the system better than anyone, and all of that was true, but being forgiven is not the same as having learned anything.

What made me accept the lesson was that it arrived a second time from the opposite direction. Off the CRM, I put my hands on advertising and content, generating at volume, and after a dozen iterations I went text-blind and picture-blind, skimming instead of reading, sending colleagues work with errors and machine tells still in it. That is drift meeting a tired reviewer, the exact failure mode I now build against. One burn is bad luck. The same lesson from two directions is when you accept it is about you, and it is why I hold the two judgments at the top of this piece as operating principles rather than opinions.

What I would do now

What both failures taught me is that these models multiply output, not productivity. They multiply productivity only when the design going in is good and the review coming out is good. Otherwise they make more of whatever you gave them, faster than anyone can inspect it.

So I rebuilt the system I had broken, backwards on purpose: starting at the end of the funnel, at customers, where the answer is already known and a mistake costs nothing, and working forward toward live sales. I had Claude map every workflow in the instance first, including the damage from my own failure. Records that needed fixing got moved by hand; enrollment is for what happens next, not for repairing the past. The rebuild went well past the original brief: lifecycle automation rewritten from scratch, lead status doing one job instead of five, cold records aging into nurture on their own, and, once the platform opened up enough, a small app I built with Claude Code against the HubSpot API to calculate each clinic’s launch window and flag the salesperson two months out, which is the capability that thousand-a-month tier had been holding hostage.

What matters most in that list is what I left alone. The salesperson updates three things by hand: the size of the clinic, when they intend to open, and what he makes of their chances. Those three cannot be automated, because the only way to know them is to have been on the call and formed a view. Everything downstream of them runs itself. That is judgment one, built into a live system: the human at the front designing, the human at the back holding quality, and the machine carrying everything in between.

And the quality seat has stayed filled, because the drift never stops. The checking gets lighter as parts of the system prove themselves, exactly as it should. It has never gone away, and I no longer expect it to. That is the honest shape of this technology as I have found it by operating it: a division of labor, not a headcount reduction, demonstrated here at the smallest possible scale. One person holding design and judgment, machines carrying six functions of production, and quality control sitting where it belongs, as structure. I know where the humans have to be because I have run the experiment on myself.