AI realism

AI's Most Dangerous Trick Is Making You Feel Faster

I will put my hand up. In the summer and fall of 2025, I moved too fast with AI, and it cost me.

I had not yet understood how much editing and checking AI output actually needs. So I outsourced part of my thinking to it and trusted what came back. I built workflows that looked right, that matched my experience of the work, and I presented them as finished solutions before I had verified they functioned. One of them fell apart on contact with reality. The calculations I needed in HubSpot turned out to live only in a premium add-on tier that made no sense for an organization that size to buy across the board. I had designed a clean system on top of a capability that was not really there, and I found out after I presented it, not before. I spent the following weeks building workarounds for a problem I should have caught on day one. A small tool I built later handles those calculations now. The scar is why it exists.

Here is what that taught me, and it is the whole point of this piece. The most dangerous thing about AI is not that it fails. It is that it fails while making you feel faster. I felt productive the entire time I was building toward a wrong answer.

And it is not just me.

Researchers at METR took experienced software developers, the kind who maintain serious open-source projects, and had them do real work with and without AI tools. Before starting, the developers expected AI to make them about 24 percent faster. After finishing, they believed it had made them roughly 20 percent faster. In fact, the AI tools made them 19 percent slower. They were slower and they could not feel it. They walked out of the experiment convinced of the opposite of what had happened to them.

Sit with that, because it is the whole skeptical case in one finding. The problem is not that AI is stupid. The problem is that your sense of whether it is helping is unreliable, and it is unreliable in the direction that flatters the tool.

The frontier is jagged, and you cannot see the edge

The best map of where AI helps and where it hurts came from a large experiment with management consultants. Inside the range of tasks the AI was good at, the effect was enormous. Consultants using it finished 12 percent more tasks, 25 percent faster, with over 40 percent higher quality. If the study stopped there it would read like a commercial.

It did not stop there. The researchers also gave people a task that sat just outside the AI’s real capability, one that looked similar but required reasoning the model could not actually do. On that task, consultants using AI were 19 percentage points less likely to get the right answer than the ones working without it. The tool did not just fail to help. It pulled competent people toward the wrong answer, because it produced something confident and plausible and they trusted it.

The researchers called this the jagged frontier. AI is spiky, brilliant at one task and useless at the task next to it, and the two are not labeled. Nothing tells you which side of the line you are standing on. The model sounds exactly as confident when it is right as when it is wrong. That is the part the demos never show, and it is the part that matters most for anyone about to put AI in front of a customer or a decision. It is also exactly the trap I walked into. My workflow sat one task past the edge, and it looked identical to one that would have worked.

What the skeptic gets right

So let me give the skeptic their due, because they have earned it. Even purpose-built AI, sold specifically for high-stakes professional work, is unreliable in ways the marketing does not mention. When Stanford tested specialized legal research tools, the kind that cost real money and promise accuracy, one still produced wrong or fabricated information more than 17 percent of the time, and another more than a third of the time. These were not free chatbots. These were the professional-grade tools, and they invented things in one out of six answers or worse.

The consequences are not hypothetical either. When Air Canada’s website chatbot gave a customer wrong information about a fare, the airline argued in a tribunal that the chatbot was a separate entity responsible for its own statements. The tribunal rejected that outright and held the company liable for what its AI told a customer. The lesson for any business is blunt. Your AI’s mistakes are your mistakes. You do not get to blame the model.

And the aggregate picture backs the skepticism. Near-universal adoption, very little bottom-line impact so far, a large share of ambitious projects predicted to be abandoned or canceled. If you have felt like the hype outran the results, you were not being a Luddite. You were reading the data correctly.

The barbell is what I use now

Here is where I part ways with the pure skeptic. Everything above is an argument for using AI with your eyes open, not for avoiding it. The gains inside the frontier are not small, and they are available right now to anyone who stops trusting the tool blindly and stops refusing to touch it.

My mistake in 2025 was not using AI. It was letting AI into the middle of the work, where the judgment lives, and stepping back from both ends. So I flipped it. Now I run what I think of as a barbell. Humans design the system. AI does the work. Humans edit and confirm at the end. The weight sits on the two ends, where judgment matters most, and AI carries the load in the middle, where it is strong.

That structure would have saved me the whole HubSpot mess. If I had confirmed the system actually functioned before I presented it, the confidence the AI handed me at design time would have hit a human check at the other end of the bar, where output meets reality. I skipped that end because the work felt done. It looked done. Done and confirmed-done are not the same thing, and only a person can tell them apart. That is the entire job the far end of the barbell exists to do.

None of this is glamorous. Know what your AI is reliably good at, and what it only appears good at. Keep a person at both ends of anything that reaches a customer or a number that matters. The enthusiast trusts the confident wrong answer. The skeptic misses the real gains sitting one task over. Both are half right, and both halves are dangerous alone.

So the question is not whether you believe in AI. It is sharper than that. On the work that actually matters in your business, do you know where your AI stops being reliable, and is there a human standing at that line right now, checking? I learned where mine was the expensive way. You do not have to.