A content strategist at Previsible ran the experiment most agencies talk about and never actually measure. She took three ways of producing an article / written entirely by AI, written by AI and cleaned up by a human, and written by a human with AI only in the wings / then watched all three in Google Search Console for the better part of a year.

The pure-AI pages, published in April 2025, did fine for a moment. By January 2026 they had all but vanished from search. The hybrid pages needed a full human rewrite before they moved at all. The human-written page climbed all winter and kept climbing.

We build AI marketing systems for a living, so read the finding carefully, because it is not "AI is bad at SEO." It is narrower and far more useful: efficiency helps right up to the sentence where original thinking was supposed to start, and past that line it does damage. The entire question is where you stop the machine.

Where the machine is genuinely excellent

Start with the part the article gets right and most AI skeptics skip. In Jenny Lee Lynch's own workflow, Gemini took more than 2,000 declining page-one keywords, sorted them into topical clusters, then mapped those clusters against Search Console data to produce a short list of URLs actually losing visibility. Hours of spreadsheet work, compressed into minutes, with no quality cost whatsoever.

That is the shape of the good use case. Sorting, clustering, cross-referencing, summarizing, first-pass pattern-finding across a dataset too big to eyeball. Nobody is going to out-tab-a-spreadsheet a language model, and pretending otherwise is just expensive nostalgia.

Note what the machine produced in that example, though. It produced a list of pages worth a human's attention. It did not produce the pages. The output was leverage, not inventory. Her line for it is the cleanest summary of the whole discipline: when AI supports the background strategy, it buys you the time to make a better page. When AI writes the page, the strategy is what you spent.

The benchmark, three ways

Pure AI, no human review. Three pages, published April 2025 from simple prompts. They generated some organic performance out of the gate and then decayed, and by January 2026 they were effectively gone from search results. Worth naming what that actually costs: not just the missing traffic, but nine months of a URL, an internal link budget, and a crawl allowance spent on something the site is now quietly worse for hosting.

AI-generated, human-edited. The middle ground where, by the figure Lynch cites, more than 86% of marketers now live. A human tidies the headings, fixes the grammar, adjusts the formatting. Five articles started life this way. They only moved after being rewritten entirely by hand, and even then the gain was modest: clicks up 12% and impressions up 27% year over year across the following three months.

That is the most instructive result in the set, and it is the one that will get skipped. Editing did not save the hybrid pages. Rewriting did. Grammar and formatting are not what was missing, so fixing grammar and formatting did not fix them. Most "human in the loop" content programs are polishing the exact layer that was never the problem.

Human-written, AI in the wings. Built from a real perspective / messy internal notes, a documented customer problem / with AI used for brainstorming and proofreading only. It gained visibility through the winter and peaked with substantial click growth in early April. Lynch pairs it with a separate finding that human-written content ran roughly eight times more likely to hold the number-one position.

One practitioner's client set, small sample, honest about it. But the direction matches everything Google has shipped for two years, which is the part that should move your budget.

Google stopped grading effort and started grading novelty

The helpful content system is no longer a bolt-on / it is folded into core ranking, and the scaled content abuse policy exists specifically to describe publishing at volume with nothing new in it. The March and May 2026 core updates pushed the same direction again: original research, firsthand experience and a distinctive voice up; large volumes of low-value pages down.

Strip the policy language and the mechanism is almost boring. Google's systems are built to surface information that is not already indexed. A language model, by construction, produces a fluent average of what is already indexed. You are asking a machine trained on the existing corpus to generate an addition to the corpus, then asking a ranking system to reward it for being new.

That is not a prompt problem. There is no prompt for firsthand experience. The model does not have your service records, your last four quarters of customer complaints, the reason your best technician quit, or the pattern you noticed in September that nobody else in your category has noticed yet. That inventory exists in exactly one place, and it is the thing search is now paying for.

The tell your readers hear before Google does

Lynch publishes a negative-prompt list: words to explicitly ban when you use an LLM to organize your thinking. Delve, tapestry, paramount, pivotal, synergy, harness, holistic, robust, seamless integration, transformative, game-changing, unwavering commitment, navigate the complex, pave the way. Dozens more.

We are sympathetic to this in a way that costs us nothing to admit: we banned the em-dash on this site before it was fashionable, and everything you are reading obeys that rule. So we will say the quiet part with some authority. The word list is a symptom sheet, not a cure. Banning "delve" does not add a perspective; it removes a fingerprint. The reason those words cluster in machine writing is that they are what fluent, confident, contentless prose reaches for, and the machine is very good at fluent and confident.

Her point about humanizer tools is the same argument one layer up. Running machine output through a second machine to disguise the first machine trades a recognizable pattern for a different recognizable pattern, usually a clumsier one. Readers still hear it. You cannot launder your way to having something to say.

Corporate marketing was already full of empty language before AI showed up. AI just made empty language free, and free things get overproduced.

The playbook

This is a process design problem, which is the work we actually do. DMAIC at the strategic layer, disciplined delivery underneath. Six things we would change on Monday.

Draw the line on the actual assembly line. Write down, per content type, exactly which stages the machine owns and which stage it never touches. Research, clustering, competitive scan, outline scaffolding, proofreading: machine. The argument, the example, the number nobody else has, the opinion someone could disagree with: human, always, no exceptions under deadline. A line that moves when the calendar gets tight is not a line, it is a preference.

Feed the human stage, not the prompt. The strongest content in the benchmark started from messy internal notes and a documented customer solution. So build the intake for those: a standing habit of capturing support tickets, sales objections, install-day surprises, the questions asked twice in one week. Most companies have overwhelming original material and no pipeline for it, then buy a subscription to generate the thing they already own.

Stop paying for edits and start paying for rewrites. The hybrid result is your budget lesson. If the plan is "AI drafts, human polishes," you are funding the one intervention the data says does the least. Either the human brings the substance or the exercise is decorative. Fund the smaller number of pieces properly.

Measure at nine months, not nine days. Every method in the benchmark looked survivable at launch. The pure-AI pages performed and then decayed for the better part of a year. If your content dashboard reports first-30-day traffic, it is structurally incapable of showing you the failure mode / it will keep certifying the approach that quietly loses you the site. Cohort your pages by production method and by publish month, and look at them at 90, 180, and 365 days.

Publish less, on purpose, and say so out loud. Volume was a strategy when volume was expensive, because expense was the moat. Generation is free now, which means volume is free for your competitors too, which means it is no longer a moat / it is a cost with a ranking penalty attached. Cutting the calendar in half is a defensible operating decision, and it needs to be made by someone senior enough that nobody quietly refills it.

Prune what already shipped. If you ran a pure-AI program in the last two years, those URLs are still sitting on your domain being evidence. Inventory them, then rewrite, consolidate, or remove. This is unglamorous and it is usually the highest-return work available on a mature site.

Our position

The uncomfortable implication of this benchmark is that AI made the cheap half of content production free and left the expensive half exactly as expensive as it was. Drafting got fast. Having something worth saying did not get any faster at all, and the market rate for it just went up, because scarcity moved.

Which is the real reason the efficiency story keeps disappointing people. Teams modeled the savings against total content cost, when the machine only ever touched the part that was never the bottleneck. Take the time the machine hands back and spend it on the stage it cannot do / interviewing your own operators, writing down the thing you learned the hard way, being specific enough to be wrong. That is what the ranking systems are now paying a premium for, and it is the one input with no vendor.

Use the machine to get to the starting line faster. Do not ask it to run the race.