A disclosure before anything else: every post on this blog moves through an AI-assisted pipeline before a human signs off on it. That's the product, and it's also why we know this topic from both sides. The tells below are the things our editorial gate exists to catch — and the things you'll see everywhere once you learn them, because most AI content ships without that gate. Here's what actually gives a machine draft away, what detection tools can and can't prove, and the part most marketers get backwards about Google.
The vocabulary tells
Language models have favorite words, and the preference is strong enough to show up in population-level data. A 2025 study in Science Advances ran an excess-vocabulary analysis across more than 15 million biomedical abstracts and found that at least 13.5% of 2024 abstracts showed signs of LLM processing — in some subfields closer to 40% — detectable purely from the abrupt frequency spike of certain style words. The poster child is "delve," but the family is bigger: tapestry, landscape, realm, pivotal, crucial, seamless, robust, leverage, foster, underscore, testament. None of these words is wrong. What's diagnostic is density — three of them in one paragraph reads like a machine because, statistically, it usually is one.
Wikipedia's editors, who deal with more unlabeled AI text than almost anyone, keep a running field guide called "Signs of AI writing," maintained by the volunteers of WikiProject AI Cleanup. Their vocabulary list overlaps heavily with the academic data: phrases like "stands as a testament," "plays a vital role," "rich cultural heritage," and "watershed moment" — importance-inflation language that asserts significance instead of demonstrating it.
The structural tells
Word choice can be edited in thirty seconds. Structure is harder to hide, and it's where the confident calls come from:
- Negative parallelism. "It's not just X — it's Y." One instance is a writing device; one per section is a template.
- The rule of three, everywhere. Models love triplets ("faster, cheaper, and more scalable") with a regularity human writers don't sustain.
- Trailing participle clauses. Sentences that end by editorializing on themselves: "…highlighting the importance of consistency," "…underscoring the need for a strategy." The Wikipedia guide flags these because they smuggle in significance claims without a source.
- Formatting on autopilot. Bolded bullet-point headers (yes, like these — done knowingly), section headings for 200 words of text, and em dashes at a density no human editor would let stand.
- The empty summary close. A final section that restates everything above and commits to nothing: "In conclusion, the landscape continues to evolve."
Tone tells travel with the structure: hedged both-sides-ism, "it's important to note," and vague authority like "industry reports suggest" or "observers have cited" with no named report and no named observer. Real writers cite; models gesture.
The technical tells
Some giveaways are absolute rather than statistical. Link URLs carrying ?utm_source=chatgpt.com mean the links were pasted straight from a chatbot session. Leftover placeholder text ("[insert client name]"), knowledge-cutoff disclaimers ("as of my last update"), or a stray "As an AI language model" mean nobody read the draft even once before publishing. And citations that don't resolve — a plausible-looking study title, journal, and year that simply doesn't exist — remain the most damaging tell, because a fabricated source doesn't just reveal the tool, it reveals that no human checked the claim.
What detectors can and can't tell you
The obvious question is whether software can just answer this for you, and the honest answer is: partially, and asymmetrically. The tools have improved — current vendors report document-level false-positive rates around or under 1% on longer text — but the history warrants caution. OpenAI shut down its own AI-text classifier in 2023 over low accuracy, and a study in Patterns found that detectors of that era flagged an average of 61% of essays by non-native English speakers as AI-generated; sentence-level accuracy still runs well behind document-level claims today. The practical rule: treat a detector score as a smoke alarm, not a verdict. It tells you where to look. The tells above tell you what you're looking at — and a human still has to make the call, especially before accusing a writer.
What Google actually penalizes
Here's the part that gets misquoted in half the briefs we see: Google's spam policies do not ban AI content. The policy that exists is scaled content abuse — "many pages generated for the primary purpose of manipulating search rankings and not helping users," with generative AI named as one way of doing that. The operative words are many, primary purpose, and without adding value. A hundred interchangeable posts stamped out of the same prompt is a spam problem whether a human or a model typed them. One well-researched post that happens to start as a machine draft is not. Which means the tells in this article matter twice: they're how readers lose trust in a page, and density of them across a whole site is exactly what "scaled abuse" looks like from the outside.
The tell behind the tells
Read the list back and notice what every item has in common: each one is an editing failure, not a technology failure. Favorite words survive because nobody swapped them. Fake citations survive because nobody clicked them. The empty conclusion survives because nobody asked "what does this paragraph do?" The dividing line in 2026 isn't AI content versus human content — it's reviewed content versus unreviewed content, and readers can increasingly taste the difference. If a draft has been fact-checked, sourced, cut by a third, and made to say something specific, it stops mattering where the first version came from. That's not a defense of the tools. It's the job description of the gate.