Ranking 12 Models by Their AI "Tells"
A selection of 36 AI-written novels finds 1,453 distinct constructions that models over-use past the human bar — and a gap of nearly 18× between the worst offender and the cleanest.
Key takeaways
- Across every AI novel in the corpus, there are 1,453 distinct word-level constructions that models collectively use more often than human authors do — and a reader meets one of them roughly once every 50 words.
- AI-ism breadth is model-specific, not a property of "AI Writing" in general. It spans an 18× range: gpt-4o alone accounts for more than half of the entire corpus's distinct tells, while gpt-5.6-sol has nearly engineered them out.
- The ranking is stable across genres. The same model tops the list in fantasy, romance, and science fiction, and the same two sit at the floor.
When writers sit down to generate and then edit AI writing, one of their core thoughts is: “Am I going to get the same AI results as everyone else?” Quickly followed by, “How bad is it?” The answer varies by model, but broadly speaking and encompassing 12 models over the past two years, the answer is 1,453 AI constructions exist in our generated sample of 11 million words.
That is the count of distinct two-, three-, and four-word constructions that the models in this study use more often than human authors do, pooled across 36 AI-written novels and deduplicated so that a phrase flagged in ten models and all three genres is counted exactly once. It is the total size of LLMs shared over-human prose mannerisms, the set of verbal habits that, collectively, mark prose as written by AI at the word level.
Note that this is different from semantic and longer AI constructions, which we discussed in this paper.
The volume of usage is common and explains why humans readily pick up on “AI writing.” Those 1,453 constructions appear 220,672 times across 11.0 million words of generated fiction. Normalized to a standard 60,000-word novel, that is 1,204 AI-isms per book—one roughly every 50 words. A reader does not have to go looking; the AI tells arrive at the pace of ordinary reading.
But the single number hides the more useful finding, which is how unevenly that vocabulary is distributed across the models that produced it.
Methodology
An AI-ism, for this study, is a 2/3/4-gram that a model uses at a higher rate than human authors use it in the same genre. The bar is per-n-gram: a construction is flagged only when its rate in the AI text exceeds max(human mean, human p80), the greater of the genre’s average human rate and its 80th-percentile human book rate. The p80 term accounts for common variance across authors: what’s rare for one author can be frequent for another. Each AI book is scored against its own genre’s human baseline, built from real published novels, so science-fiction prose is never judged against a romance bar.
The corpus is a controlled crosstab: 12 models, each writing the same slate of five fixed premises in each of three genres — fantasy, romance, and science fiction. The premises are held maximally different across genres by design, which is what lets the method separate tics from content without a hand-built stop-list: a world detail lives in one premise and one genre, but an AI construction recurs everywhere. Detection runs through our production n-gram detector and its presentation filters, which drop proper names, possessive debris, and pure function-word skeletons before anything is counted.
Three conventions matter for reading the tables:
- Unique n-grams is a type count — how many distinct constructions a model over-uses. It does not account for frequency.
- For a model’s overall figure, the three genre sets are combined as a set union (deduplicated), not summed, so a phrase that is a tell in two genres is not double-counted.
- occ/60K is the genuine frequency: how many times those constructions appear, normalized to a 60,000-word book — the same normalization used across our pipeline and human baselines.
The models, overall
Ranked by the deduplicated breadth of each model’s over-human vocabulary:
| Rank | Model | Unique n-grams | occ/60K |
|---|---|---|---|
| 1 | gpt-4o | 757 | 5,224 |
| 2 | minimax-m3 | 270 | 1,597 |
| 3 | opus-5 | 223 | 1,258 |
| 4 | deepseek-v4-pro | 200 | 1,051 |
| 5 | sonnet-5 | 185 | 1,187 |
| 6 | opus-5.5 | 167 | 958 |
| 7 | gemini-3.1-pro | 165 | 853 |
| 8 | glm-5.2 | 124 | 716 |
| 9 | qwen3.8-max | 123 | 584 |
| 10 | glm-5.3 | 118 | 776 |
| 11 | gpt-5.4 | 64 | 322 |
| 12 | gpt-5.6-sol | 42 | 181 |
The spread is the story. The oldest model, gpt-4o, is the heaviest in its usage of AI tells. Its 757 distinct tells are more than half (52%) of the entire corpus’s 1,453-construction vocabulary, and it uses them so relentlessly (5,224 per book) that it alone accounts for a third of all occurrences in a corpus of twelve models. It is the reference line for what unmitigated word-level non-human AI generation looks like.
At the other end, the OpenAI 5.x line is the cleanest by a wide margin: gpt-5.6-sol (42 constructions, 181/60K) and gpt-5.4 (64, 322) sit alone at the bottom. Within a single vendor’s lineage, from gpt-4o to gpt-5.6-sol, the breadth of word-level tells falls roughly 18×. Whatever changed between those releases did not merely trim a few favorite phrases — it reduced the entire over-human vocabulary to a near-floor.
The middle is where the models cluster, and where breadth and frequency come apart. deepseek-v4-pro carries a wider set of distinct tells than sonnet-5 (200 vs 185) but uses them less often (1,051 vs 1,187/60K): deepseek spreads a broad, genre-specific habit thin, while sonnet repeats a tighter set harder. The same divergence separates glm-5.3 and glm-5.2 — glm-5.3 has fewer distinct constructions but hits them more frequently. Breadth and intensity are genuinely different axes, and a cleanup strategy that targets one will not automatically fix the other.
By genre
Within a single genre there is no cross-genre dedup to do — each construction is already counted once — so these tables are the raw per-genre breadth and frequency.
The first thing to notice is what does not change. The ranking is similar across all three genres: gpt-4o ≫ minimax-m3 › opus-5, every time. The clean floor is identical too: gpt-5.4 and gpt-5.6-sol, every time. Only the middle of the pack reshuffles, and that reshuffling is small — which is exactly the signature of a real per-model fingerprint sitting on top of mild genre noise. Splitting by genre was what let us see that stability instead of obscuring it.
| Rank | Model | Unique n-grams | occ/60K |
|---|---|---|---|
| 1 | gpt-4o | 376 | 7,050 |
| 2 | minimax-m3 | 132 | 1,489 |
| 3 | opus-5 | 116 | 1,173 |
| 4 | sonnet-5 | 109 | 1,210 |
| 5 | opus-5.5 | 104 | 1,020 |
| 6 | deepseek-v4-pro | 103 | 1,075 |
| 7 | gemini-3.1-pro | 78 | 832 |
| 8 | glm-5.2 | 69 | 754 |
| 9 | glm-5.3 | 69 | 752 |
| 10 | qwen3.8-max | 55 | 543 |
| 11 | gpt-5.4 | 32 | 273 |
| 12 | gpt-5.6-sol | 21 | 203 |
Fantasy is where gpt-4o is at its worst — 376 distinct tells at 7,050 per book, nearly double its own romance figure. It is also the genre with the widest gap between first and second place.
| Rank | Model | Unique n-grams | occ/60K |
|---|---|---|---|
| 1 | gpt-4o | 267 | 3,831 |
| 2 | minimax-m3 | 139 | 1,611 |
| 3 | opus-5 | 126 | 1,235 |
| 4 | deepseek-v4-pro | 116 | 1,234 |
| 5 | sonnet-5 | 116 | 1,355 |
| 6 | opus-5.5 | 97 | 961 |
| 7 | gemini-3.1-pro | 94 | 934 |
| 8 | qwen3.8-max | 76 | 766 |
| 9 | glm-5.3 | 70 | 799 |
| 10 | glm-5.2 | 69 | 724 |
| 11 | gpt-5.4 | 45 | 427 |
| 12 | gpt-5.6-sol | 23 | 211 |
Romance compresses the field — gpt-4o’s lead is smallest here, and the mid-pack is tightly bunched. It is the genre where the models are most alike.
| Rank | Model | Unique n-grams | occ/60K |
|---|---|---|---|
| 1 | gpt-4o | 241 | 4,622 |
| 2 | minimax-m3 | 151 | 1,692 |
| 3 | opus-5 | 125 | 1,366 |
| 4 | sonnet-5 | 92 | 996 |
| 5 | opus-5.5 | 88 | 893 |
| 6 | deepseek-v4-pro | 87 | 845 |
| 7 | gemini-3.1-pro | 73 | 791 |
| 8 | glm-5.3 | 71 | 779 |
| 9 | glm-5.2 | 67 | 669 |
| 10 | qwen3.8-max | 47 | 442 |
| 11 | gpt-5.4 | 29 | 267 |
| 12 | gpt-5.6-sol | 14 | 130 |
Science fiction is the genre where minimax-m3 is at its worst, closing much of its gap to gpt-4o, and where the clean floor reaches its lowest absolute point — gpt-5.6-sol trips just 14 distinct tells across the whole genre.
Where’s Fable?
Readers who track the frontier will notice a model missing: Anthropic’s Fable-5. It was pulled from this comparison due to generation issues.
Fable refuses certain fiction premises outright, and we ran into that. Asked to write several of the novels in the slate, it declines rather than generates — which left its cells thin and, in science fiction, roughly 30% short on words compared with the models that wrote the full set. That is not a quality signal we can read as “fewer tells”; it is a different behavior entirely. So rather than footnote a misleading row, we removed the model. A like-for-like creative-generation study of Fable needs a premise slate it will actually complete, which we may tackle in a future study.
Limitations
This study measures one layer of the problem: word-level and short-construction over-use, the 2/3/4-gram band. It deliberately says nothing about the longer semantic constructions, the paragraph-shaped AI-isms that survive when the n-grams are cleaned up, which we have argued elsewhere are where “AI-ness” is migrating as the word-level tells recede. A model with a low score here is clean at the word level. But that does not mean it is free of AI-isms.
The corpus is five premises per genre per model, a controlled but narrow slice, and the human baselines are per-genre samples of published fiction rather than exhaustive. The p80 bar is a relax-only override on a mean baseline, not a single pooled-corpus threshold. “Above p80” means each construction cleared its own genre’s human bar in at least one cell. And none of this is a reader-perception study: it counts constructions, not whether a reader notices or minds them.
Data availability
Researchers who would like to access our raw and aggregated data files can request them at research@signalbox.studio.
From Signal Box Research. Studies on AI prose and long-form fiction from Signal Box Studio.
AI assistance was used to prepare this article for web publication.
https://research.signalbox.studio/research/aiism-by-model/