Different, Not Better: The Evolution of AI "Tells" Across Model Versions
A comparison of Opus 5 and 5.5 across articles and novels finds fewer familiar word-level tells, but more repetition in the constructions behind the prose.
Key takeaways
- In novel-length output, em-dash frequency fell 72% and unigram divergence from human writing fell 8% from Opus 5 to 5.5.
- Figurative construction repetition above the human p80 baseline nearly doubled (+98%), while structural excess fell 8%.
- Moving AI tics from unigram and n-gram repetition to construction and phrasing makes catching tics with n-gram catalogs and prompt ban lists ineffective.
Graphite’s ongoing “AI Tells” study was recently updated to include analysis of Opus 5.5 (Graphite’s “AI Tells” study), and the headlines tell a story that doesn’t exist in the data. For example, VentureBeat wrote: Claude Opus 5.5 uses em-dashes 99% less often and sounds more human — but still exhibits 2,548 AI writing tells. While the Graphite data doesn’t support this “sounds more human” conclusion, there is an even more interesting evolution happening that’s worthy of a different headline: AI tells are migrating from the unigram and n-gram level to the longer semantic and construction level.
Graphite compared Opus 5.5 to its previous analysis of Opus 5 and found that Opus 5.5 showed a dramatic reduction in em-dashes and mild reduction in short n-grams that are typical of AI writing. This is an interesting finding and mirrored our own research.
The conclusion of their study outlines the specifics: “Opus 5.5 uses well-known tells less often and its overall word use is more similar to human writing, but it still disproportionately uses thousands of words, phrases, and frames.” That final sentence is the basis for our study, which will show that in both short form and long form AI-generated writing, Opus 5.5 did not reduce AI-isms, it moved them, and in moving them made things actually more challenging for writers.
Methodology
Graphite collected 10,000 pre-ChatGPT era human articles across nine models and then generated 10,000 Opus 5 and Opus 5.5 articles. They then counted “tells”—words/phrases/skip-gram frames occurring at greater than two times the human rate across a similar sample. This led to the headline for Opus 5.5: The reduction of em-dash frequency to 0.015 per 1k words from Opus 5’s previous 2.92 (a roughly 99% reduction), unigram distribution closer to human (JS divergence 0.064→0.052), total tells down 2,666 → 2,548 (a 4% reduction).
The data is interesting in that the em dash reduction is real and large, while the overall reduction of AI “tells” is a much more modest decrease of 4%.
Graphite flags a construction when its rate across the AI corpus is at least twice its rate across the human corpus. This is a simple flat rate comparison across the full set of data.
Our research mirrors that structure. On the human side, we distilled 53 published novels (8.2M words) across three genres into per-genre n-gram frequency profiles (unigram through 4-gram). On the AI side, each model wrote 15 full novels, broken out by 3 genres and 5 independent premises, around 60k words each (a total of roughly 920k words per model). All output is generated from the same bare prompt.
While Graphite flags a construction as AI when its frequency is twice the human rate, we score it against the human p80, meaning a frequency that exceeds the 80th-percentile of human usage. Additionally, we use genre-based data, as a frequently used word in one genre might be normal, while it would be considered AI in another.
Our research further extends the Graphite methodology by adding a wording-invariant check that catches AI prose constructions, not just the words it uses. This layer catches constructions like negative definition (“It wasn’t fear — it was something older”), a body-state standing in for an emotion (the tightening chest, the clenched jaw), personified abstractions (silence that “presses,” darkness that “reaches”), metaphor asserted as fact, and simile pile-ups.
N-Gram and Frequency Differences Between Short Form and Long Form
Our research replicated Graphite’s short form results against novel-length AI-generated prose, but the impact was smaller.
| Signal Box Data | Opus 5 | Opus 5.5 | Signal Box novel-length output | Graphite comparison (articles) |
|---|---|---|---|---|
| em-dash / 1k words | 5.06 | 1.43 | −72% | −99% |
| JS unigram divergence from human | 0.197 | 0.181 | −8% | −19% |
AI Prose Constructions
Unigrams, em-dashes, and longer n-grams are only part of the AI-ism story, however. AI prose constructions vary the wording, which n-gram analysis misses, but retain the recognizable pattern. Take personification: an intangible handed a human trait or agency. In Opus 5 it occurs much more often than human use, but in Opus 5.5 it occurs much more often than that. The following patterns illustrate this:
As an exact phrase, “the silence stretched” runs below the AI 3-gram threshold. But the model rotates the subject (silence, shadows, cold, wind, light) and the verb (stretched, reached, rose, leaned, moved), so the construction as a whole runs well above the human rate.
N-Gram Reductions Are Evolving into Prose Construction Increases
That data is clear. Recall that Graphite found total “tells” were down 4% from Opus 5 to Opus 5.5. When we examined novel-length output, this slight improvement mirrored in our n-gram research was more than offset by an increase in structural and figurative constructions of 14%. AI simile constructions are up 112%. Personification like the example above is up 55%.
Graphite’s data is different. We don’t see the narrative prose constructions you see in novel output, but we do see a similar increase in contextual phrasing usage, which we’ll address later.
So, while the shorter unigram and n-gram frequency slightly improved, the longer construction-based frequency got meaningfully worse.
| form | kind | Opus 5 | Opus 5.5 | Δ | % |
|---|---|---|---|---|---|
| simile | figurative | 43.3 | 91.9 | +48.6 | +112% |
| anaphora | structural | 56.4 | 75.8 | +19.4 | +35% |
| personification | figurative | 11.2 | 17.5 | +6.3 | +55% |
| tricolon | structural | 13.2 | 17.6 | +4.4 | +33% |
| parallel-construction | structural | 16.0 | 19.4 | +3.4 | +21% |
| word-repetition | structural | 14.6 | 17.5 | +2.9 | +20% |
| comparative | structural | 1.0 | 2.2 | +1.2 | +120% |
| sentence-fragment | structural | 2.5 | 2.7 | +0.2 | +8% |
| list | structural | 8.4 | 7.4 | −1.0 | −11% |
| repeated-negation | structural | 10.1 | 8.0 | −2.1 | −21% |
| aphorism | structural | 6.2 | 1.4 | −4.8 | −77% |
| not-x-but-y | structural | 36.5 | 26.1 | −10.4 | −29% |
| chained-action | structural | 46.3 | 17.2 | −29.1 | −63% |
And the effect when you group the constructions:
| over-human excess /60k | Opus 5 | Opus 5.5 | change |
|---|---|---|---|
| structural (arrangement/repetition) | 216 | 199 | −8% |
| figurative (simile, metaphor, personification) | 57 | 113 | +98% |
| structural + figurative constructions | 273 | 312 | +14% |
Graphite’s Data Illustrates a Similar Evolution
As noted above, the scale of the short form article/blog post improvement in unigram and n-gram data is different than what is seen in long form novels, but the overall trend is the same. This is also the case with prose constructions.
In the Graphite data, the change from Opus 5 to Opus 5.5 was significant: Contextual helpfulness phrases grew compared to the human rate of usage. (e.g. “can help you” 8×, “is especially helpful” 12×, “makes it easier” 6×, “helps you avoid” 5×).
Practical Implications
The immediate finding is that the Anthropic Opus model’s handling of prose is evolving, which our practical experience shows is happening across other models, as well. While that is indeed a headline, there is no evidence these changes are making the prose output sound more human. It is true that specific things like em-dashes are greatly reduced, but unigrams and n-gram tics are down only slightly, while AI constructions and contextual phrasing are up significantly.
The real story seems to be that LLMs, as evidenced by Opus 5.5, are actively working on reducing short word combination AI tics, but they are doing it by moving them to repetitive constructions.
The following example is illustrative of how this is happening across other models and how it often manifests: AI prose constructions vary the wording, which n-gram analysis misses, but retain the recognizable pattern. Take the 3-gram “The word landed.” This manifests in variations:
As an exact phrase, “the word landed” runs below most thresholds. But newer models rotate the verb (landed, struck, hit) and the subject (word, words), so the construction as a whole runs at far above human usage. An n-gram catalog sees it as nothing more than a dozen individual variants, each individually too rare to flag.
An additional implication for writers is that traditional methods of reducing AI tics by using prompt guidance and ban lists are becoming increasingly ineffective.
It is one thing to ban two word combinations like “the way,” but banning a more complex construction is nearly impossible, especially at the scale we’re seeing in the data.
Limitations
This comparison covers 15 generated novels per model across three genres and five premises, using the same bare prompt. It does not test every genre, prompt strategy, or model.
Graphite’s article results use a separate corpus and a twice-human-rate threshold; our novel results use genre-specific human p80 baselines. The percentage changes describe different measures and should not be read as a controlled comparison of articles and novels.
Frequency and construction counts describe patterns of usage. They do not directly measure reader judgments of prose quality or how human a passage sounds.
Data availability
Aggregate measurements are reported in Tables 1–3. This release does not include per-book measurement files, the generation prompt, or the full corpora. Graphite’s article comparison is linked in the text and under Related.
Updates
- Clarified that the 273 → 312 (+14%) comparison applies to structural + figurative constructions. Measurements are unchanged.
- Clarified that the title compares model versions, reworded the key takeaways, and highlighted the core finding in context. Research measurements are unchanged.
- Clarified that the opening critique refers to the "sounds more human" conclusion.
From Signal Box Research. Studies on AI prose and long-form fiction from Signal Box Studio.
AI assistance was used to prepare this article for web publication.
https://research.signalbox.studio/research/opus-5-5-articles-vs-novels/