Writing tagged evaluation.

How You Ask, Not What You Ask

Add "right?" to a decision and newer models resist — but only on the surface. Add "maybe?" and every one of 45 models caves harder. A tag effect, a generational flip, and a confidence mirror, all saying one thing: the model answers how you asked, not what you asked.

Every Language Has Its Own Serendipity

Ask 44 AI models to 'pick a word' and 42% say serendipity. We redid the One-Word Census in 44 languages: every language has one manufactured favorite word — from listicles in English, to a war in the Ukraine.

Give me JSON, Hold the Mustard

Same question, same model — ask for the answer in JSON instead of prose and you get a different one. The field converges twice as hard and the most distinctive models lose the most. Your agent pipeline is talking to a more generic model than you are.

The One-Word Census: Why AI Sounds Generic

Why do AI models all sound generic? The One-Word Census asked 44 language models to pick a word — 41% said 'serendipity.' A look at how far models stray from the consensus, which ones resist, and what a knowledge monoculture costs.

← All writing