a16z says the better test for AI writing is whether the text actually does its job

a16z says the better test for AI writing is whether the text actually does its job

N
News Editor
2026-08-26 06:04:19
Steph Zinn of a16z Crypto argues that the current obsession with spotting "AI tells" rests on two shaky assumptions: that machine-written text can be identified through a stable set of giveaways, and that anything carrying those markers must be weak writing. In the essay, published in Chinese by TechFlow, she rejects both ideas and proposes a simpler standard: does the piece of writing actually accomplish what it is supposed to do? Zinn breaks common AI-style problems into four buckets — rhetoric, voice, structure, and punctuation — and says many of them long predate large language models. AI has not invented those habits so much as scaled them. Across the piece, she addresses founders and writers who use LLMs in practice, outlining when vague phrasing, generic tone, overbuilt structure, and repetitive punctuation get in the way of clear communication, and when some of those traits still make sense. Rather than treating AI as something to hide at all costs, she frames it as a tool that can help people draft, organize, and research, so long as the writer still edits for specificity, clarity, and purpose.

Steph Zinn of a16z Crypto argues that the debate around AI writing often gets stuck on two assumptions: that machine-generated text can be caught through a recognizable set of flaws, and that any text carrying those traits must be poor. Her answer is no on both counts. The more useful question, she writes, is a simpler one: does the piece of writing actually do its job?

The essay, published in Chinese translation by TechFlow, is not mainly about catching AI. It is aimed at founders and writers trying to decide which habits hurt communication, which ones only look suspicious, and which ones can stay depending on the format.

Rhetoric: sentences that look polished but say very little

Zinn starts with a familiar editing experience: a sentence sounds fine, maybe even better than the lines around it, but something is off. After moving words around, the real problem comes into view. The sentence never meant much in the first place.

She points to a recent self-analysis from Fable on punctuation style: 「Human punctuation has a body to it; it fidgets. Mine is uniformly deliberate, and that deliberate neatness is itself rhythm.」 For Zinn, people often react to writing like that by saying no human could have written it. She pushes back. As an editor, ghostwriter, and lifelong fantasy reader, she says humans have always been capable of producing this kind of language without any help from AI.

The trickier version is less theatrical. It shows up in lines such as 「Something truly important is happening.」 「This matters.」 or 「Its implications are profound.」 These statements pass as sentences, but they are hard to paraphrase into anything concrete. Zinn says AI generates them in bulk because they function as plausible-sounding connective tissue without carrying much meaning.

Before these habits were linked to AI, they already had another label: corporate language. She cites examples such as 「We are excited to announce...」 「We are at an inflection point」 and 「This is the next chapter in our journey.」 The phrases sound weighty, but often fail to point to anything specific enough to survive retelling.

She does not like calling this laziness. More often, she says, it is first-draft language: the stock phrases and cliches that float closest to the surface and arrive first on the page. That is not a crime. Stopping there is the problem.

Her fix has little to do with whether AI was involved. If a sentence feels wrong, uses familiar filler, cannot be immediately understood, or simply has the wrong vibe, rewrite it. If another version is clearer, keep that one. If the sentence boils down to 「things exist」 or 「something is changing,」 delete it and see whether the piece still stands.

Rhetorical tells she flags

  • Ready-made profundity. Lines like 「Something truly important is happening」 or 「The stakes could not be higher」 sound important but often work as filler.
  • Empty contrast. A structure like 「This is not just about X — it is about Y,」 where Y turns out to be even vaguer than X. She says it can work when the contrast is sharp, but often slows the sentence down.
  • Hedges. Phrases such as 「in many ways,」 「to some extent,」 or 「it could be said.」 Their job is to make the sentence harder to prove wrong. She notes that in fields such as finance and healthcare, some hedging is necessary and compliance-related.
  • Over-symmetry. Bullets that mirror each other too neatly, with the same syntax and the same length. This is a common AI tell, but it is also just dull writing.
  • Summary phrases. Expressions like 「At the end of the day」 or 「Once the dust settles.」 She says deleting them often changes nothing in the paragraph.

Her editorial stance here is blunt: most of this should be cut. Empty, meaning-light prose sits at the top of the list of writing traps. When she uses LLMs in this area, she tends to use two kinds of prompts: ones that push the text toward more specificity, and ones that push it toward more directness.

She also lists practical ways to clear filler out of a draft: build a rough style guide with positive and negative examples of what good writing looks like to you; ask AI to run a paraphrase test by producing a duller version of what you wrote; or have an LLM rewrite prose at a sixth-grade reading level as a crude but useful way to strip ornament away.

Voice: the default AI register sounds interchangeable

Zinn says the better test for AI-ish diction is not whether a writer used the latest suspect word. It is fungibility. Can a sentence be lifted word for word from your article and dropped into someone else’s piece on a totally different subject without anyone noticing? If the answer is yes, the line may not belong to you in any meaningful sense.

She notes that people did spend time tracking low-friction words associated with AI output. Researchers noticed a jump in the word 「delve」 in academic abstracts after ChatGPT launched in 2022. Blacklists followed, featuring words such as tapestry, testament, and underscore, then later load-bearing, scaffolding, and broader.

Still, her point is larger than detection. For anyone building or running a company, wording matters because so much marketing now depends on personality. The strongest founder writing feels personal and expressive. She says readers can easily imagine the prose styles of Brian Armstrong, Vitalik Buterin, and Chris Dixon: concise, sincere, and unironic in Armstrong’s case; technical, wandering, and dense for Buterin; restrained and philosophical for Dixon.

Her warning is straightforward: do not hand your personality over to the model. Your own lexical choices, including their imperfections, create a texture that a machine cannot reproduce.

She also concedes that practical advice on finding the right word is scarce. The task often has less to do with technical precision than emotional or atmospheric precision. Even E.B. White, she notes, admitted that nobody can fully explain why some words ignite readers and others do not.

Her advice is to begin by asking what exactly you like in the writing you admire. Learn a new word, or use an old one in a fresh context. Match word choice to the atmosphere you want to create. If you are explaining something obscure, smaller and simpler words may help. If you want readers to smell trouble, scheme may work better than plan.

Voice tells she flags

  • Generic warmth. 「Great question!」 in the tone of a customer support reply.
  • Reusable sentence frames. Phrases like 「A useful way to think about this is...」 「The core idea is...」 or 「This can be understood as...」 that travel too easily between contexts.
  • Low-friction diction. Every word feels just right in a suspiciously smooth way.
  • Abstract nouns. Heavy reliance on words such as efficiency, complexity, society, communication, and innovation.
  • Weak motion verbs. Navigate, leverage, unlock, nurture, shape, elevate, streamline, and similar terms that imply movement without doing much.
  • Beacon language. 「A testament to...」 「A beacon of...」 or 「A reminder that...」
  • Vague nouns. Terms like landscape, space, journey, ecosystem, and tapestry that circle around the thing instead of naming it.
  • Blurry intensifiers. Phrases like 「very important,」 「major impact,」 or 「critical role」 when nothing concrete follows.

Most of the time, she says, these should be edited. Not always. There are cases where what she calls the Alexa voice is useful: support docs, error messages, terms of service, safety notices, and apologies sent to huge numbers of strangers. In those settings, personality creates friction, while neutrality can be kinder and more effective.

Her practical suggestions focus more on detection than generation. An AI detector or any LLM can be used to mark generic words and phrases in a draft. Writers can then revise against those flags. She also recommends running a transplant test sentence by sentence: if a stranger could plausibly claim it, consider rewriting it. For vague language, she offers a concrete example. 「The Ethereum ecosystem is expanding」 can become something more precise, such as developers building more wallets, exchanges, and lending markets around Ethereum.

She also points to voice-note workflows, where people feed spoken notes into an LLM to turn ideas into draft text. One underappreciated benefit, she says, is that this can reveal personal quirks in a writer’s voice. The result still needs editing. There is no one-shot solution.

Structure: too much form can start working against the idea

When left on its own, AI often produces what Zinn calls overstructured writing: too many H2s, too many bullets, and paragraphs broken into thin slices. But she does not treat this as uniquely machine-made or wholly bad. People have long borrowed familiar structures because they work and because they are comfortable to read.

We split arguments into three points because three feels complete. We add signposts like 「first」 or 「in other words」 so readers can orient themselves quickly. We build tidy categories because they are easy to scan. Many writers learned some version of the essay hamburger at school: say what you are going to say, support it in distinct sections, then restate it at the end. That structure helps force clarity, evidence, and readable sequence.

The trouble begins when a default structure starts pressing the writing into a shape the idea does not actually want. For Zinn, structure is a set of decisions: what container to use, what matters most, which ideas belong together, and how sections should be named. The right choices depend on the job the piece needs to do.

She lays out examples. Narrative nonfiction needs discovery and tension to pull readers forward. Product launches need extreme efficiency because the reader may be skimming while scrolling. Explanatory writing needs a sequence that layers new ideas on top of concepts already learned.

Her suggestion is to ask, preferably at the start, what format best suits the idea. Then borrow from formats that already work. If you are writing an argument, study how other opinion writers organize similar pieces. If you are writing technical explanation, take apart one of the best examples you have read and see how the author arranges information. Once you have a template, adapt it to your own material.

Structural tells she flags

  • Over-organization. Too many subheads, bullets, numbered sections, or mini-frameworks.
  • Always landing on three points. She does not reject the rule of three, but rejects forcing ideas into it.
  • Familiar article shapes. Broad intro, explanation, example, warning, conclusion — the default template sequence.
  • Formulaic openings. Lines such as 「In today’s rapidly changing world...」
  • Too many signposts. First, next, finally, in conclusion, here is the breakdown, let’s unpack this.
  • Stiff transitions. For example, 「To understand why this matters, we first need to look at...」
  • Section previews. Such as 「There are three key reasons...」
  • Bullet lists. She notes that AI detectors often flag them unfairly. Lists are useful, but one recognizable move is the bold lead-in before the bullets.
  • Short fragments. Common in dramatic LinkedIn-style writing. Short. Punchy. Often in threes.
  • Restatement endings. Conclusions that simply repeat the article in different words.
  • Moralizing endings. A vague final line about progress, the future, or what we can learn.

She says small headers that do not fit the form should usually be removed, especially in opinion, personal narrative, and articles driven by rhythm and voice. The same goes for sections that are weak, redundant, or distort the meaning of the piece. She is especially skeptical of signposts attached to what she calls self-evident structure, like announcing 「three reasons」 when the line adds little beyond padding.

That does not mean structure should be avoided. It should be used when it matches the piece. Commentary often benefits from a recognizable argumentative track. Lists, explainers, and how-to guides often need headers and subheads. If an idea truly comes in three parts, then three parts is fine. She also notes that LLMs and search engines reward well-structured information. That makes clear organization especially useful for reference content people revisit rather than read straight through, including documentation, guides, and FAQs, and for readers who scan before they commit.

On the practical side, she says writers who do want stronger organization can ask models for help directly. Writers who do not want the default framework should specify what the structure needs to do, whether that is creating suspense or making the piece easier to scan. Another approach is to feed the model a published article in the same genre and ask for a reverse outline describing the role of each paragraph, then use that skeleton for a new draft.

Even when the model does not return something usable, she says the extra thinking involved in organizing and presenting an idea often pays off.

Punctuation: the em dash is not the enemy

The final section turns to punctuation, especially the em dash, which has become one of the most widely cited AI tells. Zinn writes that the mark has always been controversial. Strunk and White generally recommend restraint, using em dashes only when more common punctuation will not do. But she argues that em dashes have a relaxed, conversational quality that other marks do not easily match.

She acknowledges the criticism. Em dashes interrupt the sentence and insert a second track of thought, which can distract. Once they started getting associated with weak writing, they became an easy symbol to ban.

Her position is that it does not matter whether AI happens to favor them. Some situations genuinely call for them. She points to longer parenthetical material and sudden turns in thought or feeling as examples where the em dash may be the best tool available. She also mentions horror writer R.L. Stine as an admirer of that effect.

One of her more memorable lines is personal: Shift + Option + dash on macOS is the messy, absent-minded, often tangential child of human writing, and she is not giving it up. Trends change. Writers should not let anxiety over AI dictate punctuation choices.

She notes that as people rush to avoid em dashes, AI itself has started swerving toward colons. If no punctuation mark is permanently safe, then the only sensible option is to use the one that best serves the sentence. Her short answer is this: choose what is most correct and least distracting to the reader. Style disputes such as the Oxford comma will continue. Better to make sure punctuation supports meaning and does not irritate readers than to chase a fantasy of purity.

Another issue she raises is sameness. The problem is not any one mark but what happens when every sentence starts to look and sound alike. LLMs often generate chains of colon-led lists. Multiple sentences interrupted by em dashes can do the same kind of damage, no matter who wrote them.

Her test is simple: read the writing out loud and treat punctuation as stage direction. If it sounds unnatural to the ear, readers are likely to feel that too.

Punctuation tells she flags

  • Colon-heavy constructions. 「The problem is:」 「The result is:」 「The key point is:」 especially when a grocery-list sequence follows.
  • Dense clusters of em dashes. The issue is not the mark itself but the concentration, especially when combined with lists and asides. She says there is no hard numerical limit, but a sentence should generally not carry more than two.
  • Playful parentheses. The sort that signal self-awareness or joking aside, as in 「like this.」
  • Performative semicolons. She cites Kurt Vonnegut’s cutting line about semicolons while still acknowledging that they do have legitimate, if relatively rare, uses.

Her standard is functional rather than ideological. Edit punctuation when it repeats, distracts, or fails the read-aloud test. Keep it when it works. The goal is grammatical correctness with as little reader friction as possible.

She also advises against obsessing over whether a punctuation choice will make you look AI-assisted. If a dash, colon, or any other currently unfashionable symbol is the right choice, use it. Rather than asking a model to remove all em dashes, ask it to default to periods and commas and only reach for special punctuation when grammar requires it. She also suggests feeding a model examples of your own writing, or writing you admire, so it can estimate a punctuation fingerprint and usage frequency.

The real question is still whether the writing works

Zinn closes by saying that as more people use these tools, the question of whether something was machine-generated is starting to lose force. The answer will usually be yes, to some degree. She still argues that bad AI writing needs to be exposed, whether by ordinary editorial standards or with Proof of Person-style tools. She gives examples of cases where disclosure matters, where the named author is itself the point, or where one person is pretending to be a thousand.

For almost everything else, though, she returns to the old question: does the writing do its job?

AI, she argues, helps people express and publish ideas that might otherwise never be written down. It saves time on organization and research. It may even help people become better writers. Treating the ability to express oneself as automatically embarrassing or low quality, in her view, is mean-spirited. If a piece works as intended, she asks, how much does it really matter which part of the process showed traces of machine assistance?

Her final challenge is aimed at the fear itself. Why, she asks, are people willing to surrender an entire punctuation mark — one that has existed since the early era of print — just to avoid looking as if a machine helped? That is an absurd concession, she argues, especially when these tools are already being used frequently and will continue to be.

The acknowledgments credit the a16z crypto editorial team — Tim Sullivan, Robert Hackett, and Sonal Chokshi — for feedback, along with the many years of editing and discussion that shaped the piece.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
2000

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.