Japanese Business author Masakazu Kobayashi recently tested the AI writing detector Pangram with two Japanese-language samples and got opposite results. A Japanese article generated by Google Gemini was labeled "100% AI-generated," while a manuscript Kobayashi had written himself was labeled "100% human-written."
The test comes after Pangram drew attention through Substack's adoption of the tool, which fueled debate over whether paid newsletters were relying heavily on AI-assisted ghostwriting. ABMedia said Japanese media then tried the system with Japanese text.
Two Japanese samples, two opposite outcomes
In the first test, Kobayashi asked the free version of Google Gemini to generate a Japanese article and then pasted that text into Pangram. The result came back as "100% AI-generated."
He then submitted a document that could be identified as human-written: the afterword manuscript for his 2015 book, AIの衝撃 人工知能は人類の敵か. Since ChatGPT was not publicly released until November 2022, the report said a manuscript completed in 2015 can at least rule out the use of contemporary large language models such as ChatGPT. Pangram gave that text the reverse verdict: "100% human-written."
ABMedia said that, in this very small Japanese-language test, Pangram managed to distinguish between Gemini-generated content and human original writing produced before the current generative AI wave.
Not enough evidence to claim Japanese accuracy
The report also makes clear that two samples are nowhere near enough to prove Pangram reaches its claimed 99.99% accuracy rate in Japanese.
Because Pangram's free trial only allows two checks, Kobayashi did not buy a paid plan to run more tests. He also warned against drawing a definitive conclusion about Pangram's Japanese-language detection ability from just those two examples.
What the test does suggest, according to the report, is that Pangram can handle Japanese text even though the product is mainly offered in English.
How Pangram says it detects AI text
Pangram previously came under the spotlight after being introduced by Substack. Compared with older AI detectors that often misclassified human writing as AI output, Pangram has claimed 99.99% accuracy and said it was designed to reduce false positives.
As for how it works, the company's public explanation remains limited. The developers have only said the system combines "many weak signals" emitted by AI-generated text.
Kobayashi said that answer reveals almost no technical detail, and the real detection method is likely treated as a trade secret. The report adds that this matters because AI detectors have long faced an explainability problem: users often cannot tell what features a model is relying on when it decides a text was generated by AI. Higher stated accuracy does not remove that opacity.
The harder question is the gray zone between human and AI work
ABMedia said the broader issue has become more complicated than simply asking whether a piece was written by AI. If Gemini produces an entire article from scratch, detection may be relatively straightforward. But many real-world editorial workflows are mixed. A writer may do the interviews and research, then ask Claude to organize material. Another may build the argument and use ChatGPT to rewrite sentences. In other cases, AI may draft a first version and a human may heavily revise it.
That shifts the core question from "Was this written by AI?" to "How much of the work in this article was done by AI?" The report said Pangram stands out in part because it does not only offer a binary AI-versus-human label. It also claims it can estimate the proportion of AI and human text within the same document.
Why paid publishing is especially sensitive
Japanese Business also said that after Substack introduced Pangram, some previously popular posts were flagged as relying heavily on AI. The issue becomes more sensitive when authors charge subscription fees, because readers may object if they believe they are paying for human-created work but receiving AI-generated material instead.
The original report did not provide full statistics, a list of creators, or the scale of those cases. Because of that, it cannot be used to determine how much paid content on Substack is heavily AI-generated.
The article ends on a narrower point: now that AI has become part of many writers' daily toolkits, "using AI" no longer marks a clean boundary. Using AI to catch typos, organize interview transcripts, or help with research is not the same thing as asking Gemini to generate a full article, even if all of those cases involve AI.

