Quick Answer: While Pangram claims an exceptionally low false-positive rate, my testing produced several clear misclassifications. In a 10-document test, Pangram incorrectly classified fully human text as 100% AI-generated, and AI-generated text as 100% human. AI detectors are probabilistic statistical models, not definitive proof. Their integration into browser extensions that automatically label professional content on platforms like LinkedIn poses serious risks to professional reputations, authors, and academic standing.

I want to start this article by saying the following: I understand and appreciate what Pangram is trying to do. Being able to identify AI-generated content can be genuinely useful. Schools, publishers, companies, and individuals may need to determine if something was AI-generated or not.

But these tools can create dramatic consequences when they are incorrect. They can be used to accuse a student of cheating, question an author’s work, evaluate a candidate, or judge someone’s professional credibility.

But Pangram is different… isn’t it?

What is Pangram’s false-positive rate?

Pangram was recommended to me by a friend as one of the best AI detectors available.

And looking at Pangram’s own website, it seems great. Pangram currently claims that its detector is over 99% accurate and reports a false-positive rate of approximately 1 in 10,000. That is very impressive.

It is also interesting that Pangram specifically emphasizes false positives because, in its own words, incorrectly claiming that someone’s work is not their own can damage that person’s reputation or academic standing. So we both agree on that, which is great, but what happened next was not.

Is Pangram accurate? My testing results

I simply gave Pangram part of a book I recently wrote while on holiday. All by myself. But Pangram classified it as 100% AI-generated. Not a possibility, not even partially, but a strong orange-reddish claim that “this text is AI.”

I then decided to just write the introduction of another book I was going to write. Straight into Pangram. But that too was classified as 100% AI-generated. I even kept writing, making spelling mistakes, forgetting some words and calling the AI detector a piece of crap, but it kept flagging it as AI-generated. Maybe I am a robot after all?

Here is part of the text I tested, including the mistakes:

I think that nobody grows up dreaming to work in a quality department.

Kids want to be … well a lot of things, but nobody dreams of writing a non-conformance report or auditing a supplier. Quality is seen as the place to fill forms, that like to say no, and that slows everyone down.

If you have that point of view, I hope to change it.

Because thing with quality is that when it works, you do not see it. There are no recalls, no angry customers and no batch thrown in the bin. Nothing happens and “nothing” is the sign of success. Quality prevent fire from starting, but no one gives a medla to the person who prevented a fire that did not happened.

But the most serious question is to know if Pangram is a really good detector, or just a piece of crap. It is very possible that it does not do a proper job at detecting ai text, which would be a problem, considering how good they claim they are.

And you would not believe so, but it did. It flagged my fully human made text as an AI. Something that is hard to explain, or maybe simply mean that I am robot, in the end.

There was a possibility that if I keep writing, it would find the reality of things. But chances seemd slimmer by the minute.

After getting these results, I tried the opposite experiment. I wanted to write a LinkedIn post, so I gave ChatGPT my bullet points and asked it to write the post for me. I put the resulting AI-generated text into Pangram. The result: 100% human.

Dear Pangram Labs, I tried my best to be human, but apparently I was born a robot.

I appreciate your mission to detect AI-generated content and “create transparency.” Mine, in a way, is to test tools like this and highlight false claims—because false positives can have very real consequences for completely innocent people.

As you can see in the video below, pretty much anything I write seems to be detected as AI-generated. Not just possibly AI-generated: 100% AI-generated.

I’ve always thought my writing style was far from perfect. I forget words, make mistakes, and occasionally use expressions like “piece of crap”—which I’m not sure would be ChatGPT’s first choice of wording.

But according to the detector, robot.

And here comes the best part.

The message you just read was actually rewritten by ChatGPT.

That one?

100% human, according to Pangram.

In total, I tested 10 pieces of text that I either wrote or got an AI to write. Six were improperly classified. Three human texts were classified as AI, and three AI-generated texts were classified as human.

I agree that this is a small number of samples, but I don’t think I would need to test 10,000 times to see the final results.

This was enough for me to see what I would call blatant mistakes.

I gave them the chance to answer

I did not keep the results to myself. I posted them publicly and tagged Pangram Labs. No reply. I then cancelled my paid subscription and, when they asked why I was leaving, told them I had found too many mistakes and would be happy to talk about it. No reply to that either.

A company is not obliged to answer everyone who complains on the internet, so I am not making too much of the silence. But the offer stands. I still do not know why an ordinary user runs into this many errors when the published numbers say it should almost never happen, and I would rather have an explanation than a theory.

Why AI detection is statistical, not proof

Unless a text has a SynthID or a digital watermark (soon to come), AI detectors are based purely on statistical models. They are neither fact nor proof. I don’t expect them to be perfect, but three things in particular bother me about Pangram.

1. The discrepancy between published accuracy and my experience

Pangram advertises a false-positive rate of roughly 1 in 10,000 human documents, but in my small test I encountered three false positives, alongside three AI-generated texts incorrectly classified as human.

I genuinely don’t know why and would be interested to know how Pangram works, as it claims to be different from other AI detectors.

But if a normal user can encounter multiple false positives this quickly, I think that deserves investigation rather than dismissal simply because an aggregate benchmark reports an extremely low error rate.

2. A prediction should look like a prediction

AI detection is probabilistic, not definitive proof. A prediction about authorship should always be labeled as a prediction, not a fact.

I repeatedly received results saying either 100% human or 100% AI-generated. The interface explicitly states: “100% of this text is AI-Generated.” That is a bold statement with not an ounce of doubt. Generative AI is frequently accompanied by disclaimers like “AI can make mistakes,” but AI detectors, who often use AI themselves, lack this humility. There is a massive ethical difference between stating “Our classifier strongly predicts that this text is AI-generated” and declaring “This text is AI-generated.”

3. Browser extensions amplify the risk

This is the part that moved the issue beyond an amusing experiment. Pangram offers a browser extension that automatically scans posts on platforms including LinkedIn and X. It places a small AI bot face directly next to someone’s post. Clicking it opens a panel explaining that Pangram believes the text was generated by AI.

This means someone else can browse LinkedIn, see something I wrote entirely from scratch, and have a third-party tool attach an incorrect AI classification to my professional work.

Why false AI accusations matter

Imagine being a student accused of cheating on an essay you spent three weeks writing. Imagine an author being accused of generating a book they spent a lot of time creating. Imagine a job applicant having their expertise questioned because a detector placed an AI label next to their portfolio.

Proving that you didn’t use an AI is surprisingly difficult. The accusation takes software mere seconds to produce, while disproving it takes the human considerably more effort.

AI detection can be useful, but AI detection is not proof. If these tools are going to exist—and clearly they are—companies must make them as transparent as possible. Better detection must come with better communication about uncertainty. Especially when the result can be shown to third parties browsing someone’s professional content.

Trending