When the Question Is No Longer Just Real or Fake
AI-generated content has moved well past the point of being a nuisance in social media feeds. The problem has spread into job applications, product reviews, and insurance claims – areas where getting the answer wrong carries real consequences for businesses, platforms, and the people depending on them to make fair decisions.
Max Spero, founder of AI detection startup Pangram, has been making the case that the challenge his company is trying to solve is fundamentally different from a simple binary classification problem. Identifying whether something was written or made by an AI, he argues, is harder than it looks from the outside.

Why Detection Is Harder Than a Yes-or-No Question
The framing most people apply to AI detection – real versus fake, human versus machine – undersells how complicated the actual task is. A job applicant might draft something themselves and then run it through an AI tool to polish the language. A product review might start as genuine feedback before being restructured by a chatbot. The result sits in an ambiguous middle ground that a binary detector was never designed to handle cleanly.
Pangram is among a handful of startups that have emerged in the past couple of years specifically to address this problem. The broader field has grown alongside the rapid spread of large language models, as platforms began facing real pressure to validate what their users were actually submitting. Insurance companies reviewing claims, employers screening applications, and e-commerce sites managing review integrity are all dealing with the same underlying question: how do you build a reliable process around content you can no longer take at face value?
Spero’s position is that the detection problem demands more nuance than the technology most people imagine. A system that simply outputs a confidence score – this text is 87 percent AI-generated – may be giving users a false sense of certainty when the inputs themselves are hybrid or heavily edited. The margin for error is not abstract. A wrongly flagged insurance claim can delay payment to someone with a legitimate loss. A wrongly flagged application can discard a qualified candidate.

The Infrastructure Problem Nobody Talks About
AI slop – the term now widely used for low-effort, algorithmically generated content flooding the internet – is the visible surface of a much deeper structural issue. Platforms were built around the assumption that content submitted by users was, in some meaningful sense, produced by those users. That assumption is no longer safe to make.
Rebuilding trust infrastructure from scratch is not something any single company can do quickly. It requires detection tools that are accurate enough to act on, explainable enough to defend in disputes, and calibrated well enough to avoid penalizing legitimate users who happen to use AI tools as writing aids rather than wholesale content generators. Pangram’s pitch sits directly in that gap.
A Market That Formed Out of Necessity
The startup wave in AI detection did not happen because investors decided it was a promising space. It happened because platforms started running into concrete problems they could not solve with existing tools. Job boards were seeing application volumes spike in ways that correlated suspiciously with the release dates of major language models. Review platforms were noticing patterns in submission language that did not match how customers had historically written about products.
The commercial case for detection tools, in other words, built itself. What remains unsettled is whether the technical side can keep pace. Language models are updated constantly, and detection methods that work against one version of a model may perform worse against the next. Spero has acknowledged this dynamic, framing it as an ongoing cat-and-mouse problem rather than something that gets solved once and stays solved.
Insurance claims represent one of the sharper edges of this problem. The financial stakes are high, the fraud incentive is clear, and the consequences of both under-detection and over-detection are significant. Flag too little and fraudulent claims get paid. Flag too much and legitimate policyholders face delays or denials that can be financially devastating. Detection tools operating in that environment need a much higher bar of reliability than something filtering spam comments on a retail site.
Pangram’s approach, as Spero has described it, is built around the idea that detection needs to be contextual rather than purely statistical. The same words arranged the same way might be perfectly acceptable in one setting and suspicious in another. That context-sensitivity is part of what makes the problem genuinely difficult – and part of what separates a serious detection product from a tool that just runs text through a classifier and returns a number.

The trust problem Pangram is trying to address did not arrive with any single AI release. It has been accumulating for years, and the question now is whether detection tools can mature fast enough to give platforms something reliable enough to actually act on – or whether the gap between what detection promises and what it can deliver will itself become the next source of friction.
What happens to an insurance company that denies a claim based on a detection tool that turns out to have been wrong?








