Somewhere in your pension documentation, there is a sentence that reads: past performance is not indicative of future outcomes.

You have read it. You agreed to it. You continued anyway, because well sure the future is uncertain, what else is new? These days, it’s the flavour of snake oil that’s new.

That warning sentence is not there because someone thought it would change your behaviour. I checked and Gemini thought I was talking about musical performances, because statistically ‘historic performance’ is linked more to music than pensions…but more on that later. The sentence arrived in the small print because US regulators demanded it, in the 1930-40s, gee kids I wonder why. It got adopted internationally because the risk of suit made it necessary. Somewhere along the line a lawyer earned their stripes and protected their employer or client from suit before international law changed, the next institution’s response was disclosure. Case law accumulates. The disclaimers multiply. The happy pace of progress plodded along towards individual accountability and ignoring evolutionary reasons why we’re bad at this as a species.

You’ll pick a fund because the company brand has historically appealed to you. You’ll tell yourself that is meaningfully different from the fund’s historic performance. Remember this: they’ll get paid either way. The people most insulated from the risk of a bad decision are also the most motivated to price in the cost of being wrong before they have to be. There’s a happy little note on all the chat logs, lined up right on top of your chat input, prime digital real estate: “Ai can make mistakes”. “So can people”, you cry, but CEOs negotiate their exit packages before they need them, the employees they replaced with AI don’t. What happens after the golden parachute maroons your company brand in the consequences of “ai can make mistakes”?

We have built entire regulatory frameworks on the premise that disclosure is the same as comprehension: it isn’t. We’re deconstructing entire regulatory frameworks on the basis that it attacks free speech: it doesn’t. Retaliatory sanctions against EU policy architects punish foreign markets as if they are colonial US assets. What happens in a world where you are obliged to trust wildly complex and complexifying systems you weren’t trained to interrogate? Enron, Theranos, Nikola Corporation, Wirecard, and the innumerable purveyors of snake oil sell tech debt as if it’s gold. Impressive people on boards are not the same as people with the specific expertise to see through a specific deception.

Is your Trust and Safety function subprime? When the complexity of your safety infrastructure exceeds the comprehension of the people accountable for it, you don’t need bad intentions to produce bad outcomes.

In 2008, the mechanism that collapsed the global financial system was not, at root, greed. Greed was ambient; it always is. The mechanism was more specific: the systematic outsourcing of judgment to ratings that nobody sufficiently understood.

Collateralised debt obligations were genuinely complex. Layers of mortgages, packaged and repackaged, sliced into tranches with different risk profiles, assessed by ratings agencies whose models were themselves poorly understood by the people relying on them. When the agencies rated tranches AAA, the same rating as US government debt, portfolio managers accepted those ratings and moved on. The prospectus ran to hundreds of pages. Nobody read it.

The AAA was the summary. The summary was the decision.

What made it catastrophic was not that the ratings were wrong, though they were. It was that by the time the wrongness became visible, the system had become too interconnected to absorb the correction cleanly. The quants who built the models understood the fragility they had created; some of them said so, too quietly, and to audiences who did not have the tools to evaluate what they were hearing, or the incentive to slow down. Nobody who heard the warning valued the expertise behind it enough to act on it.

The elevator pitch on “this is going to implode” did not land. That is not how it sounds from thirty-five thousand feet. From thirty-five thousand feet, you see blue sky and a strong quarter. You see your bonus not your exit package. You trust a pipeline to surface substantial issues in a culture where middle management punishes people “who bring problems without solutions”. The people with the most granular visibility into fragility learn, fairly quickly, not to surface it.

Snake oil for $5 a bottle and the farmers market doesn’t sell. But from a keynote stage you can sell a place on the waitlist for $2000 per micro dose to be delivered 6 years from now when the tech catches up with your promises. It will catch up because you’re a better visionary than your peers. The phrase “artificial intelligence” is akin to this, even now we’re learning to distinguish it from AGI, ‘the real deal’ the thinks for itself real deal. It’s the snake oil of its day because it does two pieces of misdirection simultaneously.

It is not artificial. Every large language model is trained on human-generated text: our books, our arguments, our journalism, our court records, our creative work, our Wikipedia edits, our Reddit threads. The intelligence, such as it is, is a guess, a good one, but still a guess. At best therefore current LLM’s should be understood as “regurgitating vacuums”. Early AI adoption understands that one day it will be better than that, and probably soon, but if your entry level employees were to fail as often your AI does on business tasks they won’t last their probation period.

It is not intelligent. It is a prediction engine. When you ask a large language model a question, it does not look up the answer. It generates the statistically most likely next sequence of tokens given what you asked and everything it was trained on. This distinction matters more in practice than most people deploying these systems have absorbed.

Ask an LLM to write you a strawberry recipe, and it will produce something excellent, because enormous quantities of cooking text in the training data mean the probable outputs cluster convincingly around real culinary knowledge. Ask the same model how many letter R’s appear in the word “strawberry” and, until recently, most models would confidently tell you two. The correct answer is three. The model is not lying. It is not making a mistake in the way a human makes a mistake. It is doing exactly what it was designed to do: predict a likely answer. It just cannot count letters, because it does not process text as letters; it processes it as chunks, as tokens, and “strawberry” is not three R’s in a sequence, it is a token that probably follows “fresh” and precedes “jam.” (For more on this: Why AI can’t spell ‘strawberry’, TechCrunch, 2024)

Most people deploying these systems have not fully absorbed this. A 2025 study published in Nature Machine Intelligence found that users’ ability to assess whether an LLM response was correct performed barely above random chance, even when given explanations. They were, in the researchers’ framing, near-random at discriminating between correct and incorrect outputs. The same research found that models are 34% more likely to use language like “certainly” and “definitely” when generating incorrect information.

The more wrong the system is, the more confident it sounds.

If your AI deployment is currently human assisted, your feedback mechanisms have supervision and for the most part the feedback loop is likely immediate. If your bot has backend developer power and it decides that the coding environment is subpar and deletes it you have what AWS had. If your bot has front end engagement with a customer you risk customer effort increase, brand damage and dissatisfaction or potential loss on deal making and decisions in the best case. But it can get darker than that, as the parents of a boy who took his own life after support from a chat bot encouraged him.

But what happens when social media platforms deploy AI moderation systems? What happens when jobs with low probability of surviving AI can’t get human employees anymore because people don’t want to work emergency phone lines? What happens when calls are screened by AI services? Increasingly chained sequences of AI systems each feeding the next, and accepting their outputs with the same cognitive posture that portfolio managers brought to CDO ratings. The model card is the prospectus. The benchmark score is the AAA. The vendor presentation is the ratings agency. The complexity of these systems has begun to exceed the comprehension of the people accountable for their outcomes, and the infrastructure that would allow comprehension, meaning deep technical expertise, adversarial testing, and human oversight with genuine judgment, is being optimised away because it is expensive and slow and its value is difficult to put on a dashboard.

In Trust and Safety, the most consequential cases are not the routine ones. Most people do not hurt themselves with cars; the car has to be going at a specific speed, in a specific direction, toward a specific person, in a moment that was not predicted. Most people do not hurt themselves on livestreams. And then someone does, and the question is whether the system that was supposed to catch it was calibrated for what actually happened, or for what typically happens.

The pension disclaimer applies here with more force than it does to actual pensions. Past performance, meaning the training data, does not indicate future outcomes, because the environment in which those outcomes will occur is not the environment the model was trained in.

There is a compounding problem specific to platforms: success changes the environment. The product you build influences the world around it. The scale you reach creates conditions that did not exist before you reached it. History cannot predict your future performance because you are the leading edge. Nobody trained a model on the harm patterns of a platform that did not yet exist. The most dangerous moment for a Trust and Safety system is not at launch, when everyone is watching, but at the point where the platform has become significant enough to attract sophisticated adversaries, and the model was trained before those adversaries existed.

This gap does not stay the same size. It widens.

With traditional software, failure modes were findable. Edge cases were enumerable, in principle. You could characterise the system over time. AI systems don’t work like that; they respond differently to adversarial inputs than to ordinary ones. Prompt injection attacks, where malicious instructions are embedded in input data to hijack an AI system’s behaviour, are not a theoretical concern; they are an active and evolving attack surface, and the platforms most exposed are the ones where the AI is trusted to make consequential decisions with limited human review.

The biases AI systems inherit from training data do not always announce themselves. A study on AI bias in UK council services found that systems trained on historical administrative data reproduced and compounded existing inequalities in access to healthcare, invisibly and at scale, in ways that only became apparent through specific investigation. The system was not malfunctioning. It was functioning exactly as designed. It had just been designed on data that encoded decades of structural inequality, and nobody asked what that meant for the outputs before deployment.

And then there is the question of what, exactly, we trained these models on.

A substantial portion of the training data for many large language models includes Reddit. This is public information, and it should give any Trust and Safety professional a moment of pause. Reddit is not a curated library. It is an enormous, largely unmoderated accumulation of human expression across every conceivable subject, including the parts of human expression that Trust and Safety teams exist to manage. The irony is complete: the systems being deployed to moderate online harm were, in part, trained on online harm. Whether the humans who spent years moderating that content were ever consulted on what an AI would do with it is a question the industry has not seriously asked, and has not seriously answered.

There is one reliable corrective. It is expensive, slow, and does not show up on a cost-benefit analysis: human expertise. Specifically, the expertise to look at a system’s outputs and know what wrong looks like when the system itself does not flag it.

That is not the same as technical knowledge of how the model works. It is domain knowledge built over years of working on the hardest cases, not the high-volume routine ones; the kind that lets someone look at a decision and say: this is statistically likely to be correct, and it is wrong, and here is why.

This expertise is being optimised away. Not because platforms are malicious, but because it is genuinely difficult to quantify the value of expertise that expresses itself primarily as the prevention of things that do not happen. The absence of harm does not appear on a dashboard.

What gets measured gets managed; what cannot be measured tends to get managed out.

The quants who built the CDO models understood the fragility. Some of them said so. The system had no good mechanism for routing that understanding to the people who needed to act on it, and no incentive to slow down in the meantime.

We are building the same architecture in a different domain. And the people with the institutional knowledge to see what the model is missing are the ones most at risk of being replaced by the thing they can see the limits of.

The rating on your AI and Human Trust and Safety Operations integrations is whatever your vendor told you in the sales deck, right? The benchmark is the model on its best behaviour, evaluated against the data it was trained adjacent to, in conditions it was designed for.

What rating does it have on edge cases? Are you testing for that? Does your business insurance cover you for that? If past performance does not predict future outcomes, and if your platform is itself the leading edge of an environment that did not previously exist, at what point does “we reviewed the model card” stop being a defensible answer?

The question is not whether you trust your system.

The question is whether you are planning to fire the people who know why you shouldn’t.