Steelmanning AI Policy with the AI Rosetta Stone

As part of my fellowship at Harvard, I’ve been thinking hard about artificial intelligence (AI). Start with the physics of the problem: AI is what lawyers call an unavoidably unsafe product. No matter how carefully we build it, it is going to cause harm, and some of those harms will be unforeseeable. That is not a bug we can engineer away. It is an intrinsic property of the thing, the same way it is a property of pharmaceuticals, vaccines, knives, and power tools.

We already have plenty of laws and legal mechanisms for dealing with unsafe products and the harms they cause. The problem is what those mechanisms are built to see. A courtroom is very good at looking at one specific case and measuring the actual loss to one specific person. It is a terrible instrument for weighing that loss against the enormous benefits the same product produced for everyone else.

Tort law is a microscope for a problem that needs a scale.

But we have stepped on this rake before. 

In the 1970s and 1980s a number of plaintiffs sued vaccine manufacturers and won huge awards (Reyes v. Wyeth Laboratories & Givens v. Lederle Laboratories). The manufacturers looked at the math and said, “We’re out.”
They were going to stop making vaccines.

Congress recognized this as a public health infrastructure crisis and built a special solution: National Childhood Vaccine Injury Act, aka the Vaccine Court. If you suffer a covered injury, you bypass civil court and petition a specialized federal claims court in an adversarial proceeding against the government. If your condition matches an established injury table—or if you medically prove causation for off-table harm—you receive compensation from a consumer-tax-funded trust fund. That resolved the core dilemma: manufacturers secured statutory protection from ruinous tort liability to keep producing vaccines, while injured individuals gained a dedicated, no-fault avenue for funded relief without having to prove manufacturer negligence.

So I wondered: could we do the same thing for AI?

My original proposal was simple. AI companies that follow best safety practices would get limited liability for harms, adjudicated through an AI Harms Court. Instead of shopping it around to friends who would be polite to me, I ran it through my AI Rosetta Stone debate tool to see how the different camps would attack it.

My original proposal folded like a cheap lawn chair.

The first problem it found was that we don’t actually know how to do AI safety, and “safety” for a toy chatbot is completely different than safety for a medical dosing system. So I revised: the proposal would define a tiered set of safety practices, and a research institute funded by the participants in the program would establish and maintain those practices.

I ran it again and then it went after auditing and funding. Who pays, and who checks? So I revised again: the AI companies fund the program through an excise tax, and independent third-party auditors, not the companies themselves, verify that a firm actually followed the practices before it qualifies for coverage.

Round after round, it kept going. It pointed out that a cryptographic receipt proves a process ran, not that it was the right process. It pointed out that a deployer-only tax lets the developer who made the design choice off the hook. It pointed out that a population-scale telemetry archive is itself a surveillance asset that someone will eventually want to repurpose. Each time, the proposal got harder to knock over. The current version is at the end of this post.

But the specifics are not really the point, even though I find them interesting. The point is that the debate tool did for my thinking what a good adversarial review does for code or a red team does for a deployed system. It showed me the issues I had not seen. It showed me how people who don’t share my priors would look at those issues. And it let me strengthen the argument one step at a time, with each round building on the last, instead of presenting a finished product and hoping nobody found the cracks.

If you work in AI policy, I encourage you to use the tool the same way. Take the idea you are most attached to, put it in the ring, and see what survives. Then let me know what you think.

I recently ran the debate using Claude Fable and it brought up even more interesting points. Check it that debate HERE.

If you really want to roll up your sleeves and go into learning mode, change the mode from “Text” to “Analysis” and watch the details of how the arguments are being made and being grounded in a rich documented taxonomy of Beliefs, Desires, and Intentions.

Here is an example of what that will give you:

The current proposal for an AI Harms Court, in brief

Congress establishes a National AI Injury Compensation Program, modeled on the vaccine program but with one critical change: immunity is earned by measured deployed behavior, not by paperwork.

  • A no-fault compensation board with subpoena power pays injured claimants from a dedicated trust fund, without requiring proof of but-for causation.
  • Tiered safety standards, with tier definitions and metrics written by independent researchers through public rulemaking, never by the firms being tiered.
  • Immunity as a continuing condition. Systems must stay within pre-registered thresholds for error rates, subgroup disparity, and human escalation, verified by continuous telemetry. Breach a threshold and immunity suspends until the fix is verified in production.
  • A split excise levy: a certification charge on developers keyed to model capability, plus an execution charge on deployers, with graduated duties and rebates so the program does not become an incumbent moat.
  • Independent auditors, funded by the levy, employed by the board, barred from certified firms, with multiple audit shops using different methods.
  • High-stakes tiers carry insurance and capitalization floors, pooled liability for correlated failure, a 90-day pre-deployment safety case, and no immunity at all for irreversible harms.
  • A design-defect track, because some defects never trip a behavioral flag.
  • Recourse for individuals: when a system is flagged, every affected person gets notice, a docket number, and the logs.
  • Surveillance safeguards written into the enabling statute: purpose limitation, separate cryptographic doors for auditors and the state, a flat bar on law-enforcement repurposing, and retention limits.
  • Firms that skip the standards get nothing: no fund coverage and full tort exposure.

Leave a Reply

Your email address will not be published. Required fields are marked *