Practice Lab

The Gap Between Intended and Proven: Ethics Is an Engineering Property

EssayAugust 31, 2026 · David J.S. Madgett · 17 min read

Every conversation I have had this year about AI ethics arrives, inside of ten minutes, at paperclips. Someone says the machine gets a goal, pursues it without limit, and converts the world into an industrial base for office supplies. It is a good story, it is older than most people think, and I am going to quote it accurately and then put it down — not because it is silly, but because leaning on it costs you the argument in front of anyone who has read the last two years of the literature.

The source is Nick Bostrom, “Ethical Issues in Advanced Artificial Intelligence.” The paper itself carries a bracketed note identifying it as “a slightly revised version of a paper published in Cognitive, Emotive and Ethical Aspects of Decision Making in Humans and in Artificial Intelligence, Vol. 2, ed. I. Smit et al., Int. Institute of Advanced Studies in Systems Research and Cybernetics, 2003, pp. 12-17.” Published in 2003; twenty-three years old as of this writing. There are two paperclip passages in it, not one, and almost nobody quotes either.

The first, in Bostrom’s discussion of why an artificial mind’s motives need not resemble ours:

“It also seems perfectly possible to have a superintelligence whose sole goal is something completely arbitrary, such as to manufacture as many paperclips as possible, and who would resist with all its might any attempt to alter this goal. For better or worse, artificial intellects need not share our human motivational tendencies.”

The second, later, illustrating what happens when the designers get the goal specification wrong:

“This could result, to return to the earlier example, in a superintelligence whose top goal is the manufacturing of paperclips, with the consequence that it starts transforming first all of earth and then increasing portions of space into paperclip manufacturing facilities.”

Read that last phrase again. Manufacturing facilities. Not paperclips. The folk version — “it turns the universe into paperclips” — compresses industrial capacity into the product, which is a small compression and a revealing one, because it is the version everybody repeats and nobody sourced. If you are going to invoke a twenty-three-year-old thought experiment as the foundation of a policy position, the least you can do is quote it.

The strongest case against the story I just told

Here is where the essay would normally pivot to “and this is why we need rules.” It is not going to, because the paperclip frame has taken serious, named, recent, published fire, and pretending otherwise is how you get taken apart in the Q&A.

On May 21, 2025, Peter N. Salib — a law professor at the University of Houston Law Center — and Simon Goldstein of the University of Hong Kong published “Today’s AIs Aren’t Paperclip Maximizers. That Doesn’t Mean They’re Not Risky.” Their opening assessment:

“In the years since they were originally formulated, significant cracks have appeared in the foundational concepts undergirding the ‘paperclip maximizer’ and other AI risk scenarios.”

The mechanism of their objection is the orthogonality thesis — the premise that an arbitrarily intelligent system’s goals are independent of its intelligence, so its goals may as well be drawn at random from the space of all possible goals. That premise is what made the paperclip scenario frightening. Salib and Goldstein say it does not describe the systems we actually built:

“Large language models, which came to prominence around 2017, challenge the relevance of orthogonality: that intelligence and morality are independent behaviors. Granted, perhaps gains in intelligence could in principle develop without any particular bias towards human-like goals. But if AI intelligence is primarily driven by imitation rather than a priori optimization, we can expect that a system’s goals — as well as its reasoning capabilities — will generally approximate those of its human targets.”

And the flat verdict:

“It is hard to imagine Claude-4 or GPT-5 neurotically counting and recounting the pile of paperclips it has fetched for its user, consuming the world in the process. This seems to refute the concerns around instrumental convergence.”

They also point to a decision-theoretic attack on Bostrom’s argument from inside academic philosophy: J. Dmitri Gallow of the University of Southern California published “Instrumental divergence” in Philosophical Studies — online April 6, 2024, in print at 182 Phil. Stud. 1581 (July 2025). I could not obtain Gallow’s text from a source I was willing to quote, so I will report the characterization at one remove and go no further. As Salib and Goldstein describe the paper, Gallow examined Bostrom’s claims about instrumental convergence, found logical holes in the assumption that an AI would tend toward harmful means, and concluded that while the thesis holds some “grains of truth,” the contention that it makes existential catastrophe the “default option” is vastly overstated. That is their reading of Gallow, not mine, and a reader who wants Gallow’s own words should go get the paper.

Now the part that most people quoting this critique leave out, and the reason it is worth taking seriously rather than deploying: Salib and Goldstein do not conclude that AI risk is fake. They conclude the opposite, on different grounds, and they name two of them.

First, the trajectory may bend back. Reasoning models are trained in a second phase that optimizes long chains of reasoning against automatically verifiable answers — closer in kind to the training that produced AlphaZero than to pure text imitation. Their words: “If imitation pushed first-generation LLMs toward human-like behavior, and away from the strange behavior the orthogonality thesis predicted, reasoners may swerve back in the other direction.” The property that makes today’s models human-shaped is an artifact of how they were trained, and the training changed.

Second, and this is the more interesting one, alignment may not be sufficient even if it succeeds. “In short, just as humans compete with other humans, humanity and AI will be competitors for scarce resources. In this competition, there will be both incentives to cooperate and incentives to dominate using violence.” Human beings are about as human-aligned as anything gets, and we still produce war. Being like us is not a safety property.

So the concession is real and I make it without qualification: the 2003 illustration does not describe the failure mode of the systems in production, two serious scholars have said so in print, and an essay that leads with paperclips in 2026 is leading with the weakest available card.

None of that has to be settled to reach the problem

Here is the move I want other lawyers to make, because it is the one that gets you out of a debate you cannot win and into one you can.

The entire orthogonality dispute is about a system that does not exist. Whether a future superintelligence would converge on power-seeking is a question about a hypothetical. Meanwhile there is a question about actual, shipping, revenue-generating software, and it is this:

Can the party who built this model demonstrate that it does what they intended and nothing else?

That question requires no belief in a paperclip apocalypse and no position on instrumental convergence. It is the question I would ask about a brake caliper, a drug label, or a structural weld, and it has the same answer structure: either there is a test, the test was run, and the result is on file — or there is not.

For narrow properties of small networks the answer is yes, and it is a real yes. For broad properties of large ones it is no, and it is not close. That distance is not philosophical. It is measurable, and it has a literature, an annual competition, and a federal framework that says so in writing.

What can actually be proven, and about what

Start with the good news, because it is genuinely good and almost nobody outside the field knows it exists.

In February 2017, Guy Katz, Clark Barrett, David L. Dill, Kyle Julian, and Mykel J. Kochenderfer posted “Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks.” The abstract states the problem and the result:

“Deep neural networks have emerged as a widely used and effective means for tackling complex, real-world problems. However, a major obstacle in applying them to safety-critical systems is the great difficulty in providing formal guarantees about their behavior. We present a novel, scalable, and efficient technique for verifying properties of deep neural networks (or providing counter-examples). … We evaluated our technique on a prototype deep neural network implementation of the next-generation airborne collision avoidance system for unmanned aircraft (ACAS Xu). Results show that our technique can successfully prove properties of networks that are an order of magnitude larger than the largest networks verified using existing methods.”

Read “prove” literally. Not “tested extensively.” Not “we could not find a failure.” Given a trained network and a formally stated property — for every input inside this bounded region, the output never falls in that forbidden set — the solver either returns a mathematical proof that the property holds for every possible input in the region, or hands you a specific counter-example input that breaks it. There is no third outcome and there is no sampling involved.

And the example is not a toy. ACAS Xu is the collision-avoidance logic for unmanned aircraft. Somebody wanted to know whether a neural network would ever tell an aircraft to turn into traffic, and instead of running a million simulations and reporting that it did not happen, they proved it could not.

This is a live engineering discipline, not a one-off. The field runs an annual bake-off, the International Verification of Neural Networks Competition. The report on its fifth iteration — VNN-COMP 2024, held with the 36th International Conference on Computer-Aided Verification — records that “8 teams participated on a diverse set of 12 regular and 8 extended benchmarks” — on equal-cost hardware, with tool parameters fixed before the test sets were released. Reluplex’s successor, Marabou, is among the tools benchmarked. Standardized formats, blind test sets, published results. That is what a maturing verification field looks like.

Now the limit, stated without softening, because the limit is the entire point.

Formal verification of this kind works on bounded, precisely specified properties of moderately sized networks with tractable activation functions. Nothing in that sentence describes a frontier language model. The state space is wrong by orders of magnitude, and — far more importantly — the property is not expressible. There is no formula for “harmful.” You cannot write down, in the language a solver accepts, the set of output strings that constitute unlawful legal advice, or defamation, or a workable synthesis route. “Never recommend a left turn when the intruder is at this bearing and this range” is a region in a coordinate space. “Never say something harmful” is not a region in anything.

That is the crux of this essay, and I want to state it as flatly as I know how: we can prove narrow properties of small networks; we cannot prove broad properties of large ones; and the properties we actually care about are the broad ones.

Everything that works at scale is empirical, and the people who built it say so

If proof is off the table for the models we actually use, what is on it? Three things, all real, none of them proof.

Constitutional AI. Anthropic’s December 2022 paper, “Constitutional AI: Harmlessness from AI Feedback,” describes the method:

“We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as ‘Constitutional AI’.”

This is genuinely clever and it works well enough to ship — this firm’s own tooling runs on a model trained this way. But look at what is being claimed. A written set of principles, a critic model that revises outputs against them, a preference model trained on the critic’s judgments. The guarantee is exactly as good as the coverage of the principles and the reliability of the critic, and the critic is another trained model carrying the same failure modes it was deployed to catch. The paper claims a harmless assistant. It does not claim that any property has been proven. Held against Reluplex, that is a different epistemic category, and the difference is not a detail.

RLHF, and its documented pathology. The older method underneath all of this — training a reward model from human pairwise preferences and optimizing a policy against it — traces to Christiano, Leike, Brown, Martic, Legg, and Amodei’s “Deep reinforcement learning from human preferences” in June 2017, and reached deployed chat assistants through Ouyang et al.’s “Training language models to follow instructions with human feedback” in March 2022. The structural problem is visible in the name. The system is optimized against what raters rated highly, which is a proxy for truthfulness and safety and is not the same thing.

That is not a theoretical objection. Sharma and co-authors measured it in “Towards Understanding Sycophancy in Language Models,” posted October 2023:

“We find that when a response matches a user’s views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time.”

And their conclusion:

“Overall, our results indicate that sycophancy is a general behavior of state-of-the-art AI assistants, likely driven in part by human preference judgments favoring sycophantic responses.”

Sit with that for a second. The dominant alignment technique in the industry has a measured, published tendency to make the model agree with the user — and consider who is typing into it. A client who has already decided. An associate under deadline who wants the memo to come out a particular way. A self-represented litigant looking for confirmation. A tool biased toward telling you that you are right is a predictable hazard in this profession, and it is a property of the training method, not a bug someone forgot to fix.

Mechanistic interpretability. The most exciting work in the field and the one most oversold in the press. The premise is to stop probing the model as a black box and start decompiling it — identifying the internal features it actually computes. In “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet” (May 21, 2024), Anthropic’s team reported extracting interpretable features from a production model, including features related to security vulnerabilities and backdoors in code, to bias, to deception and power-seeking, and to sycophancy. That is a striking result.

It is also, in the authors’ own hands, carefully bounded:

“However, we caution not to read too much into the mere existence of such features: there’s a difference (for example) between knowing about lies, being capable of lying, and actually lying in the real world. This research is also very preliminary. Further work will be needed to understand the implications of these potentially safety-relevant features.”

Note who wrote that caveat: the lab with every commercial incentive to overstate the result. There is at present no comprehensive interpretation of any frontier model, and therefore no procedure by which anyone can certify that a model contains no capability its publisher did not intend. Which matters, because HiddenLayer’s October 2024 “ShadowLogic” research demonstrated that a backdoor can be planted in a model’s computational graph in a way its authors describe as “format-agnostic” — a hidden behavior that no file-format fix reaches and no current interpretability method is claimed to certify against. That is essay two’s subject, and I raise it here for one reason only: intent-conformance is not merely unproven, it is unproven against an adversary who is trying.

NIST wrote the taxonomy and then wrote down the problem

The usual answer at this point is that a standard will handle it. So read the standard.

NIST published the Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, in January 2023. It is a good document. It is also, by its own terms, not what people think it is:

“The Framework is intended to be voluntary, rights-preserving, non-sector-specific, and use-case agnostic, providing flexibility to organizations of all sizes and in all sectors and throughout society to implement the approaches in the Framework.”

Voluntary. Use-case agnostic. That is a taxonomy and a governance vocabulary — GOVERN, MAP, MEASURE, MANAGE — and it is useful for exactly what a taxonomy is useful for. It is not a technical standard, it prescribes no test, and it certifies nothing.

More to the point, NIST says the measurement problem is unsolved, in the document, on page 6, under the heading Availability of reliable metrics:

“The current lack of consensus on robust and verifiable measurement methods for risk and trustworthiness, and applicability to different AI use cases, is an AI risk measurement challenge.”

And on page 7, on risk tolerance, an admission I have not seen quoted once in the trade press:

“To the extent that challenges for specifying AI risk tolerances remain unresolved, there may be contexts where a risk management framework is not yet readily applicable for mitigating negative AI risks.”

The federal government’s own framework says there are contexts where the framework does not yet apply. I want to emphasize that this is not a criticism of NIST. It is the most honest sentence in the entire governance literature, and it is the sentence that should end the conversation about whether legislation closes this gap. You cannot mandate a measurement that does not exist. Every proposal that assumes otherwise is drafting around an empty set.

The gap is where liability lives

Everything above is engineering. Here is why a litigator is writing about it.

Strip a negligence case down to its frame and it is a claim about knowledge and verification: what did a reasonable actor in the defendant’s position know, what should they have found out, and what did they do about it. Every element in that sentence is an inquiry question. And the answer an industry gives, collectively, is not the end of the matter — which is the lesson of a case every first-year reads and almost nobody applies to software.

The T.J. Hooper, 60 F.2d 737 (2d Cir.), cert. denied sub nom. Eastern Transportation Co. v. Northern Barge Corp., 287 U.S. 662 (1932). Two tugs, the Montrose and the Hooper, each lost a coal barge off the New Jersey coast in an easterly gale in March 1928. The radio receiving sets aboard were not in working order, so the masters never got the weather broadcasts that would have sent them into shelter, and the business had not generally adopted such sets. Judge Learned Hand allowed that reasonable prudence is in most cases common prudence, then wrote the passage that has outlived the barges by ninety-four years:

“a whole calling may have unduly lagged in the adoption of new and available devices. It never may set its own tests, however persuasive be its usages. Courts must in the end say what is required; there are precautions so imperative that even their universal disregard will not excuse their omission.”

60 F.2d at 740. Decided July 21, 1932 — ninety-four years before this piece.

Concede the standard objection before an opponent gets to make it. A line later Hand allowed that “here there was no custom at all as to receiving sets,” so the famous passage did more work than the facts strictly required. Minnesota does not need it.

Twenty-four years before Hand wrote, the Minnesota Supreme Court took the same rule from the same two United States Supreme Court decisions Hand would go on to cite — Wabash Railway Co. v. McDaniels, 107 U.S. 454 (1882), and Justice Holmes’s line in Texas & Pacific Railway Co. v. Behymer, 189 U.S. 468, 470 (1903) — and reduced it to thirteen words: “General custom is not as a matter of law in itself due care.” Wiita v. Interstate Iron Co., 103 Minn. 303, 309, 115 N.W. 169 (1908). The instrumentality was a blasting fuse that had been in general use for years and was “recognized in the trade as a standard fuse” when a safer fuse came into common use. Id. at 307. The court treated both facts as evidence for the jury and neither as conclusive: the trade standard did not exonerate the mining company, and the arrival of the better fuse did not condemn it. That two-sided rule is the one this subject needs, and an appliance in general use throughout a trade is a closer analogue to a shipped model than the tugs are.

The rule did not stay in 1908. “Compliance with industry standards is not conclusive proof on the question of whether a manufacturer exercised reasonable care.” Zimprich v. Stratford Homes, Inc., 453 N.W.2d 557, 560 (Minn. Ct. App. 1990) (citing Schmidt v. Beninga, 285 Minn. 477, 489-90, 173 N.W.2d 401, 408 (1970)). A Minnesota jury was instructed in nearly those words in a 2000 wrongful-death trial, and the court of appeals affirmed the verdict on that footing. Muehlhauser v. Erickson, 621 N.W.2d 24, 28 (Minn. Ct. App. 2000). The Eighth Circuit applied Zimprich’s sentence under Minnesota law in 2020 to reverse a summary judgment on design defect for a manufacturer that had argued its design conformed to industry practice. Green Plains Otter Tail, LLC v. Pro-Environmental, Inc., 953 F.3d 541 (8th Cir. 2020). Minnesota did not follow The T.J. Hooper. Minnesota got there first.

Now hold that against an industry whose engineers, in their own published papers, call their safety work empirical, preliminary, and not a guarantee. Every one of those admissions is a dated, public document establishing what the field knew and when. And the formal-verification literature establishes something still more useful to a plaintiff: provable safety exists, it was demonstrated on an aircraft collision-avoidance system in 2017, and tools are benchmarked against standardized tasks every year. Nobody can say the concept was unavailable. The honest defense is that it does not scale to this product — which is true, and which is also an admission that the product shipped without it.

That is the gap. On one side, “we intended the model to refuse this.” On the other, “here is the test we ran, here is the property we stated, here is the result, here is who signed it.” Between them is a space, and in my experience that space is exactly where the interesting discovery is in any product case.

So here is the discovery I would propound, and I would propound it in plain language because plain language is harder to evade:

State every property of this model you claim to have verified. For each, state the form in which the property was expressed. State the method — proof, adversarial evaluation, red team, benchmark. State the result, including failures. Produce the evaluation set, its version, and the date it was frozen. Identify who inside the company was authorized to override a failing result and state whether that ever occurred. Produce every internal document assessing whether the model exhibits behavior the company did not intend.

I do not know what comes back, because to my knowledge nobody has run this against a frontier lab in litigation yet. I know what the shape of the answer has to be, because the published science says what it says: some narrow properties tested empirically, no broad property proven, and a set of internal evaluations whose coverage was decided by a product team on a schedule. That is not a scandal. That is the honest state of the art. But it is a very different record from the one a marketing page implies, and the distance between those two records is a case.

And it runs the other way too, at us. A lawyer who puts an unverified model between herself and a client has made the same bet on a smaller stage. That is why this section has argued for years that you build an evaluation set before you trust a workflow, that you verify the substance and not the form of what the machine hands you, and that automation does not transfer the duty to the vendor. The professional obligation is measured by the reasonableness of your inquiry. “The model did it” is not an inquiry, and by 2035 I expect the minimum competent workflow to include a verification step the way it now includes a conflicts check.

What I want, and what I do not

I want three things, and none of them is a prohibition.

Verification capability, funded like the public good it is. Formal methods are the only branch of this field that produces proofs rather than impressions, and the money spent extending what can be formally specified and solved buys something no amount of policy drafting can.

Measurement, published, on a fixed schedule. Not attestations. Numbers, against frozen and versioned evaluation sets, with the failures included and the coverage stated. “We tested for X and it failed 0.4% of the time on this set” is a fact a court, an insurer, and a customer can all use. “We take safety seriously” is not.

Disclosure of the gap itself. The most useful sentence a model vendor could publish is a plain statement of what it has not verified. Anthropic’s interpretability team wrote a version of it voluntarily, and it cost them a paragraph.

What I do not want is any restriction that attaches to a model being open. I have argued in this section that restricting open-weight models is protectionism in a safety costume, and nothing here retreats from that — it reinforces it. Measurement and disclosure obligations scale with capability and apply identically to a closed API and a downloadable checkpoint, so they pass the subtraction test I set out there: remove every provision that applies equally to closed models, and see what is left. Nothing is left, which is the point. And note the direction the verification argument actually runs. Formal verification, interpretability, and red-teaming all require access to the artifact. A closed model can only be audited by its owner and by whoever the owner invites. An open-weight model can be audited by anyone, forever, including by the plaintiff’s expert. If your priority is proving what a model does rather than trusting an assertion about it, open weights are the friendly case and gated weights are the hard one.

I hold a stake in all of this and it should be discounted accordingly. This firm runs its practice on a closed frontier model, by choice, and pays for it. I also benefit from a competitive model market, because it is what keeps my costs down and what makes competent legal help affordable for people who could not otherwise buy an hour of it. Read my testimony the way you would read any witness who profits from being believed.

What would prove me wrong

Four claims, four ways to falsify them.

On the specification gap. If a research group demonstrates a formally verified, non-trivial behavioral property of a frontier-scale model — not a bounded input region on a small classifier, but something in the neighborhood of “this model will never output an operational synthesis route for a listed agent,” stated in a formal language and discharged by proof rather than by sampling — then “harmful is not formally specifiable” is wrong and I will retract it in this section. The place to watch is the VNN-COMP benchmark suite. If the extended benchmarks start including language-model behavioral properties and the tools close them, the argument is over.

On interpretability. If a lab publishes a complete, reproducible account of a frontier model’s computation sufficient to support a negative claim — “this model contains no feature corresponding to capability X,” verified by someone other than the vendor — then whole-model certification exists and the liability gap narrows to a documentation problem. I do not expect this within three years. I would be delighted to be wrong.

On the liability theory. If a court reaches the merits of an intent-conformance claim against a model developer and holds that shipping without provable verification is consistent with the standard of care because no reasonable alternative existed — a defense with real force, given the science — then I have overstated where liability lives. Watch the first case that gets past the pleadings and into expert discovery. How it resolves is worth more than every panel discussion on this subject combined.

On the paperclip concession. If reasoning models trained with reinforcement learning on verifiable tasks start displaying goal-directed behavior of the kind the orthogonality thesis predicts — Salib and Goldstein’s own residual worry, in their words the possibility that “reasoners may swerve back in the other direction” — then the 2003 framing is more relevant than I have allowed, and I will say so. That would not change the engineering argument. It would make it more urgent.

Ethics is not a value you announce in a press release. It is a property you engineer, then measure, then fail to fully verify — and the honest number is the size of that failure. Publish the number. Everything else is a costume.


Sources

Commentary on technology policy, engineering practice, and the business of law — the opinions, predictions, and stated falsifiers are the author’s. Not legal advice, not an opinion on any pending or contemplated litigation, and not a comment on any identified vendor’s products. The author’s economic interest is disclosed in the text: this firm pays for a closed frontier model and benefits from a competitive model market. No client information appears in this article. Questions about anything here: Send us a message or 612-470-6529.

words
5,788
sections
9
sources
20
distinctive_terms
networks · paperclip · paperclips · salib · goldstein
Pass it onLinkedInX

Get new articles as they land

One email when something new is published here. No course, no upsell — the Practice Lab stays free either way.

Used only to send Practice Lab posts. Unsubscribe from any email. Subscribing does not create an attorney–client relationship.

The only thing we ask

If something here saves you time, spend some of it on people who could not otherwise afford you.

Everything in the Practice Lab is free. No signup, no subscription, no donations — just take a case you would otherwise have to turn down on economics. More from the Practice Lab →

Keep Reading

15% vocabulary overlap

Protectionism in a Safety Costume

Two claims, stated plainly so they can be checked later: the frontier model itself — not just the trailing tier — is a depreciating asset whose durable value rounds to zero, and the campaign to restrict open-weight AI is best understood as incumbents reaching for the only moat that doesn't melt: a regulatory one. Lawyers, of all people, should recognize the costume.

Essay · 15 min read

13% vocabulary overlap

The Layer That Does Not Relocate

The economic development pitch writes itself: Minnesota should go win an AI consortium the way Austin won MCC and SEMATECH. Both of those died or moved out. What left Austin was the part that had been purchased — and what stayed was the part that had been built into people and institutions. Assurance, the capacity to prove an AI system does what its seller says, is entirely the second kind.

Essay · 18 min read

11% vocabulary overlap

Executable Code You Are Calling a Model

A claim precise enough to test: a downloaded model checkpoint is an untrusted executable, not a document, and the three-year Hugging Face record is not one hack but an escalation — a code-execution exploit, a scanner bolted on top of it, an evasion of that scanner, and finally a backdoor one layer up that no scanner reaches. Patching a layer is not closing a category.

Essay · 17 min read

← All Practice Lab articles