← Back to Dead Reckoning

Published July 23, 2026

Issue 0007: The Altman Test and The Karp Test

Every AI investment should pass two tests: does it become more valuable as models improve, and do you own the asset that compounds?

Issue 0007: The Altman Test and The Karp Test

Issue 0007: The Altman Test and The Karp Test

In 2025, ServiceNow stood up an internal team we called the Enterprise AI team. The company had been working on AI product features for a while, but internal usage was bifurcated: Claude Enterprise accounts on one side, machine learning models for forecasting and decision support on the other. The new team's job was to ideate, test, pilot, and deploy AI for internal use cases. Move fast. Kill what doesn't work.

The proposals fell into two buckets. The first: fully custom, bespoke small language models built from the ground up. Small models trained on narrow, vertical-specific data, available at the last mile wherever the end user already worked. The second: existing models, open source and paid, layered on top of our own data and systems.

Early energy went to the first bucket. We thought we needed our own models. But as the frontier models got better, the second bucket started winning on almost every dimension. The use cases designed to sit on top of our data improved without us doing anything. What I remember most clearly is walking around for a week after a major model release with my jaw on the floor, because the combination of a vast data warehouse and a materially better model kept producing results that looked like a step-function change in capability. The model improved. We didn't have to.

For the past year and a half, that experience gave me one test for every AI proposal that crosses an operator's desk. On July 1, Alex Karp went on CNBC and handed me the second.

This issue is about why you need both, and what the four combinations tell you that neither test tells you alone.

The First Test: Compounding

The first test comes from Sam Altman. Because he is effectively talking his own book, I approached it with skepticism and pressure-tested it against what I observed inside a $10B revenue company for the better part of two years. Stated for operators:

Build in ways that model advances compound your investment. Avoid building in ways that model advances eat it.

It sounds obvious. It is not obvious in practice, because model commoditization is non-linear and feels abstract.

A vendor's demo might be extremely compelling today while the frontier model’s roadmap is opaque. Because the CFO needs a number this quarter, the closer decision wins, almost by default.

A question almost nobody in the room asks is whether the thing being approved will be worth more or less when Fable-6.3 or GPT-7.5 ships. “Proprietary” prompt engineering and custom UI layers are current-state assets that look and perform well today.

Structured proprietary data is a trajectory asset. That distinction is what the Altman Test is about.

The Second Test: Ownership

On July 1, Karp went on CNBC's Squawk Box to announce a partnership between Palantir and Nvidia: a sovereign AI operating system built on Nvidia's open Nemotron models, designed so that government agencies and critical-infrastructure operators run frontier-class AI on their own hardware, over their own data, holding their own model weights.

His language for it was blunt. His customers, he said, want to know they own the means of production. The fear underneath his observation is this: an enterprise paying for tokens while their proprietary data, their alpha, flows to a handful of labs that may learn from it, commoditize it, or eventually compete with them.

Karp is also, by definition, talking his own book. Palantir now sells the sovereignty stack he was describing; the CNBC appearance was a product launch with a philosophy angle. Both of the men these tests are named after profit from one particular answer, which is not a reason to discard either test. It is a reason to state each one in objective terms.

Here’s the Karp Test, stated for operators:

Build in ways that you own the asset that compounds, and the option to move it. Avoid building in ways that rent your differentiation on someone else’s terms.

If the exposure half of that sounds theoretical, consider what happened this spring.

In April, Anthropic launched Claude Design, a product that competes directly with Figma. Figma had built on Anthropic's models. Anthropic's chief product officer sat on Figma's board and resigned three days before the launch. Figma's stock has lost roughly half its value over the past year.

The same thing happened with Cursor.

Cursor was one of Anthropic’s biggest customers. Then Anthropic launched Claude Code, a product that competes with a company that was among its largest customers. That same product chief has said publicly that he expects the biggest AI labs to come to dominate software businesses.

Whatever you conclude about the ethics, the pattern is documented: the labs can watch where value accrues on top of their models, and then move in.

Two Tests, Four Squares

Here is the part most of the commentary misses: the tests are independent axes. Passing one tells you nothing about the other.

In land navigation you need two numbers to know where you are, longitude and latitude, and either one alone can put you confidently in the wrong grid square.

These tests work the same way. Every AI proposal that reaches your desk is standing in one of four squares.

1. Rented decay. Fails both tests. The vertical wrapper: a credible vendor, a frontier model underneath, proprietary prompt engineering, a branded UI. The capability decays as the base models absorb the vertical layer, and the intelligence underneath is commodity rental the entire time. Your competitor can rent the identical capability next quarter, and the model provider saw every workflow you ran through it.

Example: Jasper hit a $1.5 billion valuation in 2022; ChatGPT shipped that fall, and within a year Jasper had cut its ARR forecast by at least 30 percent, run layoffs, and marked down its own equity (Maginative). The vendor pitching this square today is a Jasper that has not had its November yet.

2. Compounding exposure. Passes the Altman test, fails the Karp test. This is the square Figma occupied: owned workflows, owned data, positioned to ride every model release. But, they ended up being eaten anyway - not by a capability advance, but by the counterparty. The investment genuinely compounds, and it compounds in full view of a single provider, through a conduit that would take a rebuild to swap, under data terms they didn’t read closely enough. Everything looks right until the day it doesn't, and by that day the switching cost is huge.

Example: In June 2025, Meta took a 49 percent stake in Scale AI, and Google, Scale's largest customer, moved to cut ties within days; customers had been sharing proprietary data and prototype products with a vendor that now answered partly to a rival (CNBC). Nothing about Scale's capability changed that week. The cap table did.

3. Sovereign decay. Passes the Karp test, fails the Altman test. An operator hears the sovereignty argument, rents the GPU cluster, forks an open model, and owns everything - including the growing gap between what he owns and the frontier models. Every frontier release widens it. This move allocates real capital to owning intelligence that depreciates, while the data work that would actually compound goes unfunded. This is the next generation of on-prem server farms, built by people who think they are avoiding the last one.

Example: Bloomberg trained a 50-billion-parameter model on four decades of proprietary financial data (Bloomberg), and within months a Queen's University study showed GPT-4, with no access to any of it, outperforming the bespoke model on a range of financial tasks (HFS Research). If Bloomberg's data and capital could not hold this square, a GPU closet will not.

4. Compounding ownership. Passes both. The structured, proprietary data layer, connected to whatever model is currently best through a conduit that was designed from day one to be swapped. The asset compounds with any model's advances, open or closed, rented or owned. The model provider is a supplier decision revisited quarterly, not an architecture commitment made once. Exposure is governed by terms someone actually negotiated. Nothing in this square demos well, which is why so little of it gets built.

Example: Intercom launched its Fin agent on OpenAI's models, swapped to Claude when its own evaluation said Claude performed better (Intercom), then post-trained its own model on an open-weights base using years of proprietary support data, with stated plans to keep switching bases over time (VentureBeat). Two conduit swaps, one compounding corpus.

Two Proposals, One Platform

Consider a $40M professional services platform in a specialized B2B vertical. The company manages client relationships that span years; the interaction data sitting in its CRM, its ticketing system, and its email archive is genuinely proprietary. Two proposals are on the table.

Proposal A is an industry-specific AI agent from a credible vendor. It wraps a frontier model with a custom UI, vertical-specific prompt engineering, and a trained response layer for the most common client queries. The demo is excellent. The implementation timeline is ninety days. Twenty-four months from now, the vertical workflows it was built around are handled natively by the base models, the custom UI is a maintenance obligation, and the differentiated capability has been absorbed by the layer underneath. Rented decay. It fails both tests.

Proposal B is unglamorous: six months of work to structure, tag, and normalize the customer interaction data, define the escalation logic explicitly, build pipelines from a dozen source systems into a unified data layer, and stand up retrieval over that history. As usually pitched, Proposal B passes the Altman test and quietly fails the Karp test. The corpus compounds, and it compounds through one provider's API, with high switching cost and data terms as an unexamined afterthought. Compounding exposure, on a smaller stage than Figma's.

The good news is that moving Proposal B to the fourth square adds three requirements, not a second project.

  1. The retrieval architecture is specified model-agnostic, so that swapping providers is a configuration change.
  2. The data terms are audited and negotiated as a named deliverable of the project.
  3. The model relationship is managed the way the company already manages any single-source supplier: reviewed on a calendar, priced against alternatives, held to a documented exit path.

Three sentences in the spec. They are the difference between owning the asset and parking it in someone else's garage.

The Counter-Position: Just Go Sovereign

The strongest pushback runs like this: if the labs really do watch where value accrues and move in, then halfway measures are naive. Skip the frontier APIs entirely. Fork an open model, run it on owned or rented hardware, own the weights, and be done with the counterparty question forever. Palantir’s whole architecture exists because serious institutions reached exactly this conclusion.

For some organizations, the conclusion is right.

If you are a defense agency, a critical-infrastructure operator, or a company whose data corpus effectively is the product, in life sciences, say, where decades of experimental data are the crown jewels, then the model layer is not a supplier to you. It is a potential competitor or an unacceptable risk surface, and full sovereignty can be a rational allocation. The same holds for software companies whose product sits adjacent to what the models themselves are becoming. Figma was not paranoid enough.

But notice what those cases have in common: the counterparty had something to gain by absorbing them. For a VC-backed software company, the model layer is a potential competitor. For a PE-backed operating company, the model layer is a supplier.

Anthropic is not launching a product to compete with a $40M HVAC services platform. Nobody at a frontier lab is studying your dispatch escalation logic in order to enter commercial refrigeration.

And operating partners already know exactly how to think about suppliers, because they run the playbook on every other critical input: concentration risk, contract terms, switching cost, price trajectory. You make sure you are not sole-sourced, you read the contract, and you keep the ability to re-bid.

That is the operator-scale version of the second test. Supplier risk management. Treating it as less than that puts you in Figma's square. Treating it as more than that puts you in the server-farm square. The honest reading of the Karp argument, for this audience, is that it converts a question most operators never asked into a question they already know how to answer.

The Prerequisite Both Tests Share

Neither test can be passed with expertise trapped in people's heads.

If the "unique context" your team keeps referencing has never been written down, structured, or made accessible to any system, it isn't proprietary data, it's institutional memory. And institutional memory neither compounds with model advances nor belongs to anyone in a way a contract can protect.

You cannot run frontier retrieval over an illegible mess, and you cannot fine-tune a sovereign model on one either. "We've never written this down" fails both tests at once.

"Here is the structured corpus of what we know that nobody else has" is the raw material for passing both. That corpus usually does not exist when the AI proposal arrives. Issue 0008 will be about what it takes to build it.

The Lens

The ServiceNow story I opened with has an update. The small, bespoke models from the first bucket weren't all wrong. A few still run and deliver value. The ones that stayed relevant did so because they were trained on data structures we owned and maintained, not because the models were exceptional. The model is easy to replace. The data structure is not.

That is the two-test answer compressed into two sentences: own the structure; have the option to rent or own the intelligence.

The next AI proposal that reaches your desk will arrive with an answer to the vendor's question: what can this do today? Put it on the two axes instead.

Does the investment compound or decay as the models improve? And when it compounds, do you own what compounds, on terms you control?

The placing takes fifteen minutes. The square it lands in will tell you more than the demo will, and it will tell you before the deck goes to the board instead of twenty-four months after.

Get the next issue

Subscribe to receive future issues direct to your inbox.