Reading time: 6 minutes | Issue #41 | Book a Discovery Call
Happy Tuesday. Mark here.
Anthropic retired Claude Opus 4.1 on August 5, OpenAI shut down its Assistants API on the 26th, and Moonshot pulled Kimi K2.5 at the end of the month. A deprecation tracker I keep an eye on lists 48 model lifecycle events for 2026.
I bring this up because I spent the last six weeks talking with 20 PE operating partners about AI inside their portfolio companies, and most of them are picking a model or a framework as if that decision is the transformation. It's the part with the shortest shelf life.
Inside the Issue
The wall operating partners keep hitting: a demo takes a week, proving it runs at production quality takes six months
The loop that gets you to a measured dollar in four weeks instead of eighteen months, with the one rule most teams skip
Blackstone's Ode scales to Chamberlain Group, 48 models get deprecated, the enterprise evaluation gap widens, and insurers race agents into claims

The Demo Takes a Week. Proof Takes Six Months.
The specifics of those 20 conversations are confidential, so I'll give you the patterns instead of the names. Four came up again and again, and they connect into one story.
Start with who's ahead. If you're an operating partner still writing your AI strategy for the portfolio, you're in the majority: Grant Thornton's 2026 PE survey found 45% of PE respondents still in the pilot stage, eleven points above the cross-industry rate, and only 5% with AI fully integrated into operations versus 14% everywhere else.
So the discomfort is common. The same survey removes the comfort, though, because organizations that fully integrate AI were four times more likely to report revenue growth. The firms that started six months ago are compounding that lead, and the distance gets harder to close each quarter.
The second pattern surprised me until I heard it enough times. Large portcos told me legacy systems were the barrier. Small portcos told me it was lack of scale. Same conversation, opposite diagnosis, and both are wrong in the same way.
The constraint underneath is identical: no engineering methodology for deploying AI against a live production operation. Size decides where you start, not whether you can.
The third pattern is where the money is stuck. Most portcos are using AI for email drafts and meeting notes, and most of the operating partners I spoke with couldn't describe what comes after that.
The few pushing into claims processing, underwriting, or revenue operations are all hitting the same wall, the most expensive misunderstanding in the portfolio right now.
A demo takes a week. Proving it runs at production quality takes six months.
The teams can build the agent but can't prove it runs reliably enough to go unsupervised. The bottleneck is evaluation, and the wider market put a number on it this summer. VentureBeat's VB Pulse survey of 157 enterprises found that 50% had deployed AI agents that passed internal evaluations and then caused customer-facing failures anyway. A quarter had multiple such failures.
Two-thirds already allow, or plan within a year to allow, production deployment with no human in the loop. And 5% fully trust their automated evaluations to make release decisions. The autonomy is rising faster than the assurance underneath it.
Our read: the fourth pattern is the one that separates the operating partners pulling ahead. Instead of buying AI tools for the portfolio, they're changing how engineering capacity gets deployed: smaller teams, higher seniority, agents absorbing volume, humans holding judgment. One of them described replacing a full squad with a two-person unit running against the same scope.
The team structures and vendor relationships that built these portfolios through the last cycle were designed for a different kind of work, and they won't carry the portfolio through this one.
The ones who shipped did something different, and I've now watched this run inside enough engagements to trust it. The companies that failed ran AI transformation like a consulting program: a long diagnostic, a target architecture, a projected ROI built on assumptions, and eighteen months of integration before anyone measured a single dollar. The companies that shipped ran it as a loop.
Pick one workflow. Real users, real data, and an outcome the business already cares about and already measures, not a lab task. A team I worked with this year took claim denial reconciliation inside a healthcare billing operation, a process two people had been hand-writing rules for across 300+ denial codes.
Build the evaluation harness before the agent exists. Teams skip this step more than any other, and it's the step that decides whether the numbers in step four mean anything. You can't trust a measurement you build after seeing the answers, because you'll grade toward the output you already got. On that billing build, no update reached production without clearing a curated suite of real claim scenarios from the client's own data, and the flagship agent cleared 59 of 60 on it.
Prove the smallest useful slice in two to four weeks. Build the thinnest version that touches production reality rather than the whole system, and if reaching a slice takes three months, the scope is still too big.
Measure throughput, quality, and economics. Three numbers, from the harness you built in step two, against the workflow you picked in step one.
Decide. If the numbers make sense, harden it for production and transfer ownership so the portco runs it, not you. If they don't, stop and keep the budget. A stopped workflow after four weeks is a cheap, correct decision. An eighteen-month integration nobody had the nerve to kill is the expensive one.
Pick the next workflow and run the loop again.
Now go back to those model deprecations: Opus 4.1, the Assistants API, Kimi K2.5, all retired in one month. The framework you standardize on this quarter gets deprecated next quarter, but the workflow knowledge, the evaluation criteria, and the operating ownership your team builds while proving that first slice don't expire when a model does.
That knowledge is the asset; the model is a component you swap. One workflow, one measurement, one decision based on what you observed. That's the first step to AI transformation that scales.
Who should be uncomfortable reading this: any operating partner whose AI plan is a vendor shortlist and a target architecture, with the first measured dollar scheduled for next year.


Run the loop on one workflow. A four-week plan you can start Monday.
Pick the workflow before you read the rest of this: one process with real users, real data, and a number the business already tracks. Claims, underwriting, collections, renewals, and support triage all qualify, as long as you'd feel the outcome move.
Week 0, before you write a line of agent code: build the harness. Assemble 30 to 60 real cases from your own production data, with known-correct answers your domain experts agree on. Write the scoring: exact-match where the answer is deterministic, a rubric where it's judgment.
Sequence is the discipline that makes this work, because a harness that exists before the agent can't be tuned to whatever the agent happened to produce. On the billing build I mentioned, the harness caught regressions when payers changed rules and when a provider shipped a new model version, before a live claim was touched.
Weeks 1 to 2: build the thinnest slice that touches production reality. Take the single highest-volume path through the workflow rather than the whole thing. Structure and enrich the inputs before the model sees them, so the agent reasons from ranked facts instead of hunting for context. An analysis of 100+ production deployments traces most production failures to retrieval, not to the model's reasoning.
Weeks 3 to 4: measure three numbers and decide. Throughput (volume the agent clears without a human), quality (the harness score against your real cases), and economics (cost per unit of work, model plus infrastructure). Then make the call in the open: harden and transfer, or stop and keep the budget.
The staffing change underneath all of this: the operating partners moving fastest stopped sizing teams the old way. Instead of a full squad, one senior engineer plus a delivery lead, with agents handling volume and the humans handling judgment and verification. That's the shape of our Velocity Pod, and it runs at $15,000 to $20,000 a month against scope that used to demand a squad at several times the cost.
The pod is one way to staff it. The pattern underneath, org design first and AI fitted to the design, is what the fastest operating partners share.
Run it on the one workflow your team is already arguing about, and in four weeks you'll have a number instead of an opinion.

01 Blackstone's Ode is now inside Chamberlain Group. The $1.5B implementation venture from Blackstone, Hellman & Friedman, and Anthropic, a team PYMNTS counts at roughly 160 AI experts, is now deployed at six portfolio companies with plans for 25 of Blackstone's 270+. The named example: Chamberlain Group, the garage-door maker, adding AI-built software features. Rodney Zemmel, who runs Blackstone's operating team, on why they backed it: "It's not like a bunch of kids who just learned AI." When the world's largest alternative asset manager builds a senior engineering shop to embed in portcos, the "we'll buy a tool" era is over.
02 The model graveyard keeps filling. The August deprecations at the top of this issue come from benchr's tracker, which lists 48 lifecycle events across providers for 2026. If your portco's AI architecture is welded to one model or one vendor SDK, you're carrying migration debt you haven't priced. Build the workflow and the eval criteria to outlast the model.
03 The enterprise evaluation gap, quantified. VB Pulse surveyed 157 enterprises: half had shipped agents that passed internal evals and failed customers anyway, two-thirds are moving toward deployment with no human review, and 5% fully trust their automated evals to make the release call. The best argument I've seen for building the harness before the agent.
04 PE's readiness problem is a governance problem. Grant Thornton's 2026 PE AI survey again: only 9% of PE respondents are confident they could pass an AI governance audit within 90 days, the lowest rate of any sector in the survey. Grant Thornton partner Tom Libeg's line stuck with me: "You can't walk into a portco and mandate an AI deployment. Management teams have their own priorities. What moves them is evidence." That's the loop, stated by someone outside my building.
05 NTT DATA shipped agents straight into claims and underwriting. In early August the firm, which serves 10 of the world's 25 largest insurers, launched configurable AI agents for underwriting and claims, flagging high-risk claims earlier and standardizing the process. The demo layer for regulated workflows is now a commodity, and the work moved to the proof layer: evaluation at production quality.

The Loop vs. The Program. A one-page side-by-side you can put in front of an operating partner or a portco CEO to settle how the first AI project gets run.

Most of those 20 conversations ended the same way: the operating partner wanted a team that could ship a working workflow against production data in weeks, prove the economics, and hand ownership back so the portco runs it.
That's the exact shape of our Velocity Framework, and it's how our pods work inside portfolio companies now.
If you've got one workflow where the outcome is measurable and the current answer is still a demo, that's the one to start with.
Until next Tuesday,
— Mark Ajzenstadt, Founder @ Limestone Digital
P.S. If this issue described your portfolio, book a discovery call with me. Bring one workflow, and in 30 minutes I'll tell you whether it's a good first slice and what the evaluation harness for it would look like.
