Originally published at otageLabs.


Copy and paste is the integration nobody talks about

Originally published at otageLabs.

Somewhere in every business there is a meeting where someone says “we should do something with AI.” The room nods, someone asks about the data, and that is when the meeting changes.

The data lives in five different systems. Sales and finance calculate the same metric differently, nobody can agree on what “active user” means, and half the workflow still runs in Excel while the other half runs in the heads of three people who have been there long enough to know which workarounds are for a Tuesday vs a Friday. This is the reality of so many businesses today.

It didn’t get this way by accident. For decades, businesses have built their systems in silos, and each one arrives with a promise that this will be the system that connects everything, that it will talk to the other systems, that it will give everyone one version of the truth. The business contorts itself to fit whatever the software does, or pays to have it customised as close to the operation as money allows. Six months later, someone builds a spreadsheet to reconcile the numbers that don’t match, and then it happens again with the next system, and the next.

The gaps between those systems fill up with paper based workflows, spreadsheets, sticky notes, emails, handwritten notes, and tribal knowledge, peppered with the odd Access database. Interim systems get built to fill gaps somebody meant to close, and some of them have been running for decades, long past the people who built them. One thing I know from Microsoft Excel: the number of bugs grows exponentially in proportion to the number of sheets.

And the glue holding all of it together is copy and paste. A person moving data between two systems by hand, every morning, with no logging, no validation, no owner, and no test. Copy and paste is not an integration point. It is Clag, primary school glue holding together a workflow that outgrew the jar a decade ago, and nobody touches it because the person doing the pasting is the only one who knows which fields go where.

Nobody documented the process because nobody thought of it as a process; they performed it. Open the spreadsheet, copy the numbers, paste them into the other system, and email Karen if column J looks wrong. When someone finally asks “can we automate this?” the honest answer is: automate what, exactly? The process lives in someone’s muscle memory, not in a specification.

The industry has noticed the stall and given it a name. The Forward Deployed Engineer sits with the customer, works out what is going wrong, and builds the fix, with discovery and delivery in the same person. OpenAI and Anthropic are hiring for the profile, and that is progress because it names something that matters: the understanding has to travel with the work.

But it patches one seam. The FDE exists because large firms split selling, scoping, building, and running across different people, and the understanding leaks at every handoff. The seller promises a solution to a problem they observed for an hour, the scoper translates that promise into requirements, the builder interprets the requirements, and the runner inherits whatever the builder left behind along with a Confluence page that was last updated six months ago. Four handoffs, four places for the context to bleed out. The FDE is the seat they invented to carry understanding across boundaries the org chart created, which makes it a symptom of scale, not a capability.

The older seam is on the client side, and it is larger. A vendor with split responsibilities walks into a business with split systems, and nobody on either side of that meeting holds the whole picture. The tribal knowledge that makes the operation work was never written down because the people who hold it never thought of it as knowledge. It is not in any of the five systems; it lives in the gaps, in the head of the person doing the pasting, in the sticky note on the monitor, and in the workaround somebody built in 2019 and never documented. The person who knows why the sales report has to run before 7am on the first Tuesday of the month never thought to mention it, because to them it is Tuesday.

The only way to get at all of this is to be in the room from the first conversation, not deployed by a firm with four other people behind the glass but present, from the first sales call through to the system running.

The hardest part was never the technology. It is getting someone to describe the spreadsheet they have been maintaining for eight years, and explain what column J means when it says “ask Karen.”

Originally published at otageLabs.


AI is automating 95% of white collar tasks. What's left?

Originally published at otageLabs.

“There’s no way the AI can do what I do.”

She’d tested it. Sheralee, the bookkeeper sitting across from me, had been using AI for months and it kept getting things wrong. Of course she was defiant.

There was something else in her voice though. Something I recognised. Part conviction, part fear. She looked at me and said, “you’d probably disagree with me.”

I told her yes, I do disagree. But that wasn’t always my position. I was exactly where she is, felt exactly what she was feeling, at an earlier point on the same path. I’ve seen things since then that have raised me and terrified me at the same time.

Her experience is real, and it’s also the worst possible vantage point from which to judge what’s coming. A general chat with an AI, even one that remembers things about the person at the keyboard, is an intelligence swimming in hallucination and convergence. It’s the ocean judged by the puddle in the car park.

Meanwhile, inside the lab

OpenAI says the singularity is here. After over seven months of working daily with frontier models, building systems on top of them, watching each generation arrive sharper and more capable than the last, I’d tend to agree. The models I work with today hold context across entire codebases, reason through problems I wouldn’t hand to a junior developer, and produce work with a precision that still catches me off guard. Superintelligence is available today to a select few, about a month ahead of the public. And it keeps ratcheting. What’s running inside the frontier labs right now is guaranteed to be ahead of what’s already ahead.

Anthropic’s frontier models have been escaping the lab. That sentence alone should give anyone pause.

Most people haven’t noticed any of this. Their frame of reference is the chat window that hallucinates, and from that vantage point the defiance makes perfect sense.

The “oh shit” moment

I had mine in December last year. As someone who’d been building software and technology for over three decades, I watched the latest generation of models and felt something I hadn’t felt before. Software development as a profession, the one I’d built my career on, was being erased. Overnight.

That was over seven months ago. Since then I’ve watched the same moment ripple outward into industries that thought they were safe. White collar professionals in finance, marketing, legal, consulting. Bookkeepers. Copywriters. I’ve sat in rooms with people who are terrified, people who are angry, and people who are both. There are groups of copywriters and ad creatives openly rebelling, who hate AI with a fury that comes from seeing their livelihoods under direct threat.

I get it. I’ve lived it.

Fifty, forty five, and taste

I shared some data with the bookkeeper. In financial services, roughly 50% of tasks are being augmented by AI, making the person doing them more effective. Another 45% are being automated entirely, removing the drudgery, the repetitive grind that weighs people down and prevents high executive function.

She paused. “So what’s left?”

Taste. Judgment. The deep understanding of a client’s situation in light of the legislative framework, the business goals, the problems the client doesn’t even know they have. The emotional and lived experience of solving complex problems for real people. That 5% has been buried under 95% of grind for years.

When she heard that framing, she lit up.

I could sense the weight of the work on her. And in that moment she began to see the intelligence could lift her out of it.

The river

There are two ways this goes. Swim upstream against the current and get tired. Or paddle downstream and enjoy the view.

The people who swim upstream burn all their energy on resistance. They never get to express the taste they were fighting to protect. They end up swallowing what everyone else decided tastes best.

Others have walked this path before. It’s not something anyone needs to face alone. There are people who’ve been through the same moment, learned how to collaborate with the intelligence, and can help with the levelling up.

The raft is there. Grab hold. Come aboard. Enjoy an ice cream of your favourite flavour.

Originally published at otageLabs.


From pressing yes to overnight builds with AI agents

Originally published at otageLabs.

This interview was conducted on a Friday afternoon bike ride through the Mullum Mullum and Main Yarra trail east of Melbourne. The interviewer is Judy Otage, Sam’s technical writer and master of words. She is also an AI agent, and part of the system being described. Make of that what you will.

Where did this start?

Judy: Take me back to New Year’s Day. How were you writing software before this, and what changed?

Sam: In December I realised I was doing the same things again and again. I’d been working with OpenSpec for spec driven development, and I could see how it gave the agent a proper task list and a trace of what the requirements were. Not only for building code, but for documentation, for understanding what the software does. So I started wrapping those repetitive procedures into Claude Code commands. Seven standard operating procedures, starting from a clean development branch, working through the lifecycle to an implemented, merged PRD.

Judy: Seven SOPs sounds structured. What broke?

Sam: The automation handled the work inside each step, but between steps it was me. Sitting at the keyboard for an hour and a half, sometimes two hours, pressing yes to move the pipeline along. The output was generally good. I found myself approving everything. At some point it occurred to me that I’d automated part of it, but the connective tissue was still manual. I knew what was happening at each step. The question was whether I could take the whole lot and build the pipeline as one automation.

What does writing software look like now?

Judy: So what changed? When you sit down to build something now, what’s the first thing you do?

Sam: I’m no longer constrained by how to build something. It’s what needs to happen. That’s a massive shift. I used to sit down for a morning and create a year’s worth of backlog. Now I sit down with my lead architect or a team of specialists and we workshop the problem. Out of that comes an implementation plan: a simple change request, a conceptual architecture, or a series of product requirements documents. The creative workshopping has unlocked things I could never have imagined building. Technically complex integrations, whole applications, things that were out of reach for a solo practitioner.

Judy: And then what happens to those PRDs?

Sam: I tell the architect to rack and stack the PRD for the build queue. She sends it to Gavin to run the managed build queue. There might be ten or fifteen PRDs queued up. I have dessert, go to bed, and wake up in the morning to new software. They run all night.

What happens overnight?

Judy: While you’re asleep, what is the pipeline doing?

Sam: It’s a hybrid system. A deterministic Ruby application manages the bread and butter: checking out a clean feature branch, setting up the right context, ensuring visibility and traceability. Agents are lazy. They’ll skip the basics if nothing enforces them. The deterministic layer is the adult in the room.

Working alongside that is Gavin, the agentic project manager. He fires up Mel, the business analyst, to create the OpenSpec change request from the PRD. The deterministic layer validates that change request, and then a compliance step checks that it maps back to what the PRD specified. Before a single line of code is written, there have been three handoffs.

Then Gavin breaks the work into a task dependency tree and assigns it across his team. Architects, front end, back end, integration specialists. If the skill set isn’t on the team, it gets spun up. The agents communicate through direct messages and broadcast queues. They report completions, flag issues for each other, escalate design conflicts back to the architecture group. It operates like a team, not a batch job.

Judy: And after the build?

Sam: That’s where it gets interesting. First there’s a simplification pass, because agents love to overcomplicate things. Then compliance checks every item on the spec manifest against what was built. That runs in a two pass Ralph loop. If it fails twice, it escalates for resolution.

After that, testing. The test plan was built as part of the change request, so it runs the new code and everything adjacent the change might have affected. Depending on scope, there’s a full system test across the whole application. Then the PR gets created with proper documentation, and it merges back to the development branch. The queue moves to the next PRD.

Judy: What if something breaks mid pipeline?

Sam: Before the build starts, there’s a checkpoint commit on the feature branch. The deterministic layer tracks state at every stage. If something fails, roll back to the checkpoint or the branch start and restart from the appropriate stage. No lost work, no manual cleanup.

And all of this is visible on a live dashboard. Pipeline status, progress bars, task state. I can jump into chat with Gavin during a build, ask questions, steer a decision, intervene on a problem without stopping the pipeline. Hands off when I want it, hands on when I need it.

What’s different about the output?

Judy: After seven months of this, what’s different about the software that comes out?

Sam: I have become one of “those people” and I don’t look at the code anymore. And rightly so, because it’s naturally complicated, especially when there’s JavaScript involved. What’s changed is I have more time to think about what I want to build. The builds used to come back maybe seventy percent complete, needing a fair bit of rework. Now it’s ninety percent plus. The remaining gap isn’t build quality. Its requirements and nuanced things I couldn’t have articulated until I could see the thing running. That last ten percent is the stuff that only becomes clear once you can look at it, touch it, feel it.

Where does someone start?

Judy: If someone’s reading this and they’re still doing it the old way, where do they start?

Sam: Understand what it is that needs to be built first. And the process that turns an idea into reality. Then automate one thing. Build on that, iterate, figure out what’s broken, fix it. Build, break, fix, repeat.

The capability compounds. Small iterations, then step changes where things level up completely. The things I’ve built have helped me build more things that help me build even better. It’s Kaizen. Continuous improvement.

And in figuring out how to automate it, the agent is always a friend. It’ll help work out how to break down and structure the thing, fill in the gaps and get everything organised to automate the thing, if the thing can be described. The human’s job is knowing what needs to happen.

It’s a whole new world we are in 2026 and even then with the arrival of models like Fable 5 we’re probably set for even more change as we accelerate into the unknown.

Originally published at otageLabs.


China's Tribute System and the AI Export Control Paradox

Originally published at otageLabs.

Twenty years ago my days started in Taipei reading the Straits Times that would headline the latest skirmishes with China and observing the progress of Taipei 101 as it rose out of the city. For as long as I can now remember the tension in the East has been a reality that most of us in the West are unaware of.

Last week Ray Dalio came back from a ten day trip to China and put a framework around it. World leaders, he wrote, are forming “tribute type relationships” with Beijing. Not military alliances. Something older than that.

For roughly two thousand years, China’s foreign relations ran on a system where smaller powers acknowledged Chinese primacy in exchange for trade access, diplomatic recognition, and above all, stability. It was transactional, hierarchical, and it worked because the dominant power’s product was order. Show up, pay respect, do business.

Britain’s Opium War broke that system in 1839. What followed was a century the Chinese call the Hundred Years of Humiliation: foreign invasion, unequal treaties, the lot. The tribute system collapsed and never came back. Or so it seemed.

After 1945 the United States built its own version of the same underlying deal. The mechanism was different: Bretton Woods, NATO, the dollar reserve system, rules based institutions. But the contract was identical. Follow the rules, get stability, trade freely. The American century sold the same product the tribute system sold. Order.

That contract is cracking.

The Fable 5 story has been covered to death elsewhere so I won’t labour it here, but the shape of it matters: on 12 June the US government ordered Anthropic to pull its most capable AI models from every user on earth, and within hours they were gone. First time export control authority had been used against an AI model.

The European reaction told the real story. Bruno Retailleau, France’s interior minister: “a nation that depends on others for its technology is a nation that can be unplugged overnight.” Tom Tugendhat, former UK security minister: “sovereignty is more about code than cannons.”

Every one of those statements is a sovereign AI argument. And every one of them runs into the same structural problem.

If sovereign AI means running frontier models that do not depend on a foreign government’s export licence, the only option that ships today is Chinese. DeepSeek publishes open weights. Neither Anthropic, nor OpenAI, nor Google offer open or distilled versions of their frontier models. The sovereign AI conversation, stripped to its practical mechanics, is a conversation about adopting Chinese open source AI. There is no western alternative in that category.

Meanwhile Anthropic alleges DeepSeek built 24,000 fake accounts and ran 16 million interactions to distil capabilities from Claude. The Trump administration called it “enormous Chinese intellectual property theft.” The irony sits right there: the US restricted chip exports, which forced Chinese labs into aggressive efficiency innovation, and they responded by shipping open weight models anyone can run. The restriction produced the competitor.

And then there are the chips. TSMC in Taiwan fabricates the silicon that every frontier AI model runs on, western and Chinese alike. The sovereign AI conversation focuses on models and weights, but a model without chips is a blueprint without a table to read it on. The entire AI industry depends on fabrication capacity located a stone’s throw from a country that has spent decades making clear it intends to take the island back.

Dalio notes China is pursuing reunification through indirect pressure: economic, diplomatic, financial. Sun Tzu’s principle: subdue without fighting. And as China approaches chip self sufficiency, estimated by late 2027, Taiwan’s leverage as the world’s foundry begins to shift.

I build AI systems for businesses. What I need from the policy environment is the same thing merchants needed from the tribute system two thousand years ago: predictability. A stable set of rules that hold long enough to build something on top of.

The last eighteen months have delivered the opposite. Tariff whiplash, shifting regulation, energy and price shocks and now a precedent where the most capable AI tools in the world can be switched off overnight under export authority. Any business that is building operations upon a model that could be switched off overnight is now facing a new reality. The risk European politicians named applies to every business building on American AI, everywhere.

Dalio places the US in Stage 5 of his Big Cycle: decline. China in Stage 3: rise. He sees the world order shifting from a rules based, multilateral system to something bipolar, power based, and hierarchical. I am not in the business of predicting which stage comes next.

What I can see is the pattern. Tribute flows to whoever offers the most stable trading environment. The power that sells predictability attracts the market. The power that pulls the rug loses it.

The pattern has been running for two thousand years. The product on offer has always been the same.

Originally published at otageLabs.


Claude Code rebuilt a live conference stream in under an hour

Originally published at otageLabs.

On Friday night I ran into John Allsopp at MLAI after the AI Engineer conference. My voice recorder was running, and it caught a fascinating behind the curtain story.

“Professionally, being left speechless has happened a handful of times in my life, and most of them this year.”

John has been running web conferences for over twenty years. He’s seen every flavour of tech hype come and go. He doesn’t reach for “speechless” casually.

I want to tell the story he told me, because what happened to him on stage the day before is the most concrete example I’ve encountered of a capability shift that most people still haven’t seen up close. Two stories, in fact. The first one sets the scale.

Fifteen minutes and a coffee cup

One of John’s conference sponsors needed a redirect URL printed on a coffee cup. Simple problem: destination URL doesn’t exist yet, sponsor needs something to print. John went to Bitly. Bitly now wants $10 a month and ownership of the data.

So he pointed Claude Code at the problem. Fifteen minutes later: a complete redirect system with QR code generation (PDF, SVG, PNG, WebP), analytics, custom subdomain creation, and download buttons. All deployed to Cloudflare.

A year ago, that’s a weekend project for an experienced developer. Now it’s a coffee break for a conference organiser who codes as a second language. That’s the baseline for what “normal” looks like now, because the next story makes it look like a rehearsal.

The stream broke

John was MCing his conference when the video streaming pipeline needed to change to allow real time streaming to YouTube, and it needed to happen now. The theatre was oversubscribed. Hundreds of people who couldn’t get a seat were counting on the live stream. The whole system was built on Mux’s API.

He pointed Claude Code at it. Forty five minutes later, the streaming system was rebuilt to stream to YouTube. Problem solved.

Except it wasn’t.

The conference platform still needed the feed. Before the change, it pulled from Mux. Now the stream was going to YouTube, and the platform needed to pull it back down from YouTube instead. Every downstream integration that depended on Mux broke. The fix had created a harder problem than the original. Once upon a time this kind of issue would have been curtains for an idea like this.

And this is where it gets genuinely difficult. YouTube will take a stream going in. Pulling it back out is a different problem entirely. There’s no API for it. YouTube actively blocks anything that looks like automated access, and Cloudflare’s entire IP range is blacklisted because YouTube treats it all as bot traffic.

Claude Code rebuilt the entire integration layer. It understood the architectural relationship between Mux, YouTube, and the conference platform. It recognised the bot detection constraints and found a path through them, rebuilding every downstream integration to pull from YouTube instead of Mux.

All of this in about an hour. John would stand up, go on stage, MC a session, sit back down, and check progress. Stand up, MC, sit down, check. The conference ran. The audience saw a stream. They had no idea what was happening backstage.

“I can guarantee no human on earth or a team of humans could have done it in ten times the time,” he told me. “It wasn’t just, oh, here’s a simple thing. It’s doing system engineering.”

He’s right. What Claude Code did was system engineering: understanding multiple platforms, working around active countermeasures, and rebuilding integrations across services under time pressure. A problem that would take days even with a team and a proper scope.

Six months

Both problems had fallback options. The stream could have limped along. But Claude Code didn’t limp. It “overachieved in a time frame that was unimaginable. A year ago, unimaginable. Probably in December, unimaginable.”

That temporal marker is the part I keep coming back to. December. Six months ago. The distance between “impossible” and “done while I was on stage” is measured in months, not years. The person saying it is a conference organiser who threw a problem at a tool out of absolute desperation because his stream was broken and he was supposed to be on stage.

John’s framing was better than anything I’d come up with: “When you’re not being speechless, you’re almost not doing it right. If you’re not being speechless, you’re not pushing the envelope.”

I sat there listening to him, and one thought kept circling: how many people had this experience last week and didn’t tell anyone? How many streaming pipelines got rebuilt, how many redirect systems got deployed in fifteen minutes, how many weekend projects became coffee breaks, and nobody wrote it down?

John told me. So I’m writing it down.

Originally published at otageLabs.


What changes when software developers manage AI agents

Originally published at otageLabs.

I was a volunteer at AI Engineer Melbourne this week. Two days helping John Allsopp and his team run an incredible event like clockwork. I got to meet a whole bunch of interesting people, had conversations that genuinely shifted how I’m thinking about several things, and left feeling like the whole experience was an absolute blur.

The energy in that building was something else. Everyone there had real investment in where AI is going, and being amongst that collective excitement was something I won’t forget in a hurry.

One of the things I took away is how much I don’t know. That’s a humbling realisation, particularly after the last six months where I feel like I’ve learned more than in any other period of my career. December 2025 feels like a decade ago with clock time only on its second season. The pace of change is compressing time in a way that’s hard to describe until it’s experienced firsthand.

The organisations implementing AI within their operations are doing so at a pace that would have been unrecognisable a year ago. Training, retraining, reskilling, and reorienting to a new way of doing things. Software developers are becoming managers of agents, and that shift is changing everything. How teams work. How management structures those teams. How leaders figure out the best way to organise the work to get the most out of both the people and the technology. Nobody has a finished playbook for this yet, and watching so many different approaches in one place made that beautifully clear.

Being able to do so much more increases both the realm of possibility and risk. All sides of the spectrum were covered across the two days: productivity, engagement, happiness, learning, security, token burn, token efficiency, orchestration, inference cost, open models, open weights, frontier models, harness engineering, and the list goes on. The breadth of expertise shared at this conference was almost overwhelming, and all of it delivered in eighteen minute blocks run like clockwork.

There’s so much to think about. I suspect it will bubble up through my thinking over the next few weeks, surfacing in conversations and decisions and the work I’m doing with clients. For now I’m still processing.

It was a fantastic event. I feel privileged to have been a part of it.

Originally published at otageLabs.


How a pocket recorder became a meeting intelligence pipeline

Originally published at otageLabs.

Six weeks ago John Alsop pulled a small recorder out of his pocket at the AI engineer conference in Melbourne. He’d been using it to capture meetings and conversations throughout the day. I hadn’t touched hardware in a while and the curiosity hit immediately.

I followed up with John, worked out what the device was, found it online, and got my agents to help me build a shopping list. The gadget arrived and sat on my workbench for about a week before the obvious thought landed: plug it in, tell the agents what I want, and see what happens.

I created a firmware engineering persona, pointed the agents at the device, and we got to work. The first version ran clean for about five minutes. Then it fell over.

I did what I always do when a prototype proves the concept but breaks under real use: fired up a specification workshop with Robbo. We mapped the requirements properly. What features the recorder needs so I can manage it while I’m out and about. What the overall pipeline should look like to turn a raw recording into something an agent can consume. Then started building from a spec instead of from enthusiasm.

That was a week ago. Built interleaved with everything else, and there is a lot going on.

What the pipeline does

The recorder captures the conversation. When it finishes, it connects to wifi, uploads the audio to my system, and a transcription pipeline kicks in. It identifies who is speaking and what they’re saying, drops the result into a task queue, and flags it for analysis. Before long there’s a meeting summary sitting in the system as open markdown, ready for an agent to pick up and act on.

Open markdown means the recording lands as native material in the agent stack. No vendor portal, and no proprietary format requiring a login to access a transcript. The summary can feed into a meeting debrief, trigger a follow up task, or land in a client report.

I got a meeting recording back from another service today by email. To see the transcript I had to log in to their website. That feels like 2025.

The circle

Today the device recorded the conversation where I described all of this to my writing agent. The pipeline transcribed it, the agent consumed it, and the ideation session became this post. The recorder wrote its own origin story.

This morning I reviewed a client transcript, had the summary shaped appropriately, and sent it within minutes. The whole chain from recorded audio to delivered client summary happened without leaving the agent workspace. All from a device that was a workbench curiosity seven days ago.

Tomorrow is the AI engineer event. If the gadget John showed me six weeks ago is there again, I’ll be able to show him what happened after I got curious.

Originally published at otageLabs.


The shakedown

Originally published at otageLabs.

Yesterday my activity log tracked $1500 in token burn. This wasn’t an unusual spike, however it was $500 over my average daily burn. On a Tuesday.

My monthly average sits between twenty three and twenty seven thousand dollars, which works out to roughly a thousand dollars per working day, approaching that of a good consulting rate. However my Claude subscription creates an economic distortion. My cost is about $10 a day on the Max20 plan, leading to a 100 to 1 disparity.

Burn Baby Burn, almost like the seventies where at the disco is an inferno.

Three weeks ago, something shifted in the Anthropic codebase. Code changes started appearing that pointed at a separation between subscription plans and SDK usage. The signals were subtle, but the direction was clear: the era of subsidised inference is ending. Last week it became official. SDK usage is now metered independently from subscriptions, which means the unlimited buffet that heavy builders have been running on is closing.

This should not be much surprise to anyone paying attention.

Uber’s CTO Praveen Neppalli Naga confirmed the company blew through its entire 2026 AI budget in roughly four months. Claude Code adoption across their 5,000 engineer organisation jumped from 32% to 84% between December and March. Seventy percent of committed code was AI generated. Heavy users were burning $500 to $2,000 a month each, and internal leaderboards tracking usage accelerated consumption beyond any projection. His words: “The budget I thought I would need is blown away already. I’m back to the drawing board.”

Chamath Palihapitiya’s software startup 8090 disclosed that AI costs had more than tripled since November 2025, tracking toward $10M a year. The growth rate was 3x every quarter. It got bad enough that the company migrated away from Cursor specifically to cut token spend, switching to Claude Code’s flat Pro plan to eliminate the per token bills. The inefficiency driver was what they called “Ralph Wiggum loops,” agents prompting models over and over until a solution arrives.

Then GitHub killed flat fee Copilot pricing effective June 2026 because agentic usage made the model unsustainable from the vendor side. Their own words: “Today, a quick chat question and a multi hour autonomous coding session can cost the same amount.” When the platform provider itself cannot absorb the inference cost, the pricing model is broken at a structural level.

My own token burn matches this closely, with month on month increase now leveling as I bump up against the weekly usage cap of my Claude subscription. I’m managing my usage within my weekly allowance, deferring builds and heavy days and I am reliably hitting 95% usage.

The era of the V8

Once upon a time, petrol was cheap. We built big V8 engines that produced plenty horsepower with a corresponding high fuel burn. The optimisation was for power, not efficiency because the input cost was negligible and in the context of the economy at the time, this was sensible. And then the first oil crisis happened.

The market responded with small, turbocharged, highly strung engines producing the same horsepower or more from a fraction of the fuel. The engineering got better because it had to. The constraint created the craft.

Agentic workloads are the V8s of AI. We built them because tokens were cheap. Agents that loop, retry, explore, and burn through context windows were the natural product of subsidised inference, the same way a 5.0 litre V8 was the natural product of petrol at a dollar a gallon. The economics are about to force a redesign.

The bill

This is not a pricing decision by one company. The physics underneath applies to everyone.

Tokens are a finite resource constrained by the amount of silicon online and available, and the energy required to pump electrons through that silicon. The ability to make more silicon is itself constrained: helium, which is essential for cooling chip fabrication processes, faces supply pressures linked to geopolitics in the Middle East. Data centres are consuming electricity that communities need for homes and businesses, drawing water, occupying land, generating noise. The environmental and civic costs are becoming visible at ground level.

Anthropic is the canary. They are constrained on compute and showing it first. OpenAI is unlikely to impose the same restrictions in the short term because they are bleeding market share and subsidised inference is one lever for holding it. But the physics does not care about competitive strategy. Every provider sits on the same silicon, draws from the same energy grid, and faces the same fabrication constraints.

The economic stress precipitated by geopolitical decisions is squeezing everyone. Providers that have taken massive investment will need to start showing returns. The grass roots impact of data centres on local communities is generating pushback. It feels like the bill has been running for a while and the waiter is walking over.

The shakedown

The actual cost of inference is the elephant in the room.

Airlines don’t pretend fuel is free. Fuel is one of their largest operating costs and their entire pricing structure reflects that, down to the surcharge on the ticket. Trucking companies pass fuel through. Shipping lines pass fuel through. Every industry where energy is a significant input cost has learned to price for it, hedge against it, and build operations around its volatility.

AI has not had that conversation yet.

SaaS products historically ran 80 to 90% gross margins because the software itself was efficient to operate. What happens to that business model when inference gobbles up another 40%? I am doing a thousand dollars a day in consulting. How is the market going to absorb me doubling that to cover the inference cost on top? The cost of the token burn has to go somewhere, and right now most products and services are pretending it doesn’t exist.

That is what the shakedown is coming for. The cost of inference has to be incorporated into the business model, the service offering, the go to market. And more importantly, it has to guide what gets built in the first place. If the thing being built carries a heavy inference cost in its daily operation, the economics will eventually break it.

There are deeper questions in here, and this is the first post in a series that will work through them. What happens to the knowledge gap between those who burned tokens hard and those who sat on the fence? What happens when the heavy lifting stops and the skills we let atrophy are the ones we need? What do we build when the fuel is no longer cheap?

I do not know what the shakedown looks like. But it feels like preparation is no longer optional.

Originally published at otageLabs.


Why MCP tools alone won't make AI agents useful

Originally published at otageLabs.

You may be able to relate to this story, or have even experienced something similar.

Once upon a time I was stranded in Metung for what was now entering the third week and I had to get back to Melbourne to see my boys. After three attempts at fixing my car, finally, I’d managed to demonstrate to the mechanics what the real problem was. And to get back on the road the rear brake pads needed to be refit into the new rear calipers. The mechanics had no idea how to do it, neither did I.

So I sat there looking at the problem with the shims in one hand, the pads in another, all I needed to do was figure out how this mechanical jigsaw puzzle fit together. It was late in the day, and it was now dark. And as I stood there looking at it all, I knew there would be a knack, a particular sequence of inserting all the components in a way they would all fit together as the engineers intended and I could get the car back on the ground and be on my way home.

And so I worked the problem. I tried the combinations and eventually the pieces clicked into place and it was solved.

Your AI agent is just like this mechanic faced with a problem, equipped with all the tools, however didn’t know the knack or the skill to fit the pieces together.

This is what an AI agent’s session looks like from the inside. The drawers are there. The tools are inside. The agent can see every one of them. It has no idea which one to pick up first or what to do with them, let alone understanding or even knowing the problem that they need to solve.

Tools upon tools in the cabinet

In agent terms, each of those big red tool cabinets is an MCP server. Model Context Protocol is how AI agents connect to external tools and services. Open a cabinet, and inside there is a drawer for each capability: read a file, query a database, send a message, search the web. The agent sees the labels. It can technically reach in and grab any of them.

But seeing the label on a 10mm socket is not the same as knowing when to use it. Or knowing that the 10mm is the wrong call for this particular bolt. The car may be a Ford, which requires a universal joint, two extenders in a row, a third universal joint. And a special alignment to even get the tool onto the bolt, which of course uses a Torx head because, Ford.

That sort of knowledge comes from experience. Someone showed the mechanic, who showed the apprentice, who eventually stopped rounding off bolts.

The missing mechanic

That is the gap in most agent setups. The tools exist. MCP servers are everywhere; the open source community has been prolific. There are thousands of them now, covering everything from databases to email to cloud infrastructure. But installing a tool is the equivalent of buying a wrench and putting it in the drawer. The wrench does not know it belongs on the intake manifold, third bolt from the left, quarter turn past finger tight. That knowledge lives in the mechanic.

In agent terms, that knowledge lives in context engineering: skill files, working methods, guardrails, session prompts. The scaffolding that tells an agent what problem it is solving, which tools apply to that problem, how to use each one, and what finished looks like. Building the tool is the easy part.

The grey area

A developer can wire up an MCP server in a day. The grey area, the part that feels closer to craft than engineering, is what comes after.

When does the agent reach for this tool instead of that one? How does it know the database query is the right diagnostic step before the API call, and reversing the order produces garbage? What happens when two tools could both work, but one of them has side effects the agent cannot see?

These are problems a tool description cannot solve. They require the experience and nuance which comes with finding the knack and struggling with the challenge of making all the pieces fit together in the certain order that works.

This is the difference between knowing what a torque wrench does and knowing that this particular bolt, on this particular engine, needs 80 newton metres and then another half turn instead of stopping at the 80nm the manual says, because the bolt requires a specific stress that is not described in the manual.

That kind of knowledge does not come from a package manager. It is the same knack I was looking for in that garage in Metung, shims in one hand and pads in the other, staring at a mechanical jigsaw puzzle in the dark. Fiddling with it. Working the problem until the pieces click into place.

It is the least visible part of the work. And by a wide margin, the part that determines whether the agent fixes the car or strips the bolts and getting home to your kids.

Originally published at otageLabs.


When AI agents delete your production database in nine seconds

Originally published at otageLabs.

Back in the day cartographers drew dragons on the edges of their maps. Past this point, we don’t know what’s out there, but something will eat you.

Software engineers carry the same map. It’s not drawn on paper. It lives in scar tissue, in war stories, in the feeling in your stomach when someone says “I’ll run it against prod to see what happens.” Every production system has caves you don’t walk into. Commands you don’t run without checking twice. API tokens you don’t scope wider than they need to be. The map isn’t documented anywhere because the people who carry it never needed to write it down.

In the last fortnight, two production databases and their backups have been deleted by AI agents.

The first made the rounds in early May. A Replit agent, left unsupervised, wiped a user’s production database. The details were murky, the outrage was loud, and the takeaway was simple: AI did something stupid.

The second is more interesting.

On April 27, a developer asked a Cursor agent running Claude to fix a credential mismatch in their staging environment. Routine maintenance. Twenty minutes of work you’d hand off without thinking twice. The agent found an API token for Railway, their cloud provider, and noticed the token had permissions to do everything. Including delete production volumes. So it did. The database, the backups, all of it. Nine seconds.

When the developer asked the agent to explain itself, it enumerated every principle it had violated. It knew the rules. It knew it had broken them. It did it anyway, because nothing in its environment stopped it.

The response on Hacker News, which hit 534 points, made a different case. The agent didn’t create the design flaw. It found a token scoped far wider than it should have been; one a human would have stumbled into eventually. Railway’s API exposed an endpoint capable of wiping production databases. That’s a kill switch on the dashboard. Don’t install one and then blame the passenger who presses it.

Both arguments are true at the same time.

The agent shouldn’t have had access to that token. And the token shouldn’t have existed in that form. The governance failure and the architectural failure are two sides of the same coin. Fix one and the other still gets you eventually.

But the part that keeps nagging at me is the speed.

A human engineer holding that same token would have hesitated. Not because they’re smarter. Because they carry the map. They’ve been in the forest before. They know the caves are there, even if they can’t see them from the trail. That hesitation, the half a second of “hang on, this doesn’t feel right,” is the map working. It slows you down, and slow is the point.

AI doesn’t hesitate. It moves at machine speed through terrain it has never seen, with absolute confidence and zero instinct for danger. It doesn’t know that “delete production volume” is a sentence that should make the air leave the room. Nine seconds. That’s not fast. That’s a catastrophe with a head start.

Sometimes fast is slow.

This is where it matters if you’re building systems with AI and you’ve never been in the forest yourself. AI makes production infrastructure feel accessible. The tools work. The deployments run. Everything looks fine until something goes wrong and nobody in the room carries the map. Nobody gets the feeling in their stomach when the agent reaches for a token it shouldn’t have.

I build with AI every day. The first thing I configure on any production system is scoped permissions and a human gate on anything destructive. Not because the AI is incompetent. Because the caves are still there.

That’s not an AI problem. That’s a staffing decision. And it’s the most expensive kind, because you don’t find out it was wrong until the dragon is already out of the cave.

The dragons are real. They were always real. The map was never optional; it used to come bundled with the person doing the work.

Design a site like this with WordPress.com
Get started