All Episodes

Tokenomics Will Wreck Your Margins: The AI Model Routing Workflow That Cuts Spend 50%

with Tarun Raisoni · Gruve

July 27, 202601:03:34Redwood City, CA

Tokenomics Will Wreck Your Margins: The AI Model Routing Workflow That Cuts Spend 50%

0:000:00

Show Notes

Tarun Raisoni built a data center company to roughly $400 million in trailing sales without taking a dollar of outside capital, sold it to a Fortune 500 for $217 million, then walked away from the boat and the fishing rod to start over. His answer for why is four words long: “exits are just a number.”

Funding stage: Growth. Gruve has raised $87.5 million from Mayfield, Xora Innovation, and Cisco Investments, employs roughly 500 people across the US, India, Korea, Japan, and Singapore, and built its enterprise customer base in part by acquiring three operating companies outright.

What he started instead is Gruve, and what he came on the show to argue is the most uncomfortable idea in enterprise AI right now. In the cloud era, your data went somewhere else and stayed exactly what it was, because a provider could store it, back it up, and hand it back to you unchanged. Storage does not understand. The AI era broke that: push high quality proprietary data into a frontier model and the data can be parameterized, meaning the intelligence inside it can escape. Tarun's question is not who owns the file. It is who owns the understanding. He calls it the intelligence boundary, and nobody has drawn it yet.

Ryan pushed on the obvious nightmare version: install Claude Code across your team, ship ten times faster, feed in your financials and roadmap, then watch the model provider launch a product into the exact market you were building for. Tarun did not flinch. He reached for Amazon Basics instead, the cleanest available precedent: watch what sells on your platform, build your own version, undercut the seller. Except now the platform is not a marketplace, it is the thing your team talks to all day. His prescription is not “go local and hide.” It is a hybrid routing model, and he is blunt that where you draw the lines depends entirely on your business.

Then the economics arrive, and this is where founders should sit up. Tarun's word is tokenomics, and his point is that the bill you are staring at right now is not a real bill, it is a heavily investor-subsidized one that is being withdrawn on a schedule you do not control. Ryan told the story of a founder friend who worked himself into a hospital bed in January convinced the tools were about to be yanked. Tarun's answer was calmer and more useful: stop worrying about the rug pull and start measuring whether the spend is actually producing measurable productivity.

The go-to-market story is genuinely unusual. Enterprises are notoriously brutal to sell into, so Gruve bought its way to a standing start, acquiring NetServ, Lumos Cloud, and SecurView to import talent, partnerships, and customers in one motion, with Cisco serving as both channel partner and investor. And then the conversation turns sideways: asked what he would do with the abundance AI is supposed to create, Tarun said we are framing abundance wrong. Not AI abundance. Energy abundance. He nearly built a renewable energy company instead of Gruve, spent a year at the Stanford Doerr School preparing for it, then read the political weather and pivoted, still visibly annoyed that the US generates power from only about three percent of its tens of thousands of dams.

Named Frameworks

The Intelligence Boundary

Tarun's term for the undefined line between the data you hand a model and the understanding it derives from it.

  • Your data is one asset. The understanding derived from your data is a different asset.
  • Cloud providers could hold the first without acquiring the second. Frontier models can acquire both.
  • Nobody has legally or contractually defined where one ends and the other begins.
  • Until someone does, the boundary is whatever your architecture decides it is.

The Hybrid Routing Model

Tarun's prescription for where each workload should actually run.

  • Tier one: low IP density, low data gravity. Send it to frontier models and public cloud.
  • Tier two: real proprietary value. Fine-tune a large open source model and run it inside your own ecosystem.
  • Tier three: crown jewels. The frontier model never sees the data, the metadata, or the algorithm.
  • There is no out-of-the-box router. The classification work is the product.

The Tokenomics Test

Why cost per token is the wrong number to watch.

  • Cost per token is the wrong metric. Measurable productivity per dollar is the right one.
  • Match model tier to task value: a cheaper model for UI work, a top-tier model for vulnerability testing.
  • Engineering output is measurable today. Most other functions are not, which is why most AI spend can't be defended.
  • If you can't name the productivity gain, you're buying a subscription, not leverage.

The Investor-Subsidized Bill

Why today's AI pricing isn't the real number.

  • Today's AI pricing includes a venture subsidy that is being withdrawn on a schedule you don't control.
  • Model your unit economics against a future price, not the current one.
  • The real compute bill is still ahead of you.

The Time and Materials Trap

Why paying by the hour stretches the work to fill the meter.

  • Paid by time and materials, work expands to fill the meter. Ryan's roofers work Sunday afternoons when they're paid on completion instead.
  • Applies to how you prompt: constrain output, work from your agenda, cap the back-and-forth.
  • Applies to how you buy services: pay for outcomes, not hours.

The Acquisition-Led Cold Start

How Gruve solved the enterprise cold-start problem.

  • Enterprise sales cycles punish new logos with no history.
  • Acquiring a services business imports talent, partnerships, and existing customers simultaneously.
  • The acquired founders become co-builders when they also want to transform with AI.
  • Relationships, word of mouth, published technical work, and partner-led motion carry it from there.

Twice Is Enough

Tarun's personal rule for risk and mistakes.

  • Every day is a new day, and yesterday is unreachable, so don't spend today on it.
  • Making the same mistake twice is acceptable. A third time is not.
  • Roughly half of what you decide to do will fail. That's the correct rate.
  • Too few failures means too little risk.

Energy Abundance Over AI Abundance

Tarun's reframe of what AI's promised abundance actually depends on.

  • Cheap energy at scale addresses hunger, literacy, and homelessness directly.
  • Cheap energy makes AI cheap as a downstream effect.
  • It's also the better pitch. Nobody fears an abundance of electricity.

Founder Experiment

Run an Intelligence Boundary Audit this week. Pick the three workflows where you've pushed the most proprietary data into a frontier model in the last 90 days. For each, write one sentence naming what a competitor could reconstruct about your business from that data alone: pricing logic, sourcing, customer segmentation. Be specific and be honest. Sort those three into the hybrid tiers: frontier-safe, fine-tune-in-house, never-leaves-the-building. Pull your actual model spend for those workflows and next to each number write the measurable outcome it produced, not “faster,” a number. Re-price every line against a future where the venture subsidy is gone and your bill doubles. Anything that fails that test is a habit, not a strategy. Then move exactly one workflow to a cheaper model tier this week and see whether output quality actually drops. Most founders discover it does not.

Glossary

Frontier model

The largest, most capable general purpose models from the leading labs.

Intelligence boundary

Tarun's term for the undefined line between the data you hand a model and the understanding that model derives from it.

Parameterized

Absorbed into a model's weights, as distinct from stored as a retrievable file.

Tokenomics

The unit economics of AI usage, measuring cost per token against productivity produced.

Outcome-based services

Pricing engagements on a defined result rather than billable hours or headcount.

Hybrid AI

Deliberately splitting workloads across frontier models, self-hosted open source models, and fully sealed environments.

Fine-tuning

Additional training on a base model using your own data to specialize its behavior.

Data gravity

The tendency of applications and workloads to cluster around large data stores.

ERP

Enterprise resource planning software, where sourcing, pricing, and supply chain IP tends to live.

Data governance

The policies and controls determining who can access which data, and how quality is maintained.

Red teaming

Adversarial testing that attacks your own systems to find vulnerabilities before someone else does.

CISO

Chief Information Security Officer.

Inference

Running a trained model in production to generate outputs, as opposed to training it.

Neocloud

An AI company with surplus compute that resells capacity to peers, including direct competitors.

Bookings vs. run rate vs. ARR

Bookings are contracts signed, run rate annualizes a recent period, ARR is recurring subscription revenue. They are not interchangeable.

EBITDA multiple

Acquisition price expressed as a ratio of earnings. Rahi sold at roughly 7.5x.

AI factory

Framing a data center as a facility that manufactures intelligence rather than storing data.

Agentic

AI systems that take multi-step actions toward a goal rather than returning a single response.

Q&A: What Founders Ask After This Episode

Who owns the intelligence derived from enterprise data fed into an AI model?

It is legally and contractually undefined. In the cloud era, providers stored data without inferring intelligence from it, so ownership was clean. Frontier models can parameterize data, meaning the derived understanding may escape even when the raw data is protected. Tarun Raisoni argues enterprises will treat owning that intelligence as their primary competitive moat over the next three to five years.

What is a hybrid AI strategy and why do enterprises need one?

Splitting workloads by IP density. Low-value, low-gravity tasks go to frontier models and public cloud. Genuinely proprietary work runs on a fine-tuned large open source model inside your own environment. Crown jewel workloads never expose data, metadata, or algorithms to an external model. There is no default router, so the classification work is the actual project.

Can frontier AI models become competitors to their own enterprise customers?

In enterprise contexts, yes, and the precedent is Amazon Basics: observe what sells, build your own version, undercut the seller. Tarun's caveat is that this risk is concentrated where deep operational IP lives, such as ERP data covering sourcing, pricing, and logistics. For internal learning and development, frontier models are largely beneficial.

Are AI costs going to increase for businesses?

Yes. Current pricing reflects a substantial investor subsidy that has been declining over time and will continue to decline. The bill enterprises see today is not the real cost of the compute they consume. Founders should model unit economics against future pricing rather than today's.

How do you sell enterprise software or services as a brand new company?

Gruve's approach was acquisition-led: buy services businesses to import talent, partnerships, and existing customers at once. From there, relationships, customer word of mouth, published technical application notes, and partner ecosystems carry the motion. Enterprise decision cycles have also compressed because C-suites now use AI directly for strategy instead of hiring consultants first.

Should companies use frontier models for security testing and red teaming?

Tarun recommends against pointing a frontier model directly at your environment. Use purpose-built security tools that sit on top of those models and provide guardrails, so you get the findings without your information flowing into the model. He considers AI-assisted security the area where applied AI will scale first.

Will AI eliminate jobs or transform them?

His view is transformation, not elimination. Gruve has continued adding headcount while the roles themselves changed. He points to rising demand for electricians, cooling technicians, and construction trades driven directly by AI infrastructure buildout, plus entirely new categories for managing AI systems. He also expects India's technology services job count peaked in 2026, because outsourcing to manage cost stops making sense when AI does the cost reduction.

What does an outcome-based services model look like in practice?

Clients pay for a defined result rather than consulting hours or headcount. Gruve started with security because outcomes there are unusually well defined. Delivering with AI agents rather than human hours is what makes the margin structure work, and Mayfield has publicly cited software-like gross margins in the seventy to eighty percent range as the reason venture capital is now willing to fund a services company.

How much data do enterprises need to clean before deploying AI?

Tarun's rough field number is that seventy percent of enterprise data is usable and thirty percent lacks proper governance or quality. That thirty percent has to be addressed before AI deployment, alongside establishing governance practices, or the outputs cannot be trusted.

Why does energy abundance matter more than AI abundance?

Cheap energy at scale addresses hunger, literacy, and homelessness directly, and makes AI inexpensive as a consequence. Tarun nearly founded a renewable energy company instead of Gruve and studied at the Stanford Doerr School before pivoting, citing dependence on government support. He notes China added more than 430 GW of wind and solar in 2025, roughly a third of total US installed capacity, and that the US generates power from only about three percent of its dams.

URLs Mentioned in the Episode

Links & Resources