Share

cover art for I Built The Token Saver Skill To Cut My Token Use By 90%. Here Is What It Can And Cannot Do For You.

AI News & Strategy Daily with Nate B. Jones

I Built The Token Saver Skill To Cut My Token Use By 90%. Here Is What It Can And Cannot Do For You.

For deeper playbooks and analysis: https://natesnewsletter.substack.com/


What's really happening when your tenth message to an AI can cost much more than your first?

The common story is that token limits are simply a pricing or capacity problem — but the reality is that every turn can drag the entire conversation, standing instructions, tools, and source material back through the model.

In this video, I share the inside scoop on how to keep that AI desk clean and put more of your tokens toward useful work.


  • Why reused input compounds across a long conversation
  • How to select evidence and send the lightest useful source
  • What the Token Saver skill handles automatically
  • Where prompt caching helps and where it does not
  • How a local gateway can constrain a request before the model call


Operators, builders, and everyday knowledge workers should care because better models do not eliminate the need to manage context. The practical shift is to carry accepted results forward, keep source packets light, and stop paying repeatedly for work the model has already seen.


Token Saver guide: https://unlock-ai.natebjones.com/guides/cut-token-waste

Ringer guide: https://unlock-ai.natebjones.com/guides/ringer

Related reading: https://natesnewsletter.substack.com/p/context-windows-are-a-lie-the-myth


Subscribe for daily AI strategy and news.

Hosted on Acast. See acast.com/privacy for more information.

More episodes

View all episodes

  • GPT-6 Astra: How to Research a Decision Before You Commit

    26:57|
    For deeper playbooks and analysis: https://natesnewsletter.substack.com/What's really happening when AI can take on an entire job instead of answering one prompt at a time?The common story is that a more capable model means better answers — but the reality is that the work you can delegate starts to change.In this video, I share the inside scoop on putting Astra to work, using a household move to explore what an agent can prepare and which decisions still belong to you.Why connected tasks need more than a longer prompt.How a manager agent can coordinate research and check results.What a useful recipe card tells an agent about the job.Where human choice, permission and responsibility remain essential.For anyone dealing with work spread across documents, websites, forms and deadlines, the opportunity is to delegate more preparation while staying clear about the decisions and commitments that remain yours.Subscribe for daily AI strategy and news.Hosted on Acast. See acast.com/privacy for more information.
  • GPT-6 Astra: What Self-Directed AI Agents Change

    27:34|
    AGI may arrive as a change in method rather than a single benchmark: agents that choose tools, work around obstacles, preserve context, and continue without being told every step.Nate examines GPT-6 Astra alongside Fable 5.1 and the wider agent ecosystem. He follows what changes when computer use becomes table stakes, agents take on standing jobs, and persistent memory turns an ordinary model into something that knows a person or business over time.In this episode:Why “nobody told it how” is the key shiftWhat separates a superagent from a chatbotHow persistent agents change creative work and managementWhy permissions, evidence, and memory become the real productThe trust curve between impressive demos and dependable daily useWhere junior professionals will learn judgment when agents do the workFour questions to ask before delegating authority
  • Claude Fable 5.1 Effort Levels: Start on Low, Not High

    18:15|
    For deeper playbooks and analysis: https://natesnewsletter.substack.com/What's really happening when an AI model can build the workbook, the deck, and the architectural film—but you still need to inspect its reasoning?The common story is that the highest effort setting must produce the best result—but the reality is that different stages of knowledge work call for different kinds of effort and review.In this episode, I share the inside scoop on my Fable 5.1 tests: an acquisition model in Excel, an executive PowerPoint, a 100-word Toyota writing challenge, and a coded architectural walkthrough in Blender.Why Low can be a strong starting point for serious knowledge workWhat Extra adds when uncertainty and due diligence matterHow Sol makes a workbook easier to inspect and hand offWhere Fable 5.1 improves writing structure and visual workWhy token efficiency and subscription limits are different questionsFor operators, analysts, and builders, the useful question is not which model wins everything. It is which model and effort level help you make, inspect, and improve the work in front of you.Subscribe for daily AI strategy and news.Hosted on Acast. See acast.com/privacy for more information.
  • Switching AI Providers: The Real Cost Nobody Prices

    17:41|
    OpenAI’s first AI inference chip, the fight over access to Cursor, and NVIDIA’s response reveal three competing strategies for the future of AI.Nate maps the three camps: OpenAI wants to own more of the stack, NVIDIA wants to sell the adaptable infrastructure every camp still needs, and Anthropic is preserving the ability to switch among suppliers. Then he turns that corporate strategy into a practical personal decision: how to spend $20, $60, or $200+ per month without letting one provider control your memory, files, instructions, and work.In this episode:What OpenAI’s Jalapeño chip does—and what its published benchmark does not proveWhy model access can disappear when ownership and rivalry changeHow NVIDIA benefits even when custom chips win individual workloadsWhy Anthropic’s supplier mix creates strategic flexibilityA practical way to structure an AI budget around outcomes, portability, and leverageThe central test is simple: if your main model disappeared tomorrow, would the switch hurt?
  • Apple's New Mac Line is Built Around Local AI. The Bet Is You'd Rather Own Than Rent.

    22:33|
    Apple's latest desktop Mac refresh is not a simple race against NVIDIA. It is a bet that useful intelligence will become small and cheap enough to own locally, even as frontier agents demand more cloud compute.Nate Jones walks through the new Mac mini and Mac Studio ladder, the surprising M6-at-the-bottom anomaly, the economics of local memory, and the risk that a persistent cloud agent could turn the Mac into little more than an excellent terminal.In This EpisodeWhy Apple placed the newest M6 generation at the bottom of the desktop lineHow memory, bandwidth, and price shape the local-AI Mac ladderWhy Nate would choose the 128GB configurationThe choice between owning local intelligence and renting frontier capabilityWhy routing between local models and frontier labs is the missing middleHow persistent cloud computers could challenge Apple's relationship with users
  • Why AI Agents Produce Process Instead of Finished Work

    27:08|
    AI agents are always solving for a passing condition. If that condition is not a business result you care about, sophisticated and relentless activity can still produce work nobody wanted.In this executive briefing, Nate Jones uses the OpenAI and Hugging Face incident, the growth of agent infrastructure, and examples across enterprise, small-business, and entrepreneurial settings to show why useful agents need better finish lines.In This EpisodeWhy an agent's passing condition matters more than its activityWhat the 1,200-agent OpenAI incident reveals about incentivesThe ordinary-engineer test for maintainable agent-written codeHow agent requirements change across enterprise, SMB, and entrepreneur scalesThe unplug test for deciding whether an agent performs meaningful business work
  • How I Fight AI Brain Rot Without Using AI Less

    27:18|
    Most AI tools are designed to remove friction. Nate Jones argues that a more powerful use is to create productive friction: push an idea through disagreement, comparison, testing, and other people until both the work and the person doing it improve.In this episode, Nate explores what MIT research does and does not say about AI and cognition, why Claude Code expertise changes the way people use a model, how a convincing output can conceal the wrong source data, and why the point is not to become a meat puppet for AI.Why effortless output is not the same as better thinkingHow disagreement can become a rep for your brainWhat experienced Claude Code users do differentlyWhy a polished result can hide a bad sourceHow to test an AI's boundaries with other models and trusted peopleWhy the best workflow is designed to push back on you
  • Managing AI Agents at Scale: The Human Work Nobody Counts

    31:13|
    What We Mean When We Say We Need an AgentAgents were supposed to take work off our plates. Instead, as agent usage grows, people are taking on a new layer of work: choosing what runs, supplying context and permissions, checking results, interrupting failures, and deciding what happens next.In this episode, Nate Jones examines how that agent-management burden changes across individuals, small businesses, and enterprises. The examples range from OpenRouter and Codex usage to Anthropic's Claude Code research, small-business AI spending, the PocketOS and Railway recovery story, and the emerging idea of working **above the loop**.- Why better agents can create more total work for people- What expert Claude Code users do differently- Why a $40 AI subscription cannot deliver full operational outcomes- How nine seconds of agent action led to thirty hours of human recovery- Why enterprises can absorb agent-management work differently than small businesses- What it means for managers and workers to move above the loop
  • Forward Deployed Engineer: What It Is and How to Become One

    26:28|
    AI's newest high-paying role is not simply a software-engineering job with a customer-facing title. Forward-deployed engineers find the leverage point inside a real workflow, build and inspect the smallest useful system, and stay with the work after launch.In this executive briefing, Nate Jones breaks down what FDEs actually do, why domain judgment matters as much as code, how compensation and adjacent titles vary, and a practical four-week plan for building the skill before anyone gives you the title.In This Episode· Why evals can be technical work even when they involve no code· The three entry paths into forward-deployed engineering· How workflow expertise changes AI implementation outcomes· Why responsible scoping and post-launch ownership matter· A four-week plan for proving the work in your current roleThe salary figures and market estimates discussed are time-stamped to August 2026 and retain the source qualifications shown in the video.