Reduce AI spend
without reducing
what the team can do.
Most teams don't have an AI capability problem. They have expensive defaults. We go through the way your engineers actually work and fix those defaults without pushing them back to worse models or slower workflows.
For teams that moved from flat-rate subscriptions to enterprise or API billing and watched costs jump.
Your team learned on flat-rate plans. Enterprise bills every decision.
On a flat-rate plan, model choice, reasoning level, long context, and repeated work are mostly invisible. On enterprise billing, every one of those decisions has a price. The workflow did not change. The bill did.
The expensive part is usually the defaults.
We measure how the team actually works, then change the settings and habits that are costing money without helping the work.
Model selection
The cheapest model is not always the cheapest way to finish the task. We compare cost per completed task.
Cost per completed taskReasoning effort
Use high reasoning when the work is genuinely hard. Use lower effort for routine execution.
Effort should match the workHarness fit
The same model can use very different amounts of context depending on the tool around it. We test what the team actually uses.
Test the real combinationPrompt-cache discipline
Reusable context should stay reusable. Switching models mid-task often makes the provider process it again.
Keep reusable contextSubagent architecture
Subagents only save money when they get a smaller context and a cheaper model. Otherwise they multiply the bill.
Smaller context, cheaper modelCode is free, tokens aren't
Not every task needs another model call. If code or a command can do it cleanly, use that.
Use the cheapest reliable toolRetention vs. caching
Some retention policies disable caching. We make that cost explicit instead of discovering it on the invoice.
Price the tradeoffLower effort
after the hard
decisions are made.
Use high reasoning for planning and hard decisions. Routine execution runs at lower effort, without switching models mid-session and throwing away cached context. We measure that change first. Then we fix expensive defaults, bloated context, bad delegation, and repeated work.
This has to come from the team's actual work.
First, we run a live session for the whole team. Then we work alongside the person who will own the standard, using the team's actual tools and actual work. We leave them with a written model, effort, and tool matrix for the work we reviewed.
Team session
About 90 minutes with the whole team. We explain what drives cost, answer questions, and use examples from their work.
- About 90 minutes, live
- Their tools and models
- Questions about real work
Workflow review
We spend the better part of a day with the team lead and watch the real process. This is where the expensive habits usually become obvious.
- The real workflow
- Tools, tasks, and habits
- What the work costs
Fix and hand off
We change the defaults, document the decisions, and leave the team lead with something they can actually enforce.
- Team-specific settings
- Written model and tool matrix
- Clear owner after handoff
Need a BAA or your own Azure environment?
Boxwood runs in your Azure account and is built for teams that need a BAA or stricter data handling. We account for the cost and caching tradeoffs instead of pretending they are separate problems.
Start with the bill and a few representative tasks.
We will calculate the cost of the work, identify which defaults are driving it, and show the team lead what to change.