We posted a short thought this month that turned out to be the most useful thing we said all week: using the lowest model that gets the job done feels rewarding. Not because pennies matter to us more than anyone else, but because it usually means the job ran on a local option, which does not add to the grid, and when it does not, it is pennies instead of tens of dollars for tokens.
That is a nice line. It is also a habit, and habits are harder to hold than lines. So here is what sits behind ours: which jobs in a normal business week a small model can carry, where it quietly falls over, and how we decide, task by task, when something earns the expensive one.
Start with what the small lane actually is, because people picture one thing when we say it. There are two shapes of it in our world. The first is a small hosted model, the cheap tier from any of the big providers, where a routine call costs a fraction of a cent. The second is a model with downloadable weights running on a machine we already own, where the marginal cost of the next thousand calls is electricity and nothing leaves the building. Both are cheap. Only the second one is free at the margin, and only the second one keeps a client's documents off somebody else's servers.
Most of the work our clients bring us does not need the top of the line. It is classification, extraction, reformatting, a first pass at a reply, turning a messy email into a row in a spreadsheet, summarizing a document into a fixed shape, writing the glue that moves data from one system to another. We used to send all of it to the largest model available, because that was the one in the toolbar. The change that mattered was not finding a better small model. It was deciding, per job, what the job was worth, and noticing how often the answer was "not much."
What the small lane is good at, and how to tell
The jobs that survive on a small model share a shape. One task per call. One clear right answer or a small set of acceptable ones. Input you can describe in a sentence. Output you can check in a second, because a person or a script will look at it before it goes anywhere.
That is a bigger category than it sounds. If the task is "is this message a billing question or a support question," the small model will get it right nearly every time, and the occasional miss costs you a minute. If the task is "read this fifty page contract and tell me what we are exposed to," you are asking for judgment, and judgment is where the cheap models get expensive, because a wrong answer arrives in the same confident tone as a right one.
Here is the test we actually use, and it takes an afternoon. Take the last fifty real inputs you put through AI, the messy ones, not the demo ones. Run them against your smallest option and against your best one, and keep both sets of answers. Then read them side by side and count three things: how often the small model agreed, how often it was wrong in a way a person would catch, and how often it was wrong in a way nobody would catch. That third number is the only one that decides anything. The first two are speed and supervision. The third is risk.
Where it breaks, and why the break is quiet
Small models fail differently than big ones. A big model that cannot do something will often say so, hedge, or ask. A small model rarely knows the edge of its own competence, so it produces something plausible instead. That is the whole problem, and it shapes how we deploy the small lane.
Three failure shapes show up over and over in our work.
The first is ambiguity. Give a small model a task that depends on context it does not have, and it will pick an interpretation and commit. Ask for a summary of a client note where the client changed their mind halfway through, and you get a summary of the first half with the second half smoothed over.
The second is the middle. Long inputs get worse, not because small models cannot read, but because the details that matter are rarely in the first paragraph, and the cheap tier has less room to hold a long document in mind. If your task depends on the third page of an attachment, that is not a small model job.
The third is anything with more than one requirement at once. "Extract the dates, convert them to our format, and skip anything that is a draft" is three instructions. Small models tend to satisfy two and lose the third, and they lose a different one each time, which is worse than failing outright because it looks like noise instead of a bug.
None of that is a reason to avoid the small lane. It is a reason to build the check before the automation. Our pattern is the same every time. Narrow the task until it has one job. Constrain the output so it has to come back in a fixed shape. Add a step that verifies the shape, then send the exceptions up to the bigger model or to a person. The routing is not the clever part. The check is. A cheap call with a good check behind it beats an expensive call nobody reads.
The part that is not on the invoice
There is a second half to the line we posted, and it is the half we care about more than the money. When the model runs on a machine you own, the work does not add to a data center's load, and it does not sit in anyone's logs. For client work that is often the bigger win. Half the questions we get from businesses about AI are really questions about where their data goes, and the honest answer is that it depends entirely on which lane the job runs in.
We publish a free calculator for exactly this reason, so a company can put a number on its own AI usage instead of guessing. You give it a token count, a message count, or an image batch, and it estimates the energy, carbon and water behind it. No signup, no email. It lives here. We wrote it because we could not answer the question for ourselves without building something. If you are being asked about your company's AI footprint and you have no number, that is the number to start with.
The rule is not "never use the big model"
We want to be careful here, because the argument gets read as saying the expensive models are a waste. They are not. On a genuinely hard job they are the only thing that works, and the contrast is worth being specific about.
A client once asked us a strange question: what could we make with a frontier model, more than one billion tokens, and Unreal Engine. We found out. The result was an Atlanta game called The Battle of ATL, built through 97 playable versions, with real map data, bicycle physics, traffic, timed objectives and a full gameplay loop. That project was a frontier job in every sense: long context, multi-step planning, debugging across an entire production process, and a hundred small judgments with no fixed answer. We wrote up what came out of it here, and the part that matters in this piece is the shape of the work. It looked nothing like the rest of that month.
So our rule is not to avoid the big model. It is to make the big model apply for the job. The default should be the floor. A task climbs only when it has shown you why it deserves to.
How to run this in a normal business week
Write it down. Five or six tasks, named in plain language: draft the reply to a new inquiry, pull line items out of a vendor PDF, summarize the weekly meeting notes, write the first pass of a service page, tag incoming leads by industry.
Mark each one. Fixed answer or judgment call. Anything with a fixed answer starts on the smallest option that can produce it, with one real example per task so the output has something to match.
Keep a folder of failures. One file, real inputs, and the wrong outputs as they happen. If the same kind of input fails twice, that input belongs in the expensive lane or with a person, and now you can say why.
Measure per finished task, not per token. Cost per token is the number vendors advertise and the least useful number you have. A cheap model that retries four times and needs a rewrite is not cheap, and a small model that finishes in one pass on hardware you own is close to free.
Put the boundary in writing. Which jobs run where, who changes the rule, and what gets reviewed. The routing knowledge in one person's head leaves with that person, and then everything drifts back to the top of the line and the bill comes back with it.
The reward
The pennies are real, and after a year of frontier pricing they are not a rounding error. But they are not the point. A small model finishing a job is evidence that you understood the job, that you could describe it, bound it, and check it. That is the part that carries over into everything else you build, whether or not the model underneath it is the cheapest one available in six months.
The bigger tool is always there. Using it less, and on purpose, is the discipline. If you want help sorting which jobs in your business can move down a lane, and what the checks look like when they do, that is the conversation we have with Atlanta businesses all the time. Bring the list.
