Fox & SantiagoTechnology, applied.Call (877) 244-9687

Field note

The Council Method on a shop floor

The Council Method checks AI output by having several models answer separately, challenge each other with sources you can open, and keep only what survives, with a person signing the result. On a shop floor, the same discipline applies to an AI-drafted estimate or schedule: no single answer is trusted because it sounds right, every claim traces back to a record, and a named person approves before anything moves.

An estimator gets an AI-drafted quote for a job. It looks clean. The line items are sensible, the labor hours look about right, and the total lands in a range that feels familiar. She is busy. The customer is waiting. The easy move is to glance at it and send it.

That glance is where a lot of AI in operations quietly goes wrong. Not because the AI is bad, but because an answer that looks right and an answer that is right are different things, and on a busy day it is very hard to tell them apart.

The Council Method is how we deal with that. It comes from the book Let the AI Be Smart by Jason Santiago, and it was built for a harder problem than an estimate. But the discipline it teaches fits a shop floor surprisingly well.

The method in plain terms

Here is the method as it is defined: Several AI models answer the same question separately, so none can copy another. Each then sees the others' answers and must defend or correct its own, backing every claim with a source you can open, because agreement alone is not proof. One model writes a verdict that keeps only what survived, and a person signs it.

There are three ideas in that definition worth pulling out, because they carry over even when the setting changes.

First, independence. The answers are produced separately, so one mistake cannot spread by being copied. If two independent answers agree, that means more than one answer repeated twice.

Second, agreement is not proof. Even when answers agree, each claim still has to be backed by something you can open and check. Confidence does not count. Sounding right does not count.

Third, a person signs. The final result is not the machine's decision. Someone accountable reads what survived and puts their name on it.

Applying it to an estimate

Take the AI-drafted estimate. Here is how the same discipline looks when the question is "what should we charge for this job?"

Every line has to trace to a record. The labor hours should point to past jobs with similar scope. The material prices should point to the current price list or a recent purchase order, not a number the model produced from nowhere. The site conditions should point to the notes from the site visit. If the draft cannot show where a line came from, that line is a guess, and it gets marked as one.

Disagreement is information. If you have two independent ways of arriving at a number, for example a draft from the tool and a quick build-up from the estimator's own template, compare them. Where they agree, move on. Where they disagree, that is exactly where to spend review time. It is often where the job is unusual: an access problem, an old system, a customer request that is not standard.

Only what survived goes to the customer. Lines that were checked against a record stay. Lines that could not be checked are either fixed by a person or removed. The estimate that goes out is smaller in ambition than the first draft and much more trustworthy.

The estimator signs. Her name goes on it, the same as before. AI changed who wrote the first draft. It did not change who is accountable for the price. When the job comes in over, the question "who approved this?" has a clear answer. That matters more than it sounds, and it is the heart of the problem we describe in nobody knows who approved what.

This is also the practical defense against estimates that don't match actual job costs. A draft that ties each line to a real past job gives the estimator something to compare against when the actuals come in.

Applying it to a schedule suggestion

Scheduling is a different kind of problem, but the same discipline holds.

Say a tool suggests how to rearrange tomorrow's board after an emergency call. The suggestion might be excellent. It might also have missed that one tech is not certified for that kind of equipment, or that a customer only allows access before noon, or that the part for the second job has not arrived.

So the check looks like this. Each move in the suggestion should point to the records it relied on: the crew calendar, the certifications list, the customer's access notes, the parts status. Where the tool could not confirm something, it should say so plainly instead of assuming. The dispatcher reviews the moves that matter, especially the ones that change a customer's time, and approves the board. If a customer gets a call saying their appointment moved, it is because a person decided it should.

When the day falls apart anyway, and some days will, there is a clear record of what was suggested, what was checked, and who decided. That is how you learn from a bad day instead of repeating it. Our guide on the schedule falling apart by mid-morning covers the upstream causes.

A practical note on both examples: the check only works if the records exist and can be found. If past job costs live in someone's head, or the certifications list is a binder in the truck, the AI cannot cite them and the reviewer cannot check them. Sometimes the most valuable result of trying this discipline is discovering which records the company needs to keep in one reachable place. That is useful to know even before any AI is involved.

Matching the check to the stakes

Not every piece of AI output deserves the full treatment. A draft internal note about a job does not need several independent answers and a formal sign-off. An estimate going to a customer, a schedule change a customer will feel, or anything that moves money does.

A simple way to decide is to ask what a wrong answer would cost, and who would find out first. If the answer is "a customer, and it would cost us money or trust", use the full discipline. If the answer is "someone in the office, and it would take two minutes to fix", a lighter check is fine.

What we do not accept, at any level of stakes, is an answer trusted only because it sounds right. That is the failure the method exists to prevent, and it is the easiest one to fall into when a crew is busy.

What this changes about building tools

When we build a tool with AI inside it, this way of thinking shapes the design from the start. The tool has to show where each part of its answer came from. It has to mark what it could not confirm. It has to hold anything important until a named person approves. And it has to keep a record of who approved what.

Those features rarely show up in a demo. They are the ones that let a company trust an AI-assisted tool on its worst day, not just its best.

There is also a benefit that has nothing to do with AI. A process that asks where each number came from, and who approved it, makes the whole operation easier to learn from. When a job runs over, you can see which assumption was wrong. When a schedule holds up under pressure, you can see why. That record is what turns a busy company's daily decisions into something it can improve, one workflow at a time.

What to do next

Pick one place where AI-drafted output already touches your customers or your schedule, even informally. Ask three questions about it: where does each claim come from, what happens when it is unsure, and whose name is on the result. If any answer is vague, start there.

To see who we are and how the method's author approaches field work, read who we are. For the broader cost of estimates drifting from reality, start with the guide on estimates that don't match actual job costs.

Questions people ask

Do we need several AI models to use this idea?

Not always. The full method uses several models. The principle underneath it, that an answer must be backed by something you can check and signed by a person, applies to any AI-assisted output, even from one model.

Does this slow everything down?

It adds a review step where the stakes justify it, such as an estimate going to a customer. For low-stakes drafts, a lighter check is enough. The point is to match the check to what a wrong answer would cost.

What counts as a source on a shop floor?

Your own records: a past job with similar scope, the current price list, the crew calendar, the site notes, the purchase order. If the AI cannot point to one, the claim is a guess.

Who should sign off?

The person who would have owned the decision without AI: the estimator for an estimate, the dispatcher or service manager for a schedule. AI changes who drafts, not who is accountable.