The Estimate Is a Claim. Build It Like One.

Build notes on the pipeline behind our service jobs: transcript to catalog-grounded estimate, code-verified acceptance with a hard ceiling, calendar handoff, vault-backed records, and an invoice page — in about ten minutes of human time.

The Estimate Is a Claim. Build It Like One.

An estimate is a claim about the future: this is what the job is, and this is what it will cost. An invoice is a claim about the past. Most small-business tooling treats both as documents to generate. I've come to think of them as claims to trace, and that changes how the whole pipeline is built.

Here is the one I run on our service jobs now, built end to end.

One real job, start to finish: the estimate, the acceptance, the service record, the site record and the invoice page. Customer details redacted.

Intake: three doors, one record

A tech can upload a call transcript, forward an email thread, or dictate the scope. All three end up as text, and that text is the record. Everything downstream has to trace back to it.

Extraction and pricing are separate passes

The first model pass does one thing: it turns the conversation into strict JSON — summary, line items, assumptions, out-of-scope, client dependencies. It never returns prose. The catalog is injected into the prompt, and each line item has to carry a real catalog ID or null. That's the difference between a model that estimates and a model that makes up a price list.

The second pass takes that clean structure and refines it against the catalog. Keeping extraction and pricing apart does two things: extraction isn't distracted by arithmetic, and pricing always works from structure, never from free text.

The pricing itself sits on a stored set of factors, not model judgment:

  • catalog prices, with anything unverified in 60 days flagged for a re-check
  • labor roles and bill rates, with mobilization on every job
  • owner oversight as a tier, from review-only to hands-on
  • contingency for technical unknowns
  • new construction versus retrofit, with the model told explicitly to default to retrofit when unsure, because that single flag moves labor hours a lot and a wrong "true" is the expensive mistake
  • pass-through materials at cost plus tax, shipping and handling
  • a locked set of core assumptions, so every estimate states them the same way
  • the quote type: fixed, cost-plus, or time and materials

Then comes the first check: every line has to be supported by something in the conversation. Anything that isn't gets flagged before a person approves the draft.

Acceptance is identity plus a fingerprint, with a ceiling

The approved estimate is published as an expiring page. The token is the access control, and every open texts me, with link-preview bots and my own ?preview loads filtered out.

Acceptance works like this:

  • The customer confirms with a six-digit code sent to their own phone, with email as the fallback. It expires in ten minutes, allows five attempts, and only its hash is stored.
  • The page is fingerprinted. If it changed after they opened it, accept is refused until they reload.
  • The full snapshot of the page they accepted is stored permanently, and a PDF goes to both sides.

And it has a hard ceiling: every option must be under $2,000. That's a design decision, not a limitation I'm apologizing for. A texted code is proportionate to a small time-and-materials job. For larger or more contentious work, I'm piloting a heavier version with full PandaDoc integration and real digital signatures.

Accepting hands straight off to a calendar, so the customer books the work in the same motion.

Close-out: dictation in, three documents out

After the work, the tech talks for a few minutes: hours, parts, anything unexpected. On a real job, under four minutes of dictation produced the full summary and invoice. From that come three things:

  • A gated record. Cloudflare Access with a one-time code to the customer's email: every device, every setting, and next steps.
  • Passwords that are never on the page. They live in Bitwarden. The page's own read-only key can open only that site's entries. Reveal validates the Access token and then fetches one secret at that moment. It's never cached, never in the HTML, re-masked after two minutes, and every reveal is logged.
  • An invoice page. Line items come from the Stripe products and services catalog. Stripe never emails the customer; the customer gets an expiring branded page with the charges, a PDF download, a link to the record, and Stripe's payment element embedded. I'm texted when the payment is confirmed.

The metric

The whole thing exists to move one number: human minutes per job spent on anything other than the work. The target is about ten. Scoping to a priced estimate timed at about nine minutes on a real job. The first job run end to end came in around eleven and a half, with part of the flow still unbuilt at the time. It's built now, and I expect it to come in under ten. I'll call it under ten when a timed job says so, not before.

That last sentence is the same principle as the rest of the pipeline. Every figure should trace to something that actually happened.