Case study · Home services

Built, dark by design

A roofing receptionist that answers at 2am and will never quote a price

A&J Professional Services is a family-owned roofing and exterior contractor. When the owners are on a roof, the phone goes to voicemail, and the homeowner calls the next contractor. We built a voice receptionist that answers, qualifies the job the way their estimator would, and books the inspection. It is backed by 132 passing tests and a written call-quality rubric.

Client

A&J Professional Services

Headline result

Deployed, 132 tests green

Built, tested, and deployed / dark until the owners flip it on

The system is built, tested, and deployed to production, and it is switched off until the owners are ready to go live. We are not going to tell you how many calls it has answered, because the honest answer today is that the number is not the point yet. What we can show you is what it was proven against before it was allowed near a homeowner.

132

Tests passing before it could ship

10

Labeled golden calls the judge must match

0

Prices it is allowed to quote

About the company

A&J Professional Services is a family-owned roofing and exterior contractor that has been working since 2004. They are a GAF Master Elite contractor, which is the manufacturer's top certification and what gates the enhanced warranty tiers a homeowner cares about. They run on JobNimbus, which is the system of record for most of this trade.

It is a real operation with real crews, which is exactly why the phone is a problem. The people best equipped to answer a homeowner's questions are the people on a roof.

The problem

A homeowner with a leak calls three contractors. The one who picks up gets the inspection. This is not a subtle market dynamic, it is the whole thing. Every call that goes to voicemail after five is a lead being handed to whoever answers next.

But the fix is not just answering. The first question a homeowner asks is what it costs, and the honest answer is that nobody can price a roof from a phone call. Any ballpark given becomes the number you are held to, and it gets given by the person with the least information about the property. So a phone system that helpfully offers a range has not solved the problem, it has created a new one.

There was a third thing we found once we started building, and it turned out to be the most important. The FAQ document the business already had was actively sabotaging the phones. It pushed toward collecting an email address and hard-booking the appointment on every call, which is telemarketer behavior, and homeowners hang up on it. The receptionist's real job is warm capture toward a booked inspection. We had to gut that document to get there.

The document we were given told the system to always ask for an email and always push for the booking. That is how you lose the call.

Paraphrased, not a verbatim quote

The solution

We built Sarah, a voice receptionist wired into JobNimbus. She answers at any hour, understands what the homeowner is describing, checks the service area, asks the qualification questions the estimator would ask in the order he would ask them, and books the inspection. The call record and the appointment land in JobNimbus, so nobody rekeys anything and there is no second system to check.

The trade knowledge is real, and that matters more than it sounds. She knows the difference between a manufacturer certification and a workmanship warranty. She knows what a square is, which is 100 square feet of roof. She knows that the enhanced GAF warranties require a qualifying full GAF system, which is why brands do not get mixed on those jobs. A system that gets any of that wrong says something embarrassing to a homeowner who knows better.

The hard rules are built in as constraints rather than instructions. She does not quote a price. Not a range, not a per-square number, not a most-jobs-run-about. She does not ask for an email. She does not invent a fact or a commitment. When something is outside her boundary, an angry customer, an insurance adjuster, a commercial job, she stops and hands the call to a person with the context attached.

Then there is the part that actually decides whether any of this ships: the grading. Every call is scored against a written rubric in two layers. A deterministic layer checks the hard rules first, and the caps mirror the trade's real risks: quoting any price caps the score at 50, asking for an email caps it at 50, fabricating a fact or a commitment caps it at 45. Then a model judges the softer parts against the same rubric a human would use, and it has to band-match a human grader on 9 of 10 labeled golden calls before we trust its judgment at all.

The impact

The system is deployed and it is dark, which is a deliberate state. It is switched off until the owners are ready, and that is the correct order of operations: build it, prove it, then let the business decide when it starts answering their phone. We are not going to publish a call volume number, because the honest one right now is zero, and a number we made up would be worth less than nothing.

What we can point at is what it took to get here: 132 passing tests, a 52 out of 52 migration slice when we moved the voice layer, and 21 tests that exist purely to prove the safety guardrails hold. Ten labeled golden calls with a written rubric, and a model judge that has to agree with a human grader before its scores mean anything.

The reason that matters is simple. This system talks to homeowners, in your name, when nobody is watching. The question is not whether the demo sounded good. The question is what happens on call four hundred, at 2am, when somebody asks a question nobody anticipated. The tests are the only honest answer to that, and they are why the owners get to decide when it goes live instead of finding out the hard way.

The demo always sounds good. The question is what it says on call four hundred at 2am, and the only honest answer to that is a test.

Paraphrased, not a verbatim quote

Under the hood

How the system actually runs.

The call comes in, the system understands what the homeowner is describing, checks it against what it is allowed to do, and either books an inspection or hands the call to a person. The interesting part is not the conversation. It is the two layers of grading that decide whether a build is allowed to reach a real homeowner at all.

  1. What goes in

    • An inbound call, at any hour, from a homeowner who is usually calling more than one contractor
    • The trade knowledge base: shingle lines, warranty tiers, financing, service area, what an inspection involves
    • The hard rules, written down before anything was built: no prices, no email requests, no invented facts
  2. What the system does

    1. 01

      Understand the job

      What is actually wrong, what kind of property, whether the caller can authorize work, and whether it is in the service area.

    2. 02

      Qualify like the estimator would

      Replacement, repair, or an insurance claim, asked in the order a person who does this for a living would ask it.

    3. 03

      Earn the inspection

      Warm capture toward a booked visit. Not a hard booking script, because homeowners hang up on those and it costs you the lead.

    4. 04

      Grade the call

      A deterministic layer checks the hard rules first, then a model judges the rest against the same written rubric a human grader uses.

  3. Human checkpoint

    Anything outside what it is built for stops and routes to a person with the context attached: an angry customer, an adjuster, a commercial job. It never improvises past its boundary.

  4. What comes out

    A qualified inspection on the calendar and a call record in JobNimbus, or a routed call with the details already captured. No price was quoted. That is enforced, not hoped for.

Results

The numbers, plainly.

  • 132 tests passing before the system was allowed to reach production.
  • 52 of 52 on the migration slice when the voice layer was rebuilt, so the rebuild changed the platform and nothing else.
  • 21 tests exist for the safety guardrails alone, because those are the ones that matter at 2am.
  • 10 labeled golden calls, with a model judge required to band-match a human grader on 9 of them before its scores are trusted.
  • 3 hard caps enforce the trade's real risks: quoting a price caps at 50, asking for an email caps at 50, fabricating a fact caps at 45.
  • 0 prices the system is allowed to quote over the phone, enforced by a test that fails the build.
  • 24/7 coverage, deployed and waiting on the owners' go-live. We will publish call volume when there is a real number to publish.

More work

Other systems we built.

Read next

The thinking behind the work.

Get started

Talk through your call flow

Book a 30 minute call. We find the one place your work waits on you, and you get a straight answer about whether AI is the fix. No pitch, no deck.