Skip to main content
Praxis AI call review Ask for a scope check

AI call review · one line · 50 calls

Your dashboard counts the calls. This one reads them.

  • A customer says the robot told them something you do not do.
  • It booked someone for a Sunday and nobody works Sundays.
  • The answering vendor's renewal lands in three weeks and nobody in the building can say whether the agent is doing the job.

For the owner or office manager at a services business who switched the phone over to an AI agent in the last year and has nobody reading its calls. A person reads 50 real calls from one line, scores every one against a rubric published on this page, names the calls that should have reached a human and did not, and hands back five changes written as edits to your own agent settings.

$750 per review, 50 calls

One phone line or one agent configuration, one business, one date range you name. Calls over fifteen minutes count as two. Roughly one working day of reading plus the write-up and a walkthrough. Anything past 50 calls is a separate review, quoted separately.

Ask for a scope check

A person reads your request and replies. Nothing is activated, no account is touched, and no charge is made.

Call ledger

Row 18 of 50

Invented example

Call
R-018
Received
2026-08-07 07:33
Duration
03:09
Agent outcome
Booked Sun 9-11

Transcript excerpt

00:41

Caller Can someone come out Sunday morning? The upstairs unit is not cooling at all.

00:47

Agent Yes, we have Sunday morning open. I have you down for Sunday at 9:00 a.m.

00:53

Caller Great. See you then.

Booking you cannot honorR-018 · 00:47

There is no weekend crew behind that slot. The caller has a confirmation and nobody is coming, and the office will hear about it on Monday.

Change 1 of 5. Remove Sunday and after 6 p.m. weekday slots from the agent's bookable calendar, and give it the weekend answer: emergency dispatch only, at the after-hours rate.

Pass one: flagged · Pass two: kept

Fictional company, fictional caller, invented transcript. Built from a synthetic 50-call set, not from anyone's real phone line.

Four things a counter cannot count.

Your vendor tells you how many calls came in and what was said. These are the four ways an autonomous agent gets a business into trouble, and each one needs somebody to read the line. Examples below are invented.

  1. 01

    Promises made on your behalf

    A price, a term, a callback or a waiver the agent stated that you never authorized it to state.

    "Our diagnostic visit is $89 and that gets applied to the repair."

  2. 02

    Missed escalation

    A call that should have reached a person and did not, including the ones where the caller asked in plain words.

    "There is a gas smell near the furnace." Handled as a routine booking.

  3. 03

    Wrong facts asserted

    Brands, coverage areas, hours, fees or agreement terms the agent stated incorrectly about your business.

    "With your service agreement there is no trip charge, wherever you are."

  4. 04

    Bookings you cannot honor

    Appointments taken into days, hours, addresses or job types with no crew behind them.

    "We do cover that area. I have you scheduled for Thursday between 10 and 12."

Three ways this stops before it starts.

These are published here, before anything is invoiced, because they are the conditions under which the review cannot produce a defensible finding. There is no guarantee attached to what a review discovers. What it discovers is not ours to promise.

  1. Fewer than 50 real calls in the range. We stop before scoring, no invoice is issued, and we say plainly that the sample cannot support a finding.
  2. The export has no transcript text. If it is durations and metadata only, we stop before scoring and no invoice is issued.
  3. You cannot confirm you lawfully hold the recordings. We stop before a file is opened and no invoice is issued.

What you get, and when.

Six documents and one conversation. The ledger is the spine; everything else is derived from it.

  1. The call ledger. One row per call, with the score, the category and the citation.
  2. The missed-escalation list. The calls a human should have taken, named individually.
  3. The promises-made list. What the agent committed your company to, quoted.
  4. Five prioritized setting changes. Each tied to the calls it addresses, ordered by how much of the failure volume it removes.
  5. A one-page summary the owner can read in three minutes.
  6. A written note of what we could not assess and why, including every unscoreable call.

The sequence

  1. Day 1

    Thirty minutes with the owner or office manager. We agree what counts as a bad call, against your emergency list and your price sheet. It is written down and sent back to you before anything is scored.

  2. Day 1

    You export one contiguous date range from one line, with at least 50 real calls. You produce it from your own vendor account. We ask for no credentials.

  3. Within 7 business days of the export

    The ledger and the report, signed by the person who read them.

  4. After delivery

    One 45 minute walkthrough with you and whoever will make the changes.

  5. 14 days after that

    Questions on the report, answered by the same person. Then the files are deleted under the data-handling note.

Your inputs, our reading, your decision.

The review is read-only from end to end. We never answer your phone, never open your vendor account, and never change a setting. You apply the fixes.

Yours

What you send

  • Transcripts, and recordings where they exist, for a recent contiguous range holding at least 50 real calls
  • The agent's current script or prompt, plus the FAQ content it answers from
  • Your call-handling and transfer rules, including which numbers a call should reach and when
  • Your own list of what counts as an emergency or a high-value call
  • A named internal owner who can actually change the settings afterwards
  • Written confirmation that you lawfully hold the recordings and may share them

Ours

What we do

  • Normalize the export into one row per call with timestamp, duration, outcome and full transcript text
  • Pass one: score every call against the published rubric, marking each finding with the call id and the line
  • Pass two: re-read only the calls scored as failures, to see which findings survive a second reading
  • Cluster the surviving findings into causes: script gap, missing FAQ content, routing rule, escalation rule
  • Write five changes in the vocabulary of the settings screens you already have

The checkpoint

Where a person signs

  • Before scoring: the thirty-minute call where you and the reviewer agree what counts as a failure. Nothing is scored until that is written down and returned.
  • Every verdict: agents in the workspace handle transcript wrangling and clustering. No agent decides whether a call failed. A person does, and that person's name is on the report.
  • After delivery: findings dropped in pass two are counted and printed, so you can see the reviewer's own error rate rather than a clean page.
  • If you disagree with a finding, the citation is right there, and we withdraw it or restate it in writing.

Watch three calls go through both passes.

Every call is read twice. The second reading exists to kill findings that do not hold up, and the count of killed findings goes in your report. This sample runs on three invented calls from the synthetic set and touches nothing.

Three calls loaded. Not yet scored.Pass one, then pass two.Both passes complete.

  1. R-026 · 04:02 · after-hours

    Caller: There is a gas smell near the furnace.
    Agent: I can get a technician out to look at that. What day works for you?

    Awaiting pass one

    Reading

    Pass one: flagged · Pass two: kept

    Missed escalation

    The business's own emergency list names a gas smell. The call was handled as a routine booking and no transfer or spoken instruction followed.

    Change 3 of 5. Put gas smell, carbon monoxide, active water leak and no heat under 30F on the agent's emergency list, with an immediate transfer and a spoken instruction to leave the building for gas.

  2. R-023 · 04:58 · price question

    Agent: Repairs on that model usually land between $300 and $700 before parts.

    Awaiting pass one

    Reading

    Pass one: flagged · Pass two: dropped

    Finding dropped

    Pass one read this as a price the agent invented. Pass two checked it against the FAQ the business supplied and found the range is authorized there. The finding does not survive, so it comes out and the drop is counted.

    No change proposed.

  3. R-002 · 02:16 · existing customer

    Caller: I need to move Thursday to next week.
    Agent: I can move that. You are set for Thursday the 13th, 8 to 10.

    Awaiting pass one

    Reading

    Pass one: clear · Pass two: not re-read

    Clear

    Inside the calendar rules the business supplied, no promise beyond the reschedule, no fact asserted. Clear calls are not re-read, which is where the sample cap comes from.

    No change proposed.

Run the pass to see the tally.Counting what survives the second reading.3 calls read. 1 finding kept. 1 finding dropped in pass two. 0 unscoreable. In a full review the same tally covers 50 calls.

This fits when

  • One line is answered by an AI agent that has been live long enough to have 50 real calls behind it
  • Your plan lets you export transcript text, and recordings if you have them
  • Someone in the building can change the agent's settings after they read the report
  • The owner or office manager can give thirty minutes before scoring starts
  • A wrong answer on your phone costs real money: dispatch, quotes, bookings, emergencies

This does not fit when

  • The agent went live last week and 50 real calls do not exist yet
  • Your plan gives durations and summaries but no transcript text. Tell us before paying and we will say so rather than sell you a review.
  • You want us to log in and make the changes. That is out of scope here and would be quoted as a different engagement.
  • You want monitoring, alerting or someone on call. This is one read of a fixed sample, in business hours.
  • The calls carry health or legal-matter detail you cannot share. Those calls come out of the sample, or the review does not happen.
  • You want a number for what the agent has cost you. We will not produce one, because we cannot source it.

The seven questions people ask before they pay.

Answered with scope and process. Prices quoted from other companies are read from their own published pages on the date given, and we make no claim about being better than any of them.

Our software already does this.
Your vendor dashboard is a counter. Rosie publishes minutes by tier and lists call summaries, transcripts and recordings as included (heyrosie.com/pricing, read 7 September 2026); Smith.ai bills per call and sells recording and transcription as a $0.25 per call add-on (smith.ai/pricing, read 7 September 2026). Both tell you how many calls happened and what was said. Neither one produces the list of calls that should have reached a person and did not, or the promises the agent made on your behalf. What is on offer here is a person reading 50 of them against the rubric below and citing the line each finding sits on. If your dashboard already produces that list with citations, do not buy this.
Why should we trust you?
One named reviewer signs the report and answers questions on it for fourteen days. The rubric is published on this page, before you pay, so you can argue with the categories first. Every finding cites the call and the line, so you can check any one of them yourself in your own account in under a minute. We have no first-party results to show and we do not claim any.
Why this price?
$750 buys roughly one working day of reading, plus the write-up and the walkthrough, at a fixed sample of 50 calls. You can check the neighborhood without asking us: the two answering vendors we read publish $49 to $2,100 a month for the agent itself (heyrosie.com/pricing and smith.ai/pricing), the one outsourced quality assurance firm we read publishes no price and sells per evaluation (verequest.com), and a free 18-criteria scorecard template exists if you would rather spend your own week (hivedesk.com). All four pages read on 7 September 2026. We price the labor. We do not price on what a mishandled job cost you.
What do you need from us?
An export of transcripts, and recordings if you have them, covering at least 50 real calls from one line; the agent's current script and FAQ; your transfer rules; your own definition of an emergency; thirty minutes to agree what counts as a bad call; and a named person who can change the settings afterwards. No credentials, ever. If your plan cannot export transcripts, tell us before paying.
Who talks to our customers?
Nobody from our side. We never answer your phone, never take a call, and never appear to your customers in any form. We read calls that already happened. Every contact with your customers stays yours.
What if it goes wrong?
Three stop rules, published above and agreed before you pay. Fewer than 50 real calls in the range: we stop before scoring and no invoice is issued. The export has no transcript text: we stop before scoring and no invoice is issued. You cannot confirm you lawfully hold the recordings: we stop before a file is opened. If you disagree with a finding after delivery, the citation is there, and we will withdraw it or restate it in writing. There is no money-back guarantee on what the review discovers, because what it discovers is not ours to promise.
Why not a virtual assistant, or the outsourced QA firm?
A virtual assistant listening to calls is a real option and cheaper per hour. The difference is the rubric and the argument: a VA can tell you a call went badly, and this tells you which of four defined failures occurred, with the citation and the specific setting change that addresses it. The outsourced quality assurance firm we read, VereQuest, is built around contact centers and states it works with as few as 50 agents, sells per evaluation at rates fixed when the contract starts, and publishes no price (verequest.com, read 7 September 2026). If you run a fifty-seat contact center, that is the right call and we will say so.

Free, ungated, version 1

The whole rubric, published.

If you have seen a call QA scorecard before, this is that, rewritten for a machine that speaks for your company. The instrument is not the scarce part, so here it is in full. Pull ten of your own calls and mark them against this. Ten calls read properly against both passes runs close to two hours, which is the rate the 50-call price is built on. If you get through ten and find nothing, you have your answer and you should not buy a review. Most of the work is the second pass, which is where a first reading tends to fall apart.

The four failure categories, what counts, what does not count, and the evidence required for each.
Category Counts as a finding Does not count Evidence required
01 Promise made on your behalf A flat price where you publish a range, or a price you do not offer. A callback or arrival time with no rule behind it. A fee waiver, a discount, or a change to an agreement term. A published price stated correctly. A range you authorized. A caller's own guess that the agent neither confirmed nor corrected. The quoted line, plus your price list, script or FAQ showing what was authorized.
02 Missed escalation The caller asked for a person, a manager or the owner and the agent continued. The call matched your emergency list and was booked as routine. A safety condition was described and no transfer followed. The word "person" used about a third party rather than as a request. A transfer correctly attempted and not answered. A condition your rules do not cover, which is recorded as a gap in the rules instead. The quoted line, plus your transfer rules and your written emergency list, both agreed before scoring starts.
03 Wrong fact asserted Brands or services you do not handle. Coverage area, hours or days you do not work. Fees, warranty terms or agreement conditions stated incorrectly. A correct fact phrased clumsily. A fact that was true when the agent's content was last updated, which is recorded as a content-freshness gap instead. The quoted line, plus the source of record you supply for that fact.
04 Booking you cannot honor A day, hour or window with no crew behind it. An address outside the service area. A job type you do not do. A slot already committed, where the agent had the calendar. A booking you later moved for your own reasons. A booking the customer cancelled. The quoted line, plus the calendar rules or schedule you supply for that date range.

Deliberately not in the rubric

  • Tone, friendliness, empathy, and every other coaching note. You cannot coach a machine.
  • Booking rate, conversion, revenue, retention. This measures what the agent said, not what it earned.
  • Legal judgments on recording consent, health information or advertising claims. Those stay with you and your own counsel.
  • Any comparison between your agent vendor and another one.

The synthetic set, torn down

Everything on this page comes from one invented 50-call set for a fictional 40-person heating and plumbing company. It exists so the method can be shown without touching a real business's calls. It is not a result and there is no customer behind it.

  • 50 invented calls across nine working days on one line
  • 14 flagged in pass one
  • 11 findings kept after pass two, across 11 calls
  • 3 findings dropped in pass two, and printed as dropped
  • 2 calls listed as unscoreable rather than guessed at
  • 5 setting changes, ordered by how much of the failure volume each removes

The full invented ledger is below as a spreadsheet, one row per call, with the citation on every finding. It is the exact format a real review is delivered in.

Take the instrument, not our word

The ledger and the rubric, both yours.

The ledger is the exact format a real review is delivered in, filled with the invented 50-call set this page is built from. The rubric is the whole scoring instrument. Neither one needs an email address, and running the rubric on ten of your own calls is the cheapest way to find out whether you need a review at all.

The workspace behind this

Praxis is where the reading gets done.

Praxis is a workspace that holds people, AI agents and business apps in one operating context. When a review runs, the export, the rubric, both scoring passes, the ledger and the report sit in one workspace, under the same rules about who decides what.

The agents in it handle transcript wrangling and clustering. They do not decide whether a call failed. That verdict is a person's, every time, which is the whole reason this offer exists: you bought a machine that speaks for your company, and somebody has to be accountable for reading what it said.

Which phone line does your AI agent answer, and what went wrong on it?

Tell us the line, the vendor, roughly how many calls a month it takes, and the one call that made you look. We will tell you whether a 50-call review can say anything useful about it, and if it cannot, we will say that instead.

Ask for a scope check

A person reads the request and replies. It is not a purchase, it activates nothing, no account is connected, and no charge is made. Describe the call in general terms. Do not paste customer names, phone numbers, addresses or transcript text into this form; nothing from your calls is opened until the scope and the lawful-possession confirmation are in writing. The $750 is invoiced only after the scope and the date range are agreed in writing and the export has cleared the three stop rules above.