A quality score is useful only when reviewers interpret its questions consistently. Begin by having reviewers independently evaluate the same evidence with the same rubric, then resolve disagreements using explicit decision rules.
This method is for QA managers, client reviewers and regional-language reviewers. It can be run in a controlled worksheet; it does not require or claim a built-in HeyRik calibration feature.
Choose evidence that challenges the rubric
Select authorised samples that include a normal outcome, a correction, an unknown answer, a failed handoff and an ambiguous result. Include the recording or transcript and the downstream evidence needed for the task. Limit access to the reviewers who need it.
For language accuracy, involve a reviewer fluent in the relevant language and business terms. Do not infer poor quality from an accent or treat an English transcript as sufficient proof that a Telugu exchange was understood.
Write observable questions
- Fact accuracy: did the answer agree with the approved source available at the time?
- Field accuracy: did the captured amount, name or date match the caller’s final confirmed value?
- Correction handling: after an interruption, did the final record retain the correction?
- Unknown-answer handling: did the agent acknowledge missing information and use the configured fallback?
- Next action: does evidence show the promised request reached its destination?
- Caller preference: was a declined next step represented accurately in the record?
Run an independent first pass
- 1Freeze a rubric version and sample set. Include an example pass, fail and unverified decision for each question.
- 2Have each reviewer score the same samples without seeing the others’ answers. Retain a short evidence reference for each decision.
- 3Compare decisions question by question. Separate missing evidence from different interpretations of the rule.
- 4Ask the QA owner or relevant specialist to explain a resolution from the source evidence. Seniority alone is not evidence.
- 5Revise unclear wording and re-score the disputed examples under the new version. Preserve the original scores.
- 6Use a fresh sample set for the next check, so agreement is not simply recall of the previous discussion.
Worked disagreement: a spoken callback promise
Illustrative example: reviewer A passes a call because the agent says a salesperson will call back. Reviewer B marks it unverified because no destination record is available. The rubric asks whether the callback request was received, so the spoken promise alone does not satisfy it.
The revised question names the receiving record and required fields. The owners investigate the missing record rather than changing the score to make the pilot look better. This scenario is fictional and contains no customer results.
Report agreement with its limits
A simple starting measure is matching decisions divided by decisions both reviewers actually scored. Report the sample size, how it was selected and how unverified responses were handled. Agreement does not establish correctness: two reviewers can make the same mistake.
Amazon Connect’s calibration documentation describes comparing evaluations of the same contact and optionally using a designated expert. The operating steps here are an independent worksheet-based adaptation. They do not establish that HeyRik offers the Amazon feature or that this process will produce a particular accuracy improvement.
Make feedback useful to the implementer
Finish with a short change record: failed criterion, evidence reference, likely cause, assigned owner and retest case. A missing business fact may need a knowledge update; an incorrect amount may need a better readback; a missing destination record may need an integration repair. Changing the entire prompt for every low score makes it harder to identify what helped.
Recheck the affected cases after the change and sample other cases for regressions. Keep model, voice, integration and rubric changes visible when comparing results across weeks.
A little more before you begin.
Is reviewer agreement the same as agent accuracy?
No. Agreement measures consistency between reviewers. Accuracy requires a defensible reference and evidence that the decision was correct.
Does HeyRik include this calibration workflow automatically?
This article describes a proposed human review process. Use your authorised review tools and confirm the capabilities available in your configured account.
Start here
Discuss a reviewable calling workflow
Build a voice agent, upload your list or connect your leads, and run your first campaign on HeyRik — free to start, no credit card required.