Lesson 9 of 10 · 9 min

Measuring usage and quality

Track who uses the assistant, for what, and whether its answers are right.

ObjectiveBy the end of this lesson, you will be able to measure an assistant’s usage and quality separately, with the Team screen, a test set and the declined approvals.

An assistant that has gone live is not finished. You need to know whether it is used and whether it answers well. These are two different measures: an assistant can be heavily used and mediocre, or excellent and ignored.

Measuring usage

The “Team” screen, available to administrators, shows usage per member: runs for the day, the week and the month, tokens used and storage. If you are not an administrator, ask for this report once a month. Low usage calls for a conversation, not a sanction: the task may have been poorly chosen.

The Team screen, with each member’s usage over the day, the week and the month.

Measuring quality

  • The test set: run the workshop questions again after every change and compare.
  • Declined approvals: each one points to a proposal that was not right.
  • Corrections: note what users rework by hand.
  • Direct feedback: ask every month for three examples of good and bad answers.
Try it in LearnyaHere are the fifteen questions of our test set and the assistant’s answers before and after the change to its instructions. Compare them, say which ones improved, which ones got worse and why, quoting the relevant passages. Try in Learnya
  1. Keep the test set in a project document, with the expected answer for each question.
  2. After every change to the instructions, run the questions again in a new conversation.
  3. Attach the two sets of answers and ask for the comparison.
  4. Open the “Activity” screen to review the month’s declined approvals.
  5. Record the findings in a dated document shared with the team.

For an assistant that trains people, the most direct measure is retention. Reinforcement campaigns sent to a cohort measure what each person has retained over time. A concept poorly retained by the whole cohort often points to an explanation that needs reworking.