Measuring usage and quality
Track who uses the assistant, for what, and whether its answers are right.
ObjectiveBy the end of this lesson, you will be able to measure an assistant’s usage and quality separately, with the Team screen, a test set and the declined approvals.
An assistant that has gone live is not finished. You need to know whether it is used and whether it answers well. These are two different measures: an assistant can be heavily used and mediocre, or excellent and ignored.
Measuring usage
The “Team” screen, available to administrators, shows usage per member: runs for the day, the week and the month, tokens used and storage. If you are not an administrator, ask for this report once a month. Low usage calls for a conversation, not a sanction: the task may have been poorly chosen.
Measuring quality
- The test set: run the workshop questions again after every change and compare.
- Declined approvals: each one points to a proposal that was not right.
- Corrections: note what users rework by hand.
- Direct feedback: ask every month for three examples of good and bad answers.
- Keep the test set in a project document, with the expected answer for each question.
- After every change to the instructions, run the questions again in a new conversation.
- Attach the two sets of answers and ask for the comparison.
- Open the “Activity” screen to review the month’s declined approvals.
- Record the findings in a dated document shared with the team.
For an assistant that trains people, the most direct measure is retention. Reinforcement campaigns sent to a cohort measure what each person has retained over time. A concept poorly retained by the whole cohort often points to an explanation that needs reworking.