· Ivelin Kozarev · Sales Training · 8 min read
Getting Started with AI Roleplay for Sales Trainers: A Practical Pilot Guide
Build your first AI roleplay pilot with a sample buyer brief, a four-part scoring rubric, a two-week schedule, and a clear way to measure progress.

To get started with AI roleplay, choose one sales skill, build one realistic buyer scenario, and agree on what good performance looks like. Test the AI’s feedback yourself, then run a small pilot with a baseline, coached practice, and a final assessment.
Your first goal is to learn whether the practice helps people use what you teach. A two-week pilot can help you test that before adding it to every client programme.
This guide gives you a starting plan you can adapt to your own methodology. The scenario, schedule, and scoring examples below are illustrative templates, not customer results.
1. Choose one behaviour to improve
“Get better at discovery” is too broad for a first pilot. Pick something you can hear in a conversation:
- Ask what triggered the buyer’s search before presenting a solution.
- Explore the business impact of a problem before proposing a demo.
- Find out what sits behind a price objection before offering a discount.
Ask the client manager for a recent example where this behaviour was missing. Use an anonymised account of the conversation to shape the exercise.
Write your goal in one sentence: “By the end of this pilot, reps should explore the buyer’s problem and its impact before pitching.” That sentence should guide both the scenario and the scorecard.
For the wider picture of where practice fits into your work, see what AI actually does for sales coaches.

Keep the assessment focused on the behaviour you are teaching. This Skylar example uses pitch practice; the same principle applies to discovery.
2. Turn that goal into a buyer brief
A useful AI buyer needs a reason to talk, a problem worth exploring, and rules about what to reveal. A job title and “be difficult” leave too much to chance.
Here is a sample you can adapt. It assumes the seller offers staff scheduling software; replace that context with your client’s product and market.
| Part of the brief | Illustrative scenario |
|---|---|
| Buyer | Operations director at a 60-person field service company |
| Situation | Agreed to a short discovery call after requesting information about scheduling software |
| Opening statement | ”We have a scheduling tool already. I’m not sure replacing it is worth the hassle.” |
| Known to the rep | The company uses a spreadsheet alongside its current system |
| Reveal when asked about the current process | Two coordinators manually reassign jobs when engineers are absent |
| Reveal when asked about impact | Last month, four jobs needed a second visit because the right engineer was not assigned |
| Concern | A change could disrupt service during the busy season |
| Rep’s goal | Understand the problem, check its impact, and agree on a relevant next step |
Add instructions for the AI buyer: answer relevant questions naturally; do not volunteer the full brief at once; do not coach the rep during the call; and do not invent budgets or losses. If asked for a fact outside the brief, say you do not know.
Keep a separate final assessment version. Change the buyer and surface details while keeping the target skill and difficulty similar. This helps you check whether learners can apply the skill beyond a script they remember.

A persona gives the buyer a face and a role. Add the situation, concerns, and reveal rules from your brief to make the conversation useful.
3. Build a scoring rubric you would use yourself
Start with a few observable behaviours. Avoid scoring vague traits such as “executive presence” without defining what they mean.
For the sample discovery exercise, use this four-part rubric:
| Behaviour | 0: Missing | 1: Partial | 2: Demonstrated |
|---|---|---|---|
| Understands the current process | Pitches without asking | Asks how scheduling works but does not follow up | Establishes the process and explores where it breaks down |
| Explores impact | Does not ask about consequences | Asks generally whether the issue matters | Gets a concrete example of the effect on service, time, or cost |
| Checks understanding | Assumes the problem | Gives a summary without checking it | Summarises the problem and asks the buyer to confirm or correct it |
| Agrees on a next step | Ends without one or pushes an unrelated demo | Suggests a relevant step without agreement | Agrees on a step tied to the buyer’s problem, with an owner and timing |
The maximum is eight points. Keep the four scores visible: a total can hide the exact skill that needs work.
Ask for a quote or timestamp to support each score, plus one thing to try next. “Good discovery” gives the learner little to work with. “You asked about missed visits, but moved on before exploring their effect” points to a specific change.
If the call ends before a behaviour can be assessed, mark it as unassessed and review it. Do not quietly treat an incomplete attempt as a full assessment.

Change the buyer while keeping the skill criteria consistent. The pitch example above illustrates how to test a skill across different conversations.
4. Test the buyer and feedback before inviting learners
Run three attempts yourself: one strong, one weak, and one mixed. In the weak attempt, deliberately pitch early. In the mixed attempt, ask good questions but skip the next step.
Check whether the AI spots those differences. Read the transcript alongside its feedback. Does it cite something you actually said? Does the buyer reveal the hidden facts too easily? Does it reward using a keyword without showing the skill?
If another trainer is available, score the same attempts independently. Discuss disagreements and tighten the rubric before asking learners to trust it. Repeat an attempt to see whether scoring changes materially without a clear reason.
You are checking whether the tool supports your judgement. A tidy dashboard alone cannot establish that.

Test the learner experience yourself: hold a conversation, read what was said, and try again with a different approach.
5. Run a two-week pilot with a clear owner
Choose a manageable group, such as six to ten learners, and one client manager who will protect time for practice. This is a suggested starting size, not a statistically validated sample.
| When | Activity | What you keep |
|---|---|---|
| Before launch | Test access, microphones, the buyer brief, and scoring | Final scenario and rubric versions |
| Day 1 | Explain the purpose; run a baseline before teaching the skill | Each learner’s first complete attempt |
| Days 2–3 | Teach the skill, show an example, and review a practice attempt | One focus area per learner |
| Days 4–8 | Schedule three short practice sessions with feedback and retries | Attempts, scores, and learner questions |
| Days 9–10 | Run the assessment variant and a trainer debrief | Final scores and specific next actions |
Choose a session length that fits the task. For this single-skill exercise, you could start with a five-minute conversation and five minutes for feedback and a retry. Adjust if that cuts off useful exploration.
Tell learners who can see recordings and scores, how long they are kept, and how they will be used. Keep early practice low stakes. Use fictional customer details unless the client has approved the data and platform arrangements.

Give learners a clear exercise to work on between coached sessions. This Skylar course screen shows how a single skill can become a practice assignment.
6. Measure practice, skill, and real-world use separately
Keep three questions distinct:
- Did people take part? Report invited learners, baseline completions, practice attempts, and final completions.
- Did the target skill improve in roleplay? Compare matched baseline and final scores using the same rubric. Include the range and individual changes, not just an average.
- Did that carry into real calls? Where permitted, have the manager review later calls for the same behaviours.
Review a sample of AI scores yourself, including low scores and unusually large gains. Keep the scenario and rubric versions with the results. Changing the scoring rules halfway through makes a before-and-after comparison harder to interpret.
Be precise in the client report. An illustrative result of “six of eight assessed learners improved on the impact question” would describe observed roleplay performance. It would not prove a revenue increase or show that AI alone caused the change. Teaching, trainer feedback, and repeated practice all contribute.

Look at individual skills as well as the total. This example compares teams on a MEDDIC scenario; your pilot should use the behaviours in your own rubric.
7. Decide what to change before you expand
If learners cannot start, fix access and instructions. If they practise but the scores seem wrong, fix the rubric and feedback. If scores rise but real conversations do not change, try a less familiar scenario and more manager coaching.
Expand when the exercise is realistic, the feedback is useful, learners can complete it, and the trainer can explain the results. Add one new skill or cohort at a time so you can see what changed.
If you choose Skylar for the pilot, the Getting Started with Skylar guide for sales trainers covers the product setup and ways to use it in your training business. For the learner’s first session, share the Skylar user onboarding walkthrough.


