· Ivelin Kozarev · Sales Training  · 8 min read

Getting Started with AI Roleplay for Sales Trainers: A Practical Pilot Guide

Build your first AI roleplay pilot with a sample buyer brief, a four-part scoring rubric, a two-week schedule, and a clear way to measure progress.

Build your first AI roleplay pilot with a sample buyer brief, a four-part scoring rubric, a two-week schedule, and a clear way to measure progress.

To get started with AI roleplay, choose one sales skill, build one realistic buyer scenario, and agree on what good performance looks like. Test the AI’s feedback yourself, then run a small pilot with a baseline, coached practice, and a final assessment.

Your first goal is to learn whether the practice helps people use what you teach. A two-week pilot can help you test that before adding it to every client programme.

This guide gives you a starting plan you can adapt to your own methodology. The scenario, schedule, and scoring examples below are illustrative templates, not customer results.

1. Choose one behaviour to improve

“Get better at discovery” is too broad for a first pilot. Pick something you can hear in a conversation:

  • Ask what triggered the buyer’s search before presenting a solution.
  • Explore the business impact of a problem before proposing a demo.
  • Find out what sits behind a price objection before offering a discount.

Ask the client manager for a recent example where this behaviour was missing. Use an anonymised account of the conversation to shape the exercise.

Write your goal in one sentence: “By the end of this pilot, reps should explore the buyer’s problem and its impact before pitching.” That sentence should guide both the scenario and the scorecard.

For the wider picture of where practice fits into your work, see what AI actually does for sales coaches.

Skylar illustration comparing scoring every sales skill with scoring only the skill being practised

Keep the assessment focused on the behaviour you are teaching. This Skylar example uses pitch practice; the same principle applies to discovery.

2. Turn that goal into a buyer brief

A useful AI buyer needs a reason to talk, a problem worth exploring, and rules about what to reveal. A job title and “be difficult” leave too much to chance.

Here is a sample you can adapt. It assumes the seller offers staff scheduling software; replace that context with your client’s product and market.

Part of the briefIllustrative scenario
BuyerOperations director at a 60-person field service company
SituationAgreed to a short discovery call after requesting information about scheduling software
Opening statement”We have a scheduling tool already. I’m not sure replacing it is worth the hassle.”
Known to the repThe company uses a spreadsheet alongside its current system
Reveal when asked about the current processTwo coordinators manually reassign jobs when engineers are absent
Reveal when asked about impactLast month, four jobs needed a second visit because the right engineer was not assigned
ConcernA change could disrupt service during the busy season
Rep’s goalUnderstand the problem, check its impact, and agree on a relevant next step

Add instructions for the AI buyer: answer relevant questions naturally; do not volunteer the full brief at once; do not coach the rep during the call; and do not invent budgets or losses. If asked for a fact outside the brief, say you do not know.

Keep a separate final assessment version. Change the buyer and surface details while keeping the target skill and difficulty similar. This helps you check whether learners can apply the skill beyond a script they remember.

Skylar persona library showing different buyer profiles available for roleplay

A persona gives the buyer a face and a role. Add the situation, concerns, and reveal rules from your brief to make the conversation useful.

3. Build a scoring rubric you would use yourself

Start with a few observable behaviours. Avoid scoring vague traits such as “executive presence” without defining what they mean.

For the sample discovery exercise, use this four-part rubric:

Behaviour0: Missing1: Partial2: Demonstrated
Understands the current processPitches without askingAsks how scheduling works but does not follow upEstablishes the process and explores where it breaks down
Explores impactDoes not ask about consequencesAsks generally whether the issue mattersGets a concrete example of the effect on service, time, or cost
Checks understandingAssumes the problemGives a summary without checking itSummarises the problem and asks the buyer to confirm or correct it
Agrees on a next stepEnds without one or pushes an unrelated demoSuggests a relevant step without agreementAgrees on a step tied to the buyer’s problem, with an owner and timing

The maximum is eight points. Keep the four scores visible: a total can hide the exact skill that needs work.

Ask for a quote or timestamp to support each score, plus one thing to try next. “Good discovery” gives the learner little to work with. “You asked about missed visits, but moved on before exploring their effect” points to a specific change.

If the call ends before a behaviour can be assessed, mark it as unassessed and review it. Do not quietly treat an incomplete attempt as a full assessment.

Skylar illustration showing three buyer scenarios assessed against the same pitch criteria

Change the buyer while keeping the skill criteria consistent. The pitch example above illustrates how to test a skill across different conversations.

4. Test the buyer and feedback before inviting learners

Run three attempts yourself: one strong, one weak, and one mixed. In the weak attempt, deliberately pitch early. In the mixed attempt, ask good questions but skip the next step.

Check whether the AI spots those differences. Read the transcript alongside its feedback. Does it cite something you actually said? Does the buyer reveal the hidden facts too easily? Does it reward using a keyword without showing the skill?

If another trainer is available, score the same attempts independently. Discuss disagreements and tighten the rubric before asking learners to trust it. Repeat an attempt to see whether scoring changes materially without a clear reason.

You are checking whether the tool supports your judgement. A tidy dashboard alone cannot establish that.

Skylar practice conversation interface with a buyer message and controls to retry or start a new conversation

Test the learner experience yourself: hold a conversation, read what was said, and try again with a different approach.

5. Run a two-week pilot with a clear owner

Choose a manageable group, such as six to ten learners, and one client manager who will protect time for practice. This is a suggested starting size, not a statistically validated sample.

WhenActivityWhat you keep
Before launchTest access, microphones, the buyer brief, and scoringFinal scenario and rubric versions
Day 1Explain the purpose; run a baseline before teaching the skillEach learner’s first complete attempt
Days 2–3Teach the skill, show an example, and review a practice attemptOne focus area per learner
Days 4–8Schedule three short practice sessions with feedback and retriesAttempts, scores, and learner questions
Days 9–10Run the assessment variant and a trainer debriefFinal scores and specific next actions

Choose a session length that fits the task. For this single-skill exercise, you could start with a five-minute conversation and five minutes for feedback and a retry. Adjust if that cuts off useful exploration.

Tell learners who can see recordings and scores, how long they are kept, and how they will be used. Keep early practice low stakes. Use fictional customer details unless the client has approved the data and platform arrangements.

Skylar course progression screen showing a technical pain identification practice exercise

Give learners a clear exercise to work on between coached sessions. This Skylar course screen shows how a single skill can become a practice assignment.

6. Measure practice, skill, and real-world use separately

Keep three questions distinct:

  1. Did people take part? Report invited learners, baseline completions, practice attempts, and final completions.
  2. Did the target skill improve in roleplay? Compare matched baseline and final scores using the same rubric. Include the range and individual changes, not just an average.
  3. Did that carry into real calls? Where permitted, have the manager review later calls for the same behaviours.

Review a sample of AI scores yourself, including low scores and unusually large gains. Keep the scenario and rubric versions with the results. Changing the scoring rules halfway through makes a before-and-after comparison harder to interpret.

Be precise in the client report. An illustrative result of “six of eight assessed learners improved on the impact question” would describe observed roleplay performance. It would not prove a revenue increase or show that AI alone caused the change. Teaching, trainer feedback, and repeated practice all contribute.

Skylar scenario skill breakdown comparing two teams across MEDDIC skills on a radar chart

Look at individual skills as well as the total. This example compares teams on a MEDDIC scenario; your pilot should use the behaviours in your own rubric.

7. Decide what to change before you expand

If learners cannot start, fix access and instructions. If they practise but the scores seem wrong, fix the rubric and feedback. If scores rise but real conversations do not change, try a less familiar scenario and more manager coaching.

Expand when the exercise is realistic, the feedback is useful, learners can complete it, and the trainer can explain the results. Add one new skill or cohort at a time so you can see what changed.

If you choose Skylar for the pilot, the Getting Started with Skylar guide for sales trainers covers the product setup and ways to use it in your training business. For the learner’s first session, share the Skylar user onboarding walkthrough.

Back to Blog

Related Posts

View All Posts »