Article
AI Avatars for Training: Building Role-Plays With Scoring, State and Feedback
Learn how to design AI avatar training simulations with explicit state, meaningful scoring, multi-character scenarios and feedback tied to what the learner actually did.

AI avatars can turn training from something employees watch into something they practise. The most useful systems go beyond an animated character asking questions. They create a role-play with explicit objectives, state that changes during the conversation, measurable behaviours and feedback based on what the learner actually did.
That makes them particularly useful for skills that are hard to learn from slides alone: sales, customer service, management, interviewing, negotiation, language learning and other situations where performance depends on what somebody says and how another person responds.
Why AI avatars are useful for training
Many workplace skills are social. You cannot fully learn them by reading what a good conversation looks like.
A salesperson needs to practise handling an objection. A manager needs to experience the discomfort of giving difficult feedback. A receptionist needs to stay calm while a guest becomes frustrated. A candidate benefits from answering a follow-up question they did not know was coming.
Traditional role-play can do this well, but it is difficult to provide consistently at scale. Learners need another person, instructors need time, different partners behave differently and some people are uncomfortable practising sensitive scenarios with colleagues.
AI makes another model possible: a learner can practise with a character whenever they want, repeat the scenario and receive structured feedback immediately afterwards.
What the research says
The research base is still developing, and AI role-play should not be treated as a proven replacement for human instruction in every domain. But recent results are encouraging.
A 2025 pilot randomized trial in medical education compared an AI chatbot simulation with peer role-play for clinical-exam preparation. The differences in overall performance were not statistically significant, but learners using the AI system particularly valued the autonomy, repeatability and structured feedback, while peer role-play was valued for authenticity and exam-like interaction. Read the BMC Medical Education study.
A larger 2026 randomized controlled trial involving 346 participants training for child-helpline conversations found that simulation improved task knowledge, while adding guidance feedback produced additional gains in knowledge and conversational performance. Read the study in the International Journal of Human-Computer Studies.
The practical lesson is not that AI should replace trainers. It is that repeatable simulation plus well-designed feedback can become a useful additional practice layer.
A talking avatar is not yet a training simulation
Suppose you create a virtual customer and prompt it:
“You are an angry customer. Have a difficult conversation with the learner.”
You now have a role-playing chatbot. You do not yet have a robust training experience.
For training, you also need to define:
- what the learner is supposed to practise;
- what the character wants;
- what information the learner must discover;
- which behaviours matter;
- how the character should react to those behaviours;
- what should change as the conversation progresses;
- what counts as a successful or unsuccessful outcome; and
- what feedback the learner should receive afterwards.
Those are the pieces that turn conversation into simulation.
The most important concept: state
A training scenario should not behave the same way regardless of what the learner does.
Imagine a customer-service simulation. The customer begins frustrated. If the learner listens, acknowledges the problem and finds a workable solution, the customer's trust should improve. If the learner becomes defensive or blames another department, frustration should rise.
That change can be represented as explicit experience state:
trustfrustrationneeds_discoveredownership_shownresolution_reached
The character can react to those variables during the role-play, and the feedback system can use them afterwards.
This is much more useful than treating the entire interaction as an unstructured transcript that is only evaluated at the end.
State and scoring are not the same thing
This distinction is worth making explicitly.
State represents what is currently true inside the scenario. It can affect what happens next.
Scoring evaluates the learner's performance.
Sometimes the same variable can serve both purposes, but they do not have to.
For example, customer_trust may be state because it changes the customer's behaviour.
A separate empathy_score might contribute to the final assessment but never be visible
to the customer.
Keeping those concepts separate prevents training simulations from turning into crude point systems.
Do not score everything
One of the easiest mistakes is to create a long rubric containing every desirable behaviour and ask the AI to score all of it continuously.
That creates several problems:
- the scoring becomes difficult to explain;
- small model errors can create apparently precise but unstable numbers;
- learners may optimise for the score instead of the skill;
- feedback becomes a checklist rather than coaching; and
- authors lose track of which measures actually affect the experience.
For many scenarios, three to six well-defined dimensions are enough.
Example: a sales role-play scorecard
Suppose the learner is practising a discovery call with a skeptical prospective customer.
| Dimension | What good performance looks like | How it can affect the scenario |
|---|---|---|
| Discovery | Asks useful questions before pitching | Unlocks the customer's real business problem |
| Listening | Responds to what the customer actually said | Raises trust |
| Value articulation | Connects the product to the discovered problem | Makes the customer more receptive |
| Objection handling | Addresses concerns without becoming defensive | Reduces perceived risk |
| Next step | Ends with a clear, mutually agreed action | Determines whether the scenario succeeds |
This scorecard is simple enough that the learner can understand it, but rich enough to drive both behaviour and feedback.
Feedback should explain cause and effect
The weakest feedback looks like this:
“You scored 7/10 for empathy and 6/10 for communication.”
That tells the learner almost nothing.
Stronger feedback connects an observable behaviour to its effect:
“When the customer said they had an important call in 20 minutes, you apologised but did not ask what they needed immediately. Because that requirement was never discovered, you kept discussing the room delay instead of solving the urgent problem.”
That is useful because the learner can see the moment where the conversation changed.
The best feedback can draw from several sources:
- the transcript;
- explicit stats or state changes;
- objectives completed or missed;
- character reactions;
- important turning points in the conversation; and
- the final outcome.
Feedback does not need to be delivered by a report
One advantage of character-based training is that feedback can itself be part of the experience.
A coach can appear after the role-play. The customer can leave, the scene can change and a training coach can talk through what happened.
That makes the transition from simulation to reflection much more natural than dropping the learner onto a dashboard full of scores.
The coach can still show structured data, but the learner can also ask follow-up questions:
“Why did the customer get more frustrated there?”
or:
“What could I have said instead?”
This turns assessment into a second conversation rather than a dead-end result screen.
Multiple characters make training scenarios more realistic
Many workplace situations involve several stakeholders.
A sales simulation might begin with a supportive internal champion and later introduce a finance director. A customer-service role-play may escalate to a manager. An interview can involve a panel rather than one generic interviewer.
The different characters can have different goals and different information. That matters because a learner may need to adapt communication style rather than repeat the same answer to everyone.
Liforma supports this through multi-character and multi-scene Avatar Experiences. For the underlying design model, see Multi-Character AI: How to Build Conversations With Multiple AI Characters.
Scenes let training have phases
A good simulation does not have to be one long conversation.
Scenes are useful when the training objective changes.
A management simulation might use:
- Preparation: learner receives context about the employee and issue.
- Conversation: learner speaks with the employee.
- Escalation: another character joins if the discussion goes badly.
- Reflection: coach reviews the interaction.
- Retry: learner repeats the difficult section with a different strategy.
The learner is still inside one training experience, but each phase can have a different purpose, character cast and environment.
Retries are one of AI role-play's biggest advantages
In a classroom role-play, repeating the exact scenario three times can be awkward and expensive.
With an AI simulation, retrying can be the default.
A learner can make a poor choice, receive feedback and immediately try the situation again. The character can keep the same underlying motives while generating a fresh conversation rather than repeating a script word for word.
This is where AI can complement human-led training particularly well: instructors can focus their time on higher-value coaching while learners get more repetitions between sessions.
Consistency matters for assessment
Generative AI introduces variation, which is useful for practice but potentially problematic for formal assessment.
If two learners receive completely different levels of difficulty, comparing their scores becomes harder.
For higher-stakes use, authors should control the parts that matter:
- the character's underlying objectives;
- facts that must remain true;
- difficulty level;
- the important objections or events;
- the scoring rubric; and
- conditions for success or escalation.
The wording can remain generative while the structure stays comparable.
Training simulations should separate practice from certification
AI-generated scoring is useful for practice, coaching and formative assessment, but organisations should be cautious about using an opaque model-generated score as the sole basis for consequential employment decisions.
A sensible pattern is:
- practice: frequent AI role-play with automatic feedback;
- progress tracking: trends across repeated attempts;
- manager or trainer review: important sessions can be sampled or reviewed; and
- formal assessment: use validated rubrics and appropriate human oversight where the stakes are high.
That keeps AI in the role where it is particularly strong: providing scalable opportunities to practise.
What should an L&D team measure?
It is useful to distinguish three levels of measurement.
1. Experience metrics
- Did the learner complete the scenario?
- How long did it take?
- Did they retry?
- Where did they stop?
2. Behaviour metrics
- Did they ask the right questions?
- Did they acknowledge the other person's concern?
- Did they identify the underlying need?
- Did they reach the intended outcome?
3. Learning metrics
- Does performance improve across attempts?
- Do learners transfer the behaviour to another scenario?
- Can managers observe improvement on the job?
- Does the intervention improve a meaningful business or learning outcome?
The first two are easy to collect inside the experience. The third is what ultimately tells you whether the training works.
Example: difficult management conversation
Imagine a new manager needs to practise addressing repeated missed deadlines with a team member.
The employee character begins defensive because they believe the deadlines were unrealistic. They also have a legitimate but undisclosed problem: work has been arriving from two different managers with conflicting priorities.
The learner's task is not simply to “be empathetic.” They need to:
- state the performance issue clearly;
- ask enough questions to understand the cause;
- avoid making assumptions;
- discover the conflicting priorities;
- take appropriate ownership; and
- agree a concrete next step.
The simulation might track clarity, psychological_safety, root_cause_found and action_plan_quality.
If the learner is accusatory, psychological safety falls and the employee reveals less. If the learner asks good questions, the hidden cause becomes available. The learner's behaviour therefore changes the information they can access.
That is a meaningful simulation mechanic, not merely a score at the end.
Example: customer complaint training
A customer arrives frustrated because a delivery is late. The obvious training goal is to resolve the complaint, but the scenario can be designed around more specific behaviours.
The customer may care less about a refund than about receiving one component before an event tomorrow. If the learner asks good discovery questions, the experience reveals that need. If they immediately offer a generic discount, the customer remains dissatisfied.
A training author can therefore measure the difference between appeasing the customer and understanding the customer.
Example: language learning
The same architecture works outside corporate soft-skills training.
A learner practising Spanish can enter a café, order food from one character, ask another character for directions and then receive coaching afterwards.
State can track whether key vocabulary was used, whether the learner understood a correction and which communication goals were completed. The environment and character roles make the language practice contextual rather than abstract.
Why avatars can matter beyond the LLM
The language model generates the conversation, but the visual character affects the experience.
A learner can see who they are addressing. Different characters can have recognisable roles. A customer can visibly become calmer or more frustrated. A scene change can mark the transition from simulation to coaching.
This does not mean every training experience needs photorealistic digital humans. In many cases a stylised character can be more practical, less uncanny and easier to reuse across a programme. See Photorealistic vs Stylized AI Avatars.
Training content should be reusable too
A training team rarely wants one isolated simulation. It wants a library.
That is where reusable characters, sets and experience structure become valuable.
A single coach can appear across an entire programme. A customer character can be reused with different problems. The same meeting-room set can host a management conversation, interview and negotiation. A successful scenario can be remixed into a new one instead of rebuilt from zero.
Liforma's creator model is designed around this reuse: characters, costumes, hairstyles, sets and experiences are separate building blocks.
Subject-matter experts should be able to author the simulation
The person who understands a difficult customer conversation is often a trainer, sales leader or operations manager rather than a software engineer.
That is why no-code authoring matters.
The author should be able to define:
- the scenario;
- character roles and motives;
- the behaviours that matter;
- stats and state;
- scene transitions;
- success conditions; and
- feedback.
The underlying STT, LLM, TTS, animation and session infrastructure should not have to be rebuilt for every training module.
For the practical creation workflow, see How to Create an Interactive AI Role-Play Without Coding.
What does an AI avatar training platform need?
If you are evaluating platforms for role-play training, look beyond avatar quality.
| Capability | Why it matters |
|---|---|
| Natural conversation | The learner should be able to respond in their own words. |
| Character control | The AI needs stable motives, knowledge and behaviour. |
| Explicit state | The learner's actions should affect what happens next. |
| Scoring / stats | Important behaviours need to be measurable. |
| Feedback | Learners need to understand what happened and how to improve. |
| Multiple characters | Many real scenarios involve more than one stakeholder. |
| Multiple scenes | Training can move through phases and consequences. |
| Repeatability | Learners need to retry without rebuilding the scenario. |
| Authoring by non-developers | Subject-matter experts should be able to change the training directly. |
| Analytics | Teams need to see usage and performance trends. |
Cost matters more than it first appears
Training experiences can be much longer than ordinary support interactions.
A ten-minute practice conversation may be repeated several times by every employee in a programme. At hundreds or thousands of learners, per-connected-minute pricing becomes significant.
Liforma Live is currently priced at approximately $0.01 per generated speech minute for STT, intelligence, TTS and speech-to-animation, without a separate per-connected-minute session charge.
That means silent thinking time, reading feedback or moving between scenes does not automatically create the same cost as generated character speech.
For a detailed comparison of current avatar pricing models, see How Much Do Interactive AI Avatars Cost?.
AI role-play is strongest as a practice layer
The most useful way to think about AI avatar training is not as a complete replacement for trainers, managers or instructors.
It is a scalable practice layer between learning and real-world performance.
A learner can receive instruction, practise privately with an AI character, get immediate feedback, try again, discuss the result with a coach or manager and then apply the skill in the real world.
That combination uses AI for what it can do unusually well: provide an interactive counterpart on demand and repeat the experience as many times as needed.
Liforma's model: train through experiences, not just conversations
Liforma treats the conversation as one part of a larger Avatar Experience.
A training author can combine reusable characters, different appearances, sets, multiple scenes, state, stats and feedback into one experience. The learner can interact naturally, while the experience still has structure and outcomes.
That is the core idea: the character creates the conversation; the experience creates the training.