On July 22, 2026, Synthesia launched Roleplay Sessions as an enterprise offering for practicing sales, support, management, field, and frontline conversations with live AI avatars. Learners can rehearse a conversation and receive real-time coaching while the system scores the attempt against a role-specific rubric. Managers can inspect attempts, transcripts, and roleplay analytics instead of relying only on course completion.
That makes practice more measurable, but it does not establish that employees perform better with customers or colleagues. A team deciding whether to buy now should focus on a narrower question: can one existing training program compare the product’s scores with independently observed performance on the job? This article explains what the launch provides, what its measurements can support, what buyers still need to verify, and when a limited evaluation is justified.
Enterprise teams can use Roleplay Sessions now; wider access remains a company goal
Roleplay Sessions is available as an enterprise product today. TechCrunch reported that it is the first release in a broader planned Sessions platform and that Synthesia aims to extend access to small and mid-sized businesses, prosumers, and education as the cost of running the underlying models falls. The company described that expansion as a goal for the coming months, not a committed release date.
Synthesia lists roleplays in English, German, Spanish, and French. Its separate Studio product supports video creation in more than 160 languages, but that larger figure does not describe the current language range for live roleplay. Enterprise teams can deliver the sessions through a learning-management system using SCORM, a common standard for packaging and tracking training content. Smaller organizations should not plan around an availability date Synthesia has not committed to.
Learners rehearse a conversation while managers inspect the attempt
A learner enters a scenario and speaks with an interactive avatar. Synthesia describes use cases such as cold calls, discovery calls, objection handling, customer complaints, performance feedback, and leadership conversations. During the session, the system can provide coaching and score the attempt across several skills instead of issuing a single pass-or-fail result.
The rubric is the set of criteria used to judge the attempt. Depending on the role, it might distinguish how a learner opened a conversation, handled an objection, gathered information, or closed the exchange. The public product page says every roleplay is scored against a role-specific rubric, but it does not explain enough about who creates the criteria, how they are calibrated, or how differences among reviewers are resolved.
Synthesia says managers can inspect key performance indicators, score distributions, skill breakdowns, individual attempts, and transcripts, then export results or send them to an LMS. A training owner can see who practiced, how the system scored each attempt, where coaching focused, and whether scores changed over repeated sessions.
A rubric score measures the exercise, not performance on the job
Scored practice provides more information than passive completion. A course-completion record shows only that someone reached the end of the material. A roleplay attempt shows how the person responded under a defined scenario, which parts of the rubric the system judged weak or strong, and whether a manager agrees after reading the transcript.
The score still belongs to the simulation. It reflects the scenario, rubric, model behavior, and scoring thresholds that produced it. A higher result does not by itself show that a salesperson qualifies opportunities more accurately, a support representative resolves cases better, or a manager handles a difficult conversation well with a real employee. Synthesia’s public materials do not provide independent evidence connecting its roleplay scores to those outcomes.
The team using the product has to establish that connection. Sales training could compare score changes with an existing manager review of relevant calls. Support training could use its established quality review or a customer outcome tied to the skill being taught. Manager training may require structured observation by an experienced leader because a simulation cannot capture every consequence of trust, history, authority, and follow-through.
The direction of a scoring error matters too. A weak rubric may rate someone as ready when live work shows otherwise, or hold back a capable employee because the simulation rewards a narrow style. Those errors carry different costs, much like false positives and false negatives in an approval process. Managers need access to attempts and transcripts so they can inspect disagreements rather than treating the score as the final decision.

The buying decision depends on an independent measure of the same skill
Begin the evaluation with an existing training program that teaches a repeatable conversation a manager can recognize in real work. The comparison must sit close to the skill being practiced. A general revenue target is too far downstream for an objection-handling exercise; a manager’s review of relevant call behavior is closer to what the session claims to measure.
The training owner and frontline manager should review the proposed rubric before employees rely on it. The criteria should describe the actual job, allow different communication styles to satisfy the standard, and include the difficult cases the training is meant to address. People who know the work should review early attempts so disagreements between the automated score and human judgment become visible.
Before employees begin, the team should decide what result would justify continuing. Scores should align with experienced reviewers, repeated practice should improve the behavior under review, and the same change should appear in the existing measure of live work. If only the product score rises, employees may be getting better at the exercise rather than the job.
Buyers still need answers about price, scoring controls, data, and results
Synthesia does not publish Roleplay Sessions pricing on the public page; buyers are directed to sales. The decision depends on a quote and on the work required to create scenarios, review rubrics, inspect disputed scores, support learners, connect the LMS, and compare results with live performance. Qualified people still have to decide whether the measurement deserves trust.
Rubric calibration, bias testing, and review are unresolved in the public material. Buyers should ask how Synthesia tests scoring consistency, how customers change criteria and thresholds, how the system handles accents and varied communication styles, and whether an employee or manager can challenge a result. The score may influence coaching, certification, or access to customer-facing work even when it was designed only as a training signal.
Data handling also needs a direct review. TechCrunch reports that OpenAI provides the reasoning layer, while Synthesia provides the avatar and voice experience, rubrics, performance data, and analytics. Buyers should confirm which conversation content reaches each provider, where audio and transcripts are stored, how long they are retained, who can access them, and what deletion controls apply. Synthesia’s product page lists SOC 2 Type II, ISO 27001, ISO 27701, ISO 42001, GDPR compliance with EU data residency, single sign-on, and automated user provisioning through SCIM. Legal and security teams still need to verify the scope and current status of those claims for the proposed deployment.
Evidence of business results remains limited. According to TechCrunch, Synthesia says it has early customers with scaled commercial rollouts, but describes them only by category and size. The public sources do not name the customers, quantify outcomes, explain the comparison used, or provide independent validation. Deployment alone does not show what changed in sales, support, or management performance.
Add it to one existing program only when the outside comparison is ready
An enterprise team should evaluate Roleplay Sessions now when it has a repeatable conversation to teach, a manager who can judge the rubric, and an existing way to observe the same skill in live work. Keep the current training and human assessment in place. Use one conversation type with a limited group, review a sample of attempts and disputed scores, and compare score changes with the chosen job measure before expanding access.
Wait when the team cannot name that outside measure, when sensitive conversation data cannot pass the vendor review, or when the price and internal review effort cannot be justified. Smaller organizations should also wait for documented access and pricing instead of planning around Synthesia’s broader availability goal. In either case, reconsider the product when those conditions change rather than treating the July launch as a deadline.
BaristaLabs can help connect one training exercise to the manager review or job measure your team already uses through process automation and integration. If your team has one conversation type and a real-world measure in mind, review how you will compare practice with live work.
Sources
Training evidence review
Compare one practice score with the work it is meant to improve
Bring one repeatable conversation, the proposed rubric, and the manager review or job measure already used by the team.
Best fit when the team can observe the same skill in practice and in real work.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before booking a call.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.
