The cost of multilingual product onboarding video is usually calculated wrongly, and the error explains why so many teams abandon it after the first attempt.
The calculation that gets made covers production: a presenter, a studio afternoon, a voice artist per language, subtitling, editing. That number is large but finite, and teams with a case for international expansion approve it.
The number that ends the project is the second one. The product changes, four sentences of the script become wrong, and the entire set has to be produced again in every language. What was a one time cost turns out to be a recurring cost tied to release cadence, and release cadence is not something a growth team controls.
The reshoot is the whole problem
Almost every failed internal video programme fails at the same point, and it is not the first production.
Consider a five minute onboarding video produced in five languages. Three months later, a settings menu moves and a feature is renamed. The change affects perhaps forty seconds of narration. Re-recording forty seconds requires the original presenter, matched lighting, matched wardrobe and matched audio conditions, in five languages, and the voice artists are on other contracts.
In practice teams either add a supplementary video that nobody watches, or leave the set wrong with a text note attached, or stop maintaining it, at which point it degrades into a collection of confident recordings describing a product that no longer exists.
None of these are tooling failures. They follow from the fact that filmed video is expensive to amend and cheap to leave alone.
What changes when the presenter is generated from a still image
Photo based avatars remove the amendment cost, which is a different claim from removing the production cost.
The mechanism is that the presenter is not a recording. Platforms that create talking avatar from photo input generate mouth movement, head movement and expression around the facial landmarks in a single still image, driven by a script. Leadde calls this a Photo Avatar and requires no video footage to build one.
The consequence for maintenance is that amending forty seconds of narration means editing forty seconds of text and regenerating. Lighting, wardrobe and audio conditions are identical because they were never conditions in the first place.
The input requirements are worth reading before choosing the photograph, because they constrain the whole programme afterwards. Leadde asks for a clear image that reflects the person’s current appearance, and excludes group photographs, hats, sunglasses, pets in frame, heavy filters, low resolution images and screenshots. A team standardising on one presenter image should treat it as infrastructure and pick something that will still look current in a year.
One constraint affects script structure directly. Leadde’s expressive animation mode, which infers head movement, gesture and posture from the script, is capped at sixty seconds per video. Longer material has to be built as a sequence of segments, which is usually the right structure for onboarding content anyway.
What 88 languages and 175 dialects actually decides
Leadde supports 88 languages and 175 dialects, and the dialect count is where the practical decisions sit. Spanish appears with more than twenty regional variants. Arabic appears with more than fifteen. English appears with fifteen, including Indian English. For a team selling into several Spanish speaking markets, the choice between neutral Spanish and a specific regional variant is a positioning decision, not a technical one.
The general principle from localisation practice is that language identification is insufficient on its own, which is why the Unicode CLDR project maintains locale data rather than a language list. A product team should be making the same distinction.
Market selection should follow evidence rather than language population. Support ticket language distribution shows where users are already struggling in a language the product does not serve, and organic search data by locale shows where demand exists ahead of investment. Teams often find that a smaller market with strong search position and weak conversion is a better first localisation than a large market with no position at all.
Version management is the part nobody plans
This is the section that determines whether the programme survives its second year, and it is procedural rather than technical.
The script lives outside the video platform, in whatever system the team already uses for content, with the source language version treated as authoritative. Every localised version is derived from it. When the product changes, the source script changes once.
Translated versions are generated from a finished source video rather than produced independently. Leadde translates a completed video into a new language, handling the narration and the on screen text together and producing a translated draft rather than overwriting the original. The advantage over parallel production is that the versions cannot drift structurally, because they share a structure.
Superseded versions are duplicated rather than overwritten during a transition period. Copying a video to produce an amended version keeps the previous one available for as long as someone may have acted on it, which matters when the change concerns a procedure rather than a name.
A fourth practice matters once the set grows: a review interval decided before the library exists, since generation is cheap enough that teams produce more videos than they can check.
What still needs a real shoot
Being clear about the limits protects the credibility of everything else.
Anything conveying that real people work at the company needs real people. A generated presenter is a delivery mechanism for information and does not carry that signal, and using it for a culture or recruitment video produces the opposite of the intended effect.
Anything demonstrating physical interaction with a product needs footage of the product being used. Generated presenters describe; they do not demonstrate.
Anything responding to a specific event needs a person, because the value of the message is that someone chose to say it.
Consent is not a formality
A photo based avatar is a synthetic likeness of a real person, and two obligations attach regardless of jurisdiction.
The person in the photograph has to agree, specifically, to their likeness being used to deliver scripted narration they did not record. An employment contract permitting use of a photograph does not obviously cover this, and the person should be told what the avatar will be used to say and for how long. Departure terms should be agreed at the same time, since a company presenting an ex employee as its onboarding presenter is a problem that is easier to prevent than to resolve. Leadde requires explicit authorisation for avatars of public figures, and the same standard is worth applying internally.
The audience has a claim too. Viewers should be able to tell that the presenter is synthetic, and the cost of disclosing is far lower than the cost of a viewer working it out.
Captions belong on every version from the start. Onboarding video is often watched in an office without headphones, and captions for prerecorded video are a WCAG success criterion.
Start with the segment that changes most
The strongest case for this approach is not the whole onboarding sequence. It is the two minutes covering the part of the product that changes every quarter, in the two languages with the clearest evidence of demand.
If that survives a release cycle without going stale, the approach works and the rest can follow. If it does not, the problem was the script discipline rather than the tooling, and finding that out cost two minutes of video instead of a full programme.
Add Business Connect magazine to your Google News feed




