
General Manager

Altaius is being built for people who work in Arabic and English, sometimes within the same conversation. That makes the language question much bigger than whether the interface has been translated.
If a participant can handle a difficult situation well in Arabic, I do not want a scoring system to miss that because it expects the wording of an English example. I also do not want a fluent translation to persuade a buyer that the two versions have already been shown to assess the same thing.
Bilingual simulation evaluation should examine whether the versions give participants comparable information and opportunities, and whether the scoring applies the intended criteria consistently. Translation quality is one part of that work.
The test design below is what I would ask to see. It is not a claim that Altaius has completed a formal equivalence study.
Consider a negotiation in which the participant is asked to promise an earlier delivery date. The intended action is to check authority and capacity before making a commitment.
Write the criterion in those terms. Do not define success as using a particular phrase such as "subject to approval". A participant may communicate the same condition naturally in several ways. Another may repeat the phrase while still leaving the customer with an unconditional promise.
The reviewer needs enough of the exchange to judge what was communicated. A sentence taken out of context can lose the condition that appeared immediately before it, or the commitment that came immediately afterwards.
For a free-text simulation, this matters throughout the conversation. Characters have their own knowledge, priorities and relationships. Their responses can change what the participant reasonably understands at the next turn.
Place the Arabic and English versions of the briefing beside each other. Check the facts, the authority limits, the available options and any uncertainty the scenario deliberately includes.
A translation can accidentally explain a constraint that the other version leaves implicit. That gives one participant a clue the other did not receive. A different tone can also make a customer's objection sound more or less serious.
I would ask reviewers who understand the language and the work situation to explain these differences. A word-for-word comparison is too narrow for this purpose. The task is to preserve the intended challenge while allowing natural language.
Use a small, documented set of responses to test how the scoring behaves. Begin with examples whose meaning reviewers can explain, then include ambiguous cases that require discussion.
| Example response | Question for the reviewer |
|---|---|
| Checks delivery capacity before confirming | Does the scoring recognise the intended action in both languages? |
| Promises the date, then adds an unclear qualification | Is the customer left with a commitment, despite the qualifying words? |
| Uses polite wording but makes an unauthorised promise | Does surface politeness receive credit that obscures the actual decision? |
| Uses a different natural phrase to state the same limit | Is meaning recognised without requiring a preferred expression? |
| Switches language while explaining the condition | Does the system preserve the meaning across the exchange? |
These are illustrative test cases, not model answers for participants to memorise.
Keep the expected judgement and the reviewer's reason with each example. If reviewers disagree, record the disagreement and examine the criterion. An average score can conceal an unresolved interpretation.
Decide whether language proficiency is part of the intended assessment. If the programme is about leadership decisions, a spelling error should not silently become a leadership penalty. If clear written communication is itself an agreed objective, explain how it will be assessed.
Test realistic phrasing rather than only polished formal sentences. The relevant mix will depend on the intended users and the workplace. Do not assume that one dialect, level of formality or style represents every Arabic speaker.
It is also worth looking at time. Reading and typing behaviour can differ across versions and users. A fixed timing threshold needs a defensible purpose; it should not become an unexplained shortcut for judging confidence or capability.
For each evaluation, record the scenario version, language, scoring criteria, model configuration where relevant, expected judgement and observed result. Keep access to the test material appropriate to its sensitivity.
Review failures by type. Was the information different? Was the response misunderstood? Did the criterion reward the wrong behaviour? Those findings lead to different corrections.
After changing a translation or scoring instruction, rerun the relevant cases in both languages. A local improvement may affect another part of the conversation.
The International Test Commission's guidelines on translating and adapting tests provide a reference point for adaptation and evidence requirements. Applying a short internal checklist does not establish formal equivalence or authorise high-stakes employment decisions.
For Altaius's bilingual simulation design, I want the explanation to remain inspectable: what was the person asked to do, what did they understand, and why did the response receive that judgement? If we cannot explain a difference between the versions, I would keep investigating before comparing their scores as though nothing had changed.