
General Manager

The number I most want to inspect in a learning report is the one that looks easiest to accept. A precise score can make an interpretation feel settled long before the underlying evidence deserves that confidence.
In Practice-Native Learning, the methodology I am developing at Altaius, the practice record should help someone understand a decision and improve it. Measuring that record responsibly means being clear about which question the evidence can answer.
I use four evidence layers when discussing an evaluation. This is a practical organising model, not a newly validated measurement framework.
A completed session can establish participation. It does not establish improved judgement. A stronger second attempt can reveal better performance in that case, but changes in assistance or familiarity may help explain it. Workplace observation and business results need additional evidence.
The 2014 Standards for Educational and Psychological Testing treat validity as evidence supporting score interpretations for specified uses. They do not treat it as a permanent stamp attached to a test regardless of how it is used. [1]
That distinction matters here. “This response supports a coaching discussion about checking authority” is a narrower use than “this person is ready for promotion.” Evidence sufficient for one should not be silently reused for the other.
For any proposed score, I would write down the intended interpretation, the intended user, the decision it may inform and the situations in which it should not be used. If those statements cannot be made clearly, adding decimal places will not solve the problem.
An illustrative evidence record could show that the participant promised a date, had not consulted the capacity note and received a coaching prompt before revising the promise. The scenario version and scoring rule should be identifiable. The reviewer should be able to distinguish the participant's actions from the character's statements and the coach's help.
“Not observed” also needs a place in the record. If the exercise never created an opportunity to demonstrate a behaviour, treating missing evidence as poor performance would add a claim the session cannot support.
I would want a process for reviewing disputed cases, checking agreement between appropriate reviewers and examining inconsistent feedback. The required evaluation depends on the interpretation and its consequences; a development conversation and a consequential assessment do not have identical evidence needs.
An organisation may reasonably ask whether the programme contributes to fewer avoidable escalations or better handovers. Define the measure, observation window and other plausible causes before the programme starts. Include the time and support needed to run it.
A before-and-after improvement can be encouraging without isolating the programme's contribution. Changes in staffing, workload, policy or manager attention may also matter. Where feasible, a suitable comparison or staged rollout can strengthen the evaluation. Where it is not feasible, state that limitation.
I would not convert a simulation score into a financial return by assigning it an arbitrary currency value. The bridge from learning to an operational measure has to be demonstrated, not filled in with a confident-looking formula.
“Participants improved on the assessed scenario under these conditions; workplace transfer has not yet been evaluated” can be an informative result. It gives the next stage of work a clear question.
Altaius's methodology does not make our scores automatically valid or establish client outcomes. I want the reporting to make the evidence inspectable enough that a reader can tell what is supported, what remains uncertain and which observation would change the conclusion.
Read the Altaius definition of Practice-Native Learning
Previous in this series: Taking Practice-Native Learning into the Workplace