Measuring Development of Judgement, Creativity and Critical Thinking

A practical approach for business and professional learning

A well-trained manager may learn your decision-making framework and recite it at will, yet still struggle to apply it when under pressure - when the decision is complex, the client is waiting, the facts are incomplete, and the team is divided.
For learning leaders, this creates a measurement challenge: how do we establish whether a program has improved someone’s ability to think creatively or critically, and act appropriately or decisively in a difficult situation? How do we measure critical thinking? How do we measure judgement, creativity and emotional intelligence under pressure?
The answer begins with observable performance.
Ask people to tackle unfamiliar problems, capture the reasoning behind their responses and examine how that reasoning develops over time. Knowledge remains essential, but a knowledge test alone cannot demonstrate that someone can weigh competing priorities, create useful alternatives or challenge a persuasive assumption.
Define what better thinking looks like
We've recognized the need to make broad claims, such as “improve critical thinking” into specific assessment criteria. For critical thinking, we're looking to develop the ability to distinguish evidence from assumptions, evaluate sources, consider competing explanations and revise conclusions when warranted. For judgement, we want to examine how participants weigh trade-offs, acknowledge uncertainty and choose actions proportionate to the situation.
For creativity, we need to assess whether participants generate meaningfully different possibilities, develop ideas that are both original and useful, and improve them in response to constraints or feedback.
Established educational approaches offer useful starting points. OECD’s PISA creative-thinking assessment uses open-ended tasks to examine the generation of diverse and original ideas and the evaluation and improvement of ideas. AAC&U’s VALUE rubrics provide frameworks for evaluating authentic student work. These approaches inform assessment design; they do not automatically validate a corporate assessment adapted from them. [1, 2]
For example, The Critical Thinking VALUE rubric examines five areas, paraphrased here:
Area | Plain-language question |
Understanding the issue | Have you explained the problem clearly? |
Using evidence | Have you evaluated the information supporting your answer? |
Context and assumptions | Have you examined what you and others are taking for granted? |
Your position | Does your recommendation account for complexity and other perspectives? |
Conclusions and consequences | Does your conclusion follow logically, with its implications considered? |
Capture reasoning before the outcome is known
Use a realistic task with enough ambiguity to require thought. Before revealing consequences, ask participants to record their recommendation, supporting evidence, assumptions, alternative perspectives and anticipated effects. Ask what information would change their decision.
Assess the reasoning against what was knowable at that point. A favorable outcome may involve luck, while a defensible decision can still have an unfavorable result. Where several options are reasonable, the assessment should recognize different well-supported responses. Where professional obligations constrain the options, make those constraints explicit.
The assessment concerns the reasoning - not whether the person gets the "right" answer.
Look for development across unfamiliar tasks
A strong answer at the end of a workshop is useful evidence, but it does not establish lasting development. Participants may be reproducing a recently demonstrated approach or responding to facilitator prompts. Assessment needs to distinguish supported practice from independent performance.
Start with a brief unfamiliar task before the program.
Afterwards, use a different task requiring similar skills at comparable difficulty. Follow up several weeks later with another challenge and a real example of workplace application. This schedule is a practical design proposal, not a universal assessment standard.
Use several samples where feasible.
Familiarity with the industry, background knowledge and task difficulty can all affect performance. Provide necessary information and keep assessment conditions reasonably consistent. A participant’s success on one scenario should not become a sweeping claim about their judgement in every context. The National Academies’ assessment work explicitly addresses the importance of context and the transfer required by assessment tasks. [3]
Make scoring specific and consistent
A rubric describes what different levels of performance look like. For evaluating evidence, a simple progression might run from asserting a preference without support, to citing relevant information, to weighing conflicting evidence, to explaining both a conclusion and its limitations. Each level needs examples so assessors can apply it consistently.
Imagine the question: “Should our company introduce an AI customer-service system? The vendor says it will reduce costs by 30%.”
We will assess just one area: how the person evaluates evidence.
Illustrative level | Participant’s response | What it reveals |
1 | “We should introduce it because the vendor says it will save 30%.” | Accepts the claim without evaluating it. |
2 | “The savings look attractive, but we should check whether implementation costs are included.” | Begins to question the evidence but examines only a narrow part of the claim. |
3 | “We need to compare the vendor’s estimate with our service volumes, implementation costs and pilot results. Savings elsewhere may not apply here.” | Evaluates the claim using relevant evidence and context. |
4 | “The pilot suggests savings on routine enquiries, but complex cases required more staff time. The vendor’s estimate excludes that cost. We should reconcile those findings before estimating total savings and state the uncertainty in our recommendation.” | Integrates conflicting findings, examines limitations and reaches a qualified conclusion. |
Agree the criteria before reviewing results. Ask assessors to practice on sample responses, discuss disagreements and independently score a proportion of the work. Where possible, conceal whether responses came from before or after the program. Reward the quality of reasoning rather than length, vocabulary or confidence of delivery.
Provide accessible response formats and distinguish independent performance from work completed with AI or other assistance. Both can be relevant to the workplace, but they answer different assessment questions. Report a profile of strengths and development needs rather than forcing every skill into one overall number.
Assess revision as well as the initial answer
After participants commit to a position, introduce information that challenges it. An important assumption may prove uncertain, a stakeholder may raise an overlooked concern, or a resource constraint may change. Examine whether participants recognize the significance and respond with a justified revision.
Changing a decision is not automatically a sign of good thinking. Maintaining it may be reasonable if the new evidence does not materially alter the balance. The assessable skill is the ability to explain why an adjustment is, or is not, warranted.
Connect development to workplace evidence
Ask participants to document a decision or project in which they applied the skill. Useful evidence includes a proposal revised after testing assumptions, a decision record comparing competing options, or an idea improved through feedback. Capture what the participant did, why it mattered and what changed. A manager’s observation can strengthen this record when it describes a specific behavior.
Keep three questions separate: did performance improve, did the program contribute to that improvement, and did it make a difference at work? Before-and-after assessments address the first. A credible comparison group, ideally with random assignment where practical, strengthens the second. Workplace evidence and relevant operational measures help address the third.
Business outcomes such as fewer errors or better project delivery can support the evaluation, but other factors influence them. Avoid attributing every improvement to training.
Equally, satisfaction, completion and greater confidence should be reported for what they show; they do not independently demonstrate stronger judgement or critical thinking.
For CLOs, a useful dashboard could show the proportion of participants meeting defined performance criteria, changes across comparable tasks, retention at follow-up and verified examples of application. Include sample sizes, follow-up participation and assessor agreement. Describe movement between rubric levels rather than converting it into an unsupported claim that “creativity increased by 40%.”
How we capture higher-level cognitive skills development with Decision Rehearsal™
Decision Rehearsal™ uses facilitated, audio-driven scenarios and participant worksheets to make decision-making processes visible. Our current worksheets capture the initial choice, assumptions, perspectives that could challenge it, anticipated consequences and information that might change the decision.
These responses create a record of thinking before the scenario reveals what happens next.
What our measurement approach can establish
Our current approach supports qualitative assessment and structured self-reflection. Participants can examine the assumptions and pressures in their initial response and identify more considered ways to act. Where responses are voluntarily shared, and there is a clear agreement about participation and use, facilitators can discuss the evidence behind those reflections.
The strongest current evidence concerns critical thinking, judgement, creativity, foresight and reflective awareness. Prompts about tone, posture and bodily sensations capture awareness and intended responses; observed rehearsal would provide additional evidence of behavior.
To demonstrate development over time, the next measurement step is to compare responses across successive unfamiliar scenarios using common criteria.
We look for
stronger questioning of presented options
stronger examination and use of evidence
clearer recognition and questioning of assumptions
more plausible anticipation of consequences
justification of rationale behind decision making
emotional intelligence and self-regulation under pressure
clear communication and expression of thoughts, ideas and emotions
Decision Rehearsal™ currently provides a structured record from which development in these areas can be examined and compared over time.
For learning leaders, its measurement value lies in making reasoning available for reflection and comparison, with a clear route towards repeated assessment of the decisions people are better equipped to make.
Please comment, connect, ask questions or challenge anything in this article.
Russell Cullingworth, CEO, ProDio
Creator of Decision Rehearsal™
Sources
1. OECD. PISA 2022 Creative Thinking. https://www.oecd.org/en/topics/sub-issues/creative-thinking/pisa-2022-creative-thinking.html
2. American Association of Colleges and Universities. VALUE Rubrics. https://www.aacu.org/value/rubrics
3. National Research Council (2001). Knowing What Students Know: The Science and Design of Educational Assessment. Chapter 4, Contributions of Research on Human Learning to the Design of Assessment. https://www.nationalacademies.org/read/10019/chapter/6



Comments