This is a demonstration, not a client engagement. The ratings and outcomes below come from a synthetic-data simulation, not a real organization. Every number is the script's actual output — nothing was adjusted to look better. Unlike the promotion-readiness case study, this one isn't about prediction — it's about the measurement step that has to come first: does the label even mean the same thing to two different observers?
The question
Ask five HR leaders what "high potential" means and you'll get five different answers — some mix of ambition, intelligence, likability, and whoever a VP happened to notice in a meeting. Most companies still run their HiPo process on a single manager's holistic nomination: one person, one gut-feel rating, no second opinion, no definition anyone wrote down.
That's not a measurement process. As covered in an earlier piece on what measurement actually requires, the test is simple: can you describe what the world looks like when the construct is high versus low, in terms specific enough that two people watching the same person would agree? If not, you don't have a measurement problem — you have a definition problem, and it has to be solved first.
Step 1: Define the construct
"High potential" was split into three separable facets, following the framing Lombardo and Eichinger popularized in their work on learning agility as the core of potential identification (Lombardo & Eichinger, 2000):
- Learning agility — seeks out unfamiliar situations, extracts lessons from them quickly, and applies those lessons under new conditions rather than repeating what worked before.
- Drive for scope — consistently seeks larger scope and ambiguity rather than stability, and sustains effort against a demanding stretch goal over time.
- Interpersonal influence — builds followership and earns buy-in from people outside their direct authority, adapting style to influence upward, across, and outward.
Each facet was written as a behaviorally anchored 1–5 rubric — concrete, observable behaviors at each scale point, not adjectives like "strong" or "exceptional" left to the rater's imagination.
Step 2: Build the rubric, use two raters
Every employee in the talent-review population was rated on all three facets by two independent observers — their manager and a calibration-panel rater who'd seen their work in a cross-functional setting. Two raters isn't bureaucratic overhead; it's the only way to find out whether the rubric measures something real or just one person's impression.
Step 3: Check whether the raters agree
Reliability was assessed with the intraclass correlation coefficient (ICC) — the standard tool for rater agreement (Shrout & Fleiss, 1979) — using the absolute-agreement form, the stricter version that penalizes raters who rank people similarly but land on different actual scores.
The three facets aren't equally measurable. Interpersonal influence is the easiest to observe consistently — two raters land close to the same score (single-rater ICC 0.81, "good" territory on its own). Drive for scope is the hardest — it's more inferential, and a single rater's judgment of it is only moderately reliable (ICC 0.50, right at the "poor/moderate" boundary). Learning agility falls in between (ICC 0.70).
Averaging the two raters' scores together lifts every facet, and lifts the composite from 0.78 to 0.87 — crossing clearly into "good" reliability. This is the practical case for a calibration panel rather than a single nominating manager: it's not about catching bias in any one rater, it's that two independent, moderately-reliable judgments combine into one meaningfully more reliable score.
Step 4: Check whether it predicts anything
A rubric can be perfectly reliable and still measure the wrong thing. To check, the structured composite score was correlated against an independent criterion collected 12 months later — an assessment from a source not involved in building the rubric — and compared against what a naive, unstructured single-rater nomination would have produced on the same population.
The structured composite correlated more strongly with the later outcome (r = 0.45, 95% CI [0.37, 0.53]) than the naive single-rater nomination (r = 0.34, 95% CI [0.25, 0.43]). Worth saying plainly: those confidence intervals overlap. This sample doesn't let us declare with certainty that structured measurement is better in general — only that, in this data, it pointed that direction and by a meaningful amount. Overclaiming precision here would be exactly the mistake this whole exercise is designed to avoid.
What's more interesting than the size of the gap is why the naive nomination underperforms: it isn't just noisier, it's partly measuring the wrong thing. It was simulated with real sensitivity to a recent performance-halo effect — a visible recent win or loss — that has nothing to do with a person's actual trajectory. A single unstructured judgment doesn't just add noise to a good measurement; it can substitute a related-but-different construct for the one you meant to assess.
What this demonstrates — and its limits
The synthetic numbers above aren't a claim about what reliability or validity your organization would get — they illustrate the method: name the construct's facets explicitly, write behavioral anchors, use enough independent raters to check agreement, and validate against something the rubric wasn't built to predict.
A real engagement would go further than this demo: testing whether the rubric or its outcomes show adverse impact across demographic groups before it touches a real decision, training and calibrating raters rather than assuming naive independence, and validating prospectively against actual promotion or performance data rather than a single simulated criterion. None of that is optional in practice; all of it is out of scope for an illustrative write-up.
Want a construct in your organization actually defined and validated? This is the same measurement-design process used in real engagements — construct first, instrument second, validity checked rather than assumed. Start a conversation or read more about measurement design.