Your "High Potential" Raters Probably Don't Agree With Each Other

A synthetic-data walkthrough of turning "high potential" into a measured construct — and what happened when two independent raters scored the same people.

Most "high potential" nominations come from a single manager's gut feel — one person, one rating, no second opinion, no written definition of what "potential" even means. In a demo analysis using synthetic data, I split "high potential" into three defined facets (learning agility, drive for scope, interpersonal influence) and had two independent raters score the same population on each.

The facets weren't equally measurable. Interpersonal behavior was easy for two raters to agree on. "Drive for scope" — the most inferential of the three — barely cleared the reliability bar for a single rater (ICC 0.50). Averaging both raters' scores together lifted every facet, and pulled the overall composite from "moderate" to "good" reliability territory.

Reliability isn't the whole story: the structured composite also correlated more strongly with an independent 12-month outcome than a naive single-rater nomination did (r = 0.45 vs. r = 0.34) — though honestly, the confidence intervals on that comparison overlap, and the write-up says so rather than overselling the gap.

Read the full case study →