Last year I wrote that I had let go of the idea that validation had to mean scientific rigour alone, and that measurement is not the point but a portal: the starting point for a deeper conversation. I still believe that. This paper is not a retreat from it.
But a portal has to hold. Coaches build development plans on these numbers and organisations make decisions from them, so the instrument behind the conversation had better be sound. We now have 11,015 completed assessments from 10,417 people, which is finally enough to test it properly. What follows is what we found. Some of it is good. Some of it calls for caution about how the numbers are used rather than about the instrument itself. There are also some specific opportunities to improve the instrument, all of which we lay out below.
The assessment measures something real and repeatable at the level of the whole profile, it detects development in proportion to how much development work was done, and its individual skill scores are noisier than the framework’s structure implies, which is why the profile, not any single number, is what we ask coaches to work from.
What follows on this page is a summary. It carries the headline findings and the reasoning behind them, but not the full analysis: the complete white paper runs to fifteen pages and includes the method, the per-scale figures, the statistical tests behind each claim, every limitation we are aware of, and the mistakes made while producing it. If you intend to rely on any of this, or to challenge it, read the paper rather than this page.
The Full Analysis
How reliable and valid is the IDG assessment?
Fifteen pages. The complete study: method, every figure, all limitations, and the corrections we had to make. This page covers roughly a third of it.
Two different questions
“Is it validated?” bundles together two things that are tested differently and can come apart. Reliability asks whether the instrument measures consistently: if several questions capture the same quality, people should answer them consistently. Validity asks whether it measures the right thing. Ten questions about shoe size would have excellent reliability and tell you nothing about presence.
This is the distinction most assessment marketing blurs. A high reliability score is often quoted as though it settled validity. It does not, and we have kept them apart throughout.
Part one: what works
Everything in this section is a finding we would stand behind, and each one is stated with what would have counted against it.
The instrument as a whole is reliable
| Scale | Questions | Reliability |
|---|---|---|
| The whole instrument | 83 | 0.91 |
| Being | 54 | 0.89 |
| Thinking | 37 | 0.85 |
| Relating | 30 | 0.81 |
| Collaborating | 30 | 0.81 |
| Acting | 41 | 0.87 |
| Individual skills reaching 0.70 | — | 5 of 25 |
What the instrument gets right
It behaves as the theory predicts. If inner development genuinely accumulates with life experience, scores should rise with age. They do, in the right order, with no exceptions:
| Generation | Mean overall score | People |
|---|---|---|
| Baby Boomers | 81.3 | 150 |
| Generation X | 78.8 | 612 |
| Millennials | 76.2 | 799 |
| Generation Z | 73.4 | 1,470 |
It also carries a warning. Because the age effect is this large, comparing any group against a general average mostly measures its age composition. Every cohort comparison we publish is matched on generation for that reason, and any comparison that is not should be treated with suspicion, including comparisons of our own published averages.
Everyone’s profile is uneven, and consistently so. The distance between a person’s highest and lowest skill averages 26.7 points, and 97% of people exceed 15 points. An uneven profile is the normal condition, not a sign of imbalance, which is worth saying to anyone reading their own report for the first time.
The same organisation gives the same reading. One client was assessed three times over more than two years, 2,113 people in total. Taking each wave’s profile relative to everyone else on the platform, the shapes correlate at 0.93, 0.91 and 0.98. The same skills stand out every time.
Does it detect development?
Stability is only half an argument. An instrument that never moves is not stable, it is blind. At that same organisation, 293 individuals could be matched across two assessments by employee number: the same people, the same instrument, two points in time.
| Measure | Change | Beyond chance? |
|---|---|---|
| Overall score | +1.30 | Yes |
| All five dimensions | +1.16 to +1.64 | Yes, all five |
| Individual skills moving beyond chance | 18 of 25 | About 1 expected by chance |
An honest note on method: compared at organisation level rather than person level, the same change reads +0.51 and is indistinguishable from noise, because the people taking part differed between waves. Pairing individuals removed that and revealed the effect.
Most telling is that the size of the change tracks what was actually done:
| What was done | Change |
|---|---|
| Two years of individual development work | +6.6 |
| Comparable period, no targeted development | +3.0 |
| Structural redesign aligning people to their strengths, no inner-development work | +1.30 |
A structural intervention that never targeted inner development produced about a fifth of the movement seen in dedicated individual development. The instrument does not simply move; it moves in proportion to how much of the relevant work was actually done. That pattern is difficult to produce by accident.
Do the five dimensions hold?
They do, but weakly, and the first test we ran said otherwise. Compared directly, two skills in the same dimension are no more related than two skills in different ones: 0.547 against 0.544, a difference of 0.003. Read alone, that says the dimensions are decorative.
But the 83 questions are shared between skills. Each question feeds 2.7 skills on average, and two skills built partly from the same answers are bound to look alike. Across all 300 pairs of skills, correlation rises directly with how many questions the pair has in common: 0.461 with none, 0.568 with one, 0.664 with two, 0.748 with three or more.
So the honest test is to compare only skills built from entirely separate questions.
| Pairs sharing no questions | Pairs | Average correlation |
|---|---|---|
| Same dimension | 34 | 0.494 |
| Different dimensions | 122 | 0.452 |
| Difference | +0.042, beyond chance |
The dimensions are real, and the question overlap was hiding them. Shared questions turn out to be spread across dimensions more often than within one, so the overlap inflates the cross-dimension side more, cancelling the genuine within-dimension excess. We had assumed the confound ran the other way and first read the null as ‘cannot tell’. It was masking a real effect.
Stated at its strongest and no further: the separation is small. Three of the five dimensions carry it clearly; two do not. All 25 skills remain strongly related to one another whichever way you cut it, and one broad factor still explains 57% of all the variation. The framework’s structure has support, not proof.
Interconnected, or indistinct?
Everything above treats skills correlating across dimension boundaries as a measurement problem. There is a serious argument that it is not a problem at all.
The framework groups 25 capacities into five dimensions. That grouping is a way of talking about inner development, not a claim that the psyche is built from five separable modules. In practice these capacities are deeply entangled: work on presence tends to surface in courage, and someone who becomes more self-aware rarely stays the same in how they relate to others. On that reading, capacities correlating strongly across the boundaries is not a defect in the instrument. It is the thing being measured.
That reading is attractive, which is exactly why it needs testing. If the instrument were really measuring one quality over and over, the data would look a particular way. It does not. Five checks, each with what the opposite result would have meant:
- 1.One broad quality does not explain everything. The strongest single pattern running through all 25 skills accounts for 57% of the variation between people. 43% comes from somewhere else. Had the instrument been measuring one thing 25 times over, that figure would sit close to 100%.
- 2.Four separate patterns are strong enough to count, not one. The standard rule of thumb asks how many independent patterns are substantial enough to treat as real. One would mean everything collapses into a single trait wearing 25 different labels.
- 3.Nobody’s profile is flat. The gap between a person’s highest and lowest skill averages 26.7 points, and 97% of people differ by more than 15. If people were simply high or low across the board, these profiles would be level lines.
- 4.That unevenness is reliable, not noise. An organisation’s profile shape reproduces at r = 0.91 to 0.98 across three assessments spanning more than two years. A profile that reshuffled between assessments would mean the differences were random.
- 5.When people develop, the capacities move together. In the 293 individuals measured twice, all five dimensions shifted, not one or two. This speaks to the argument most directly: separable modules would show development concentrated where the work was done.
So the instrument is not measuring one undifferentiated quality several times over. There is a strong shared component and stable, reproducible differentiation on top of it. That is the shape you would expect if inner capacities were genuinely interconnected but not identical, which is what the framework has always claimed and what developmental theory predicts.
An interpretation that can absorb any result is not an interpretation, it is an immunity. So it is worth saying what would have counted against this one. If the capacities were genuinely modular, development in those 293 people would have concentrated in one or two dimensions instead of moving all five. If the five groupings were arbitrary, the overlap-free test above would have found no within-dimension excess; it found +0.042. Both could have gone the other way. Neither did.
What this does not rescue. Whichever reading you prefer, the practical caution is unchanged: the 25 skills are too interrelated to be read as independent measures of one person, and a coach working from a single skill score is over-reading the instrument. Interconnection explains why that is so; it does not make it go away. It also does not excuse the question overlap, which is a measurement choice rather than a property of the psyche, and which is still being reduced.
Part two: what calls for caution
None of this says the instrument is broken. It says where a number should be read lightly, and where we do not yet have the evidence.
A single skill score carries little weight on its own
Only 5 of the 25 individual skill scales reach 0.70. That sounds alarming and mostly is not: alpha depends on how many questions a scale has as well as how well they agree, and a six-question scale cannot reach 0.70 here however well written it is. Measured on agreement alone, 19 of the 20 scales that fall short are short rather than incoherent. One is genuinely weak, and it is named in part three.
The overall score and the five dimension scores can carry weight. A single skill score for a single person cannot, on its own. The shape of a profile is more trustworthy than any one number in it, and that is how our reports are written and how coaches should read them.
What this does not establish
- It is a self-assessment. Every score reflects how a person sees themselves. We have no data linking scores to behaviour observed from outside the instrument, and until we do, no claim of that kind should be made on our behalf.
- The 25 individual skills are still not shown to be distinct from each other. The five dimensions separate; the 25 skills within them are a finer cut than the current question set can resolve.
- The change evidence rests on one organisation. 293 paired individuals is a reasonable sample, but it is one client, one sector, one country. No second organisation has yet been assessed twice at a size that could confirm it, so the finding is currently unreplicated. That is the biggest gap.
- Repeat measurement is not neutral. Growing self-awareness can lead someone to rate themselves lower because they see more clearly what a skill involves. A flat result after genuine development work is not necessarily a failure.
- A small repeat assessment cannot settle anything. Individual change varies widely, so detecting a shift of the size we observed needs roughly 75 people measured twice. Under about 20, the numbers cannot carry a verdict whatever they appear to say.
Part three: what we are improving
The genuine defects, and what we are doing about them.
What we are changing
- 1.Reduce question sharing. 83 questions currently produce 221 skill assignments, and only 17 questions belong to a single skill. That overlap did not just obscure the dimension structure, it inverted the test for it.
- 2.Rebuild the one scale that does not hold together. Inclusive mindset is the single scale whose weakness is not explained by its length, and three individual questions do not move with the scale they belong to. Those are the genuine defects the analysis found.
- 3.Report confidence honestly in the product. Overall and dimension scores carry weight; individual skill scores are indicative, and we will make that distinction explicit.
- 4.Build the evidence we do not have. Replicate the change finding elsewhere, and design a study linking scores to something observed from outside the instrument.
- 5.Republish these figures as the population grows. Every number here is generated from the platform data by script, so this can be regenerated rather than rewritten.
Who did this analysis, and what that is worth
The analysis was carried out by an AI system (Claude, made by Anthropic), working directly against the platform database under my direction. I am saying so for the same reason we published the unflattering findings: you should be able to judge the work by how it was done.
What that buys is reproducibility, since every figure traces to a script that can be re-run and nothing was transcribed by hand, and an analyst with no career, grant or citation record riding on the answer. What it does not buy is independence. The work was commissioned by the company that sells the instrument, run on that company’s data, and reviewed by me before publication. An AI engaged by a vendor is not a disinterested third party and I am not going to claim it is.
Nor does it buy freedom from error. Three conclusions reached while producing this paper were wrong:
| First concluded | Why it was wrong | How it was caught |
|---|---|---|
| The five dimensions do not separate; the data cannot settle it | The test was confounded by shared questions, in the opposite direction to the one assumed. | I asked whether the framing was more negative than the evidence warranted. |
| Individuals cannot be matched across repeat assessments | Stated three times from incomplete checks. An employee number was populated for every participant, and it is the basis of this paper’s strongest finding. | I asked for it to be looked at again. |
| A client’s three assessments formed one series | Only two of them did. Treating the third as part of the series produced an apparent decline that was partly an artefact. | I knew the client’s history. |
In every case the error was caught by a person who knew the business, not by the system that made it. So the honest description of what an AI analyst contributed here is speed, consistency and reproducibility inside a process that still needed someone able to say ‘that does not sound right’. Anyone claiming an AI is more objective than a human institution should look at that table first.
We would rather this were checked than believed. If you work on measurement in this field and want to re-run any of it, we will help. The findings we would most like someone else to attack are the change result and the dimension structure.
Read the full study
Everything above is the short version. The white paper sets out how each figure was produced, reports the numbers this page leaves out, states the limitations in full, and shows the workings well enough for someone else to disagree with them properly. Anyone evaluating the instrument, rather than reading about it, should start there.
The Full Analysis
Download the complete white paper
Fifteen pages, including the method, the per-scale figures, the statistical tests behind every claim, and what we are changing as a result.
