The math of stack ranking

What standardized performance distributions are trying to fix, where they break, and why small groups make the math fall apart.

Most federal employees are rated on a five-level scale, from Unacceptable (1) to Outstanding (5). Right now, a supervisor can give any rating to any employee — there's no rule saying only a certain number of people can get a 5. About 64% of federal employees receive one of the top two ratings.

There's a proposal to change that: require agencies to cap how many employees can receive top ratings — a standardized distributionA system that requires performance ratings to follow a predetermined statistical distribution — e.g., only 20% can receive the top rating. Also called "stack ranking" or "forced distribution.", sometimes called stack ranking. For example, a policy might say no more than 20% of employees in a group can be rated "Outstanding." If your group has 10 people, that means at most 2 can get the top mark — regardless of how many the supervisor thinks deserve it.

What the distribution looks like today

3
4
5
1: Unacceptable (0.1%)
2: Minimally Successful (0.3%)
3: Fully Successful (35.1%)
4: Exceeds Fully Successful (21.7%)
5: Outstanding (42.7%)

Levels 1 and 2 are too small to see at this scale.

Here's what the federal performance rating distribution looks like today.

FY 2024, non-SES employees on a five-level rating system (Pattern H).

Nearly two-thirds of the workforce is rated in the top two tiers.

Less than 1% is rated in the bottom two.

Why might this be a problem?

Grade compression

The problem a standardized distribution is trying to solve: how do you identify the top performers? If the top mark goes to nearly half the workforce, the ratings don't have enough resolution to tell you.

It's like a school where 43% of students get an A+ — the grade stops telling you who the strongest students are.

A note on forced attrition

There's another reason organizations sometimes use standardized distributions: to force a certain percentage of employees into the lowest tier, creating a mechanism for firing the bottom tier.

This was famously championed by GE's Jack WelchWelch advocated a "20-70-10" model: the top 20% get rewarded, the middle 70% get developed, and the bottom 10% get managed out. He called it the "vitality curve." as the "vitality curve."

For the rest of this discussion, we're talking about caps at the top — limiting how many people can receive the highest ratings. Not quotas at the bottom.

Rating people in groups creates two separate mathematical problems — one that any rating system faces, and one that's specific to standardized distributions.

Problem 1: The comparison problem

Not all groups are created equal

Some teams are genuinely better than others. Some classes really do have more top students. This isn't a hypothetical — it's the normal state of any large organization or school system.

A standardized distribution assumes every group's performance looks the same. It doesn't.

Under a standardized distribution: a high performer on an already-strong team gets forced into a 3, because top ratings are capped — not because their work isn't excellent, but because too many of their colleagues are excellent too. Your rating isn't a measure of your performance. It's a measure of your performance relative to whoever else happens to be in your group.

Without one: a tough grader on a high-performing team gives out the same "Level 5" as a lenient grader on a weaker team. The ratings look identical, but they don't mean the same thing.

There's no clean answer to this. There aren't really objective, comparable metrics across teams. The work is different, the contexts are different, and no formula can account for that. This problem exists in any system that groups people and then compares within those groups — schools, companies, agencies.

But there's a second problem that's separate and purely mathematical: the size of your group dictates the expected value of your rating.

Problem 2: The small group problem

The size of your group dictates your rating

100 employees

The math of a standardized distribution works fine for large groups — percentages translate cleanly into whole people. But what happens when the rating group is small? When a supervisor only rates eight people, or five, or three?

Say you set a cap: no more than 10% of employees can receive "Outstanding."

Start with 100 people. That's 10 slots — the math translates cleanly.

Shrink the group to 20. That's 2 slots — still a whole number.

Shrink to 10. That's 1 slot.

Shrink to 8.

10% of 8 is 0.8. You can't give 80% of a top rating.

0 slots.

Under a rigid cap, nobody on this team can get the top rating. Not because nobody deserves it — because the math doesn't produce a whole person.

But what if you relax the rule — "always allow at least one"?

Now that team of 8 gets 1 slot. That's an effective cap of 12.5% — not 10%.

The smaller the group, the more generous the relaxation.

Here's the full picture with a 10% cap:

A team of 8 under a rigid 10% cap: 0 slots. Nobody can get the top rating — not because of performance, but because of headcount.

Under a relaxed rule, a team of 3 gets an effective cap of 33% — more than triple the intended rate.

For large groups, the rounding error is small — a group of 52 gets 5 slots, an effective cap of 9.6%. Close enough. But the smaller the group, the bigger the distortion. By the time you're down to single digits, the gap between the intended cap and the actual outcome is enormous.

Now imagine you're an employee.

Under the rigid rule, you'd rather be on a team of 50 than a team of 8 — not because the work is better, but because the team of 8 can't give anyone a top rating.

Under the relaxed rule, you'd rather be on the team of 3, where your odds are three times better than on a team of 50.

It feels unfair: no matter how well you perform, your manager isn't allowed to give anyone on your team a top rating — not because of performance, but because of headcount.

Either way, the system is making a promise it can't keep: that your rating reflects your performance, not your team's headcount.

Look at the zeros. Under a rigid 10% cap, any group smaller than 10 is locked out entirely — no one can receive the top rating.

And it gets worse the tighter the cap. Here's the same table at a 5% cap:

At 5%, any group under 20 is locked out under the rigid rule.

Under the relaxed rule, a group of 5 has an effective cap of 20% — four times the intended rate.

Precision and small groups are fundamentally in tension.

The gaming incentive

Under a relaxed cap, this isn't just an abstract fairness problem. It creates a concrete incentive to fragment teams.

A group of 20 with a 10% cap gets 2 top-rating slots.

Now reorganize those 20 people into four groups of 5. Each group of 5 gets at least 1 slot under the relaxed rule.

4 slots.

Same people. Same work. Double the top ratings.

Managers figure this out. The policy rewards fragmentation. The more generous the relaxation, the more the system undermines itself.

Three options, three costs

There's no clean solution to the small group problem.

Option A

Relax the cap for small groups

Small groups get an advantage. The system incentivizes fragmentation.

Option B

Apply the cap rigidly

Small groups are locked out. Team assignment becomes a career penalty.

Option C

Enforce at a higher level

The person rating you isn't the person managing you.

You have three options. Each has a cost.

Option A: Relax the cap

Let small groups exceed the cap. For example, always allow at least one person in any group to receive the top rating, regardless of group size.

This seems humane — but as we just saw, it incentivizes fragmentation. The more generous the relaxation, the more the system undermines itself.

Option B: Apply the cap rigidly

10% means 10%, rounding down. If your group is too small to produce a whole person at 10%, nobody gets the top rating.

This doesn't just disadvantage employees in small groups — it can defeat the purpose of the system entirely. If enough people are on small teams, the system doesn't produce 10% top ratings overall — it produces far fewer, pushing people into 3s not because of performance but because of headcount.

Remember, the problem we're trying to solve isn't that ratings are too high — it's that they're too compressed to be useful. Forcing high performers into a 3 because their team is too small creates a different kind of compression, not less of it.

Option C: Enforce at a higher level

Instead of rating within teams, rate within larger pools — divisions, directorates, or agency-wide groups.

A pool of 200 people at a 10% cap gives you 20 slots. The math works fine.

But the person deciding how to allocate those 20 slots across the pool may not be the person who manages each employee day-to-day.

In practice, this means a "second reviewer" — someone higher up who distributes ratings across teams they don't directly supervise.

The supervisor who works with you every day thinks you deserve the top rating, but the allocation decision happens two levels up, based on a comparison between you and people doing completely different jobs in completely different offices.

Pick your poison

Every rating system has failure modes.

Without caps: Ratings cluster at the top. A "Level 5" doesn't tell you much when nearly half the workforce has one. Managers can't differentiate, and neither can the system.

With caps: Small groups get locked out or gamed. The person rating you may not be the person who manages you.

The math is the math. No implementation choice makes these tradeoffs disappear — it just determines which ones you live with.