Module 6 — Metrics That Matter · Lesson 6.2
People Metrics
Capability fit, throughput, growth, and load — without making it a performance hammer
~12 min
What you'll learn
- Read the people metrics on /portfolio Resources and /reports
- Distinguish capability metrics from utilization metrics
- Read the load distribution as a check-in prompt, not a burnout forecast
- Avoid the failure mode of using people metrics as ranking
People metrics tell you whether the team's capability model is accurate, whether the assignment system is producing growth, and whether load is sustainable. They are also the most dangerous class of metrics — used badly they become a leaderboard, which corrupts the data within a quarter. This lesson is the six to watch and how to read them without making the team game them.
Capability fit — is assignment producing matched work?
The single most useful people metric is capability fit: across recent assignments, how well does the assignee's skill profile match the task's skill tags?
Kavanah computes this as the mean of the assignee's strength across the task's skill tags, counting a skill they have never touched as a zero. A score near 1.0 means the assignee is well-matched; a score near 0.5 means the assignee is plausible but not ideal; below 0.4 means the assignment was a stretch.
The zero is the load-bearing part, and it encodes a deliberate choice: averaging in a zero drags a specialist who covers half the required skills below a generalist who covers all of them weakly. That is usually what you want from a router — a task needing frontend and copywriting should not go to the best frontend engineer in the company if they have never written a line of copy. But it means the score is an average, not a veto: a candidate missing one skill is penalized, not excluded. When the missing skill is the one that actually matters, you are the one who has to notice.
We recommend treating roughly 0.6–0.8 as a healthy starting band and adjusting it against your own assignment outcomes — the specific cutoffs are Kavanah defaults, not empirical constants. Below 0.6 across the board means assignment is defaulting to availability or social proximity, not capability. Above 0.85 across the board means you are over-typecasting — every task goes to the most obvious person — and nobody is growing.
The action when capability fit drifts low: audit the recommend_assignees acceptance rate. The recommender is probably surfacing better matches and being overruled. Find out why.
Stretch rate — are people getting growth assignments?
Capability fit measures match quality; stretch rate measures growth opportunity. It is the fraction of assignments per member where the assignee was not the top capability recommendation but was a deliberate stretch — usually with a stronger member available to consult.
We suggest starting from a target of roughly one assignment in five being a deliberate stretch, then calibrating against your own team's outcomes — the right fraction depends on your mix of routine versus novel work and how much senior slack you have to support stretches. Treat this as a Kavanah starting recommendation, not a measured industry norm; the general idea that people grow through stretch draws on the '70-20-10' development tradition, whose exact ratios have little empirical backing (Clardy 2018). Most assignments go to the best match (efficiency), but a meaningful slice goes to a stretch (growth). Zero stretch produces a team that ships fast and stagnates; too much produces a team that grows fast and ships unevenly.
Kavanah surfaces stretch rate per member: 'Sarah Chen — 22% stretch in the last quarter, all in skills where she's now climbing from Proficient to Expert.' That is a useful coaching artifact. It is not a performance metric — high stretch and high stretch-task completion both look healthy.
Throughput per member — with skill weighting
Raw task count per member is a misleading metric because tasks are not interchangeable. A senior contributor shipping three load-bearing platform tasks per sprint is producing more value than the same person shipping ten cleanup tasks.
The useful version is throughput weighted by task estimate: how many estimated-hours of work each member shipped per sprint. This is comparable across members in the same way a sprint plan is comparable across members.
The healthy signal is stable per-member throughput sprint over sprint, with seasonal variation matching capacity (leave, focus blocks). The unhealthy signal is a downward trend on one member with no explanation, which is usually the early sign of burnout, disengagement, or a brewing performance issue. The first action is a 1:1 conversation, not a metric review.
Do not publish raw throughput-per-member as a leaderboard. It will corrupt the data within a quarter (members game the easy tasks). It is a 1:1 input, not a public scoreboard.
Active load distribution — a load-health signal
A useful load-health signal is the distribution of concurrent active tasks across the team — in particular the p90 of active tasks per member, which shows who is carrying the most simultaneous work.
Why watch it: high concurrent load raises the cost of context-switching (interrupted work gets finished faster but at the cost of more stress, frustration and effort — Mark, Gudith and Klocke 2008), and heavy workload is one of the six organizational drivers of burnout the research identifies (Leiter and Maslach 1999; Demerouti et al. 2001). But be careful how much you load onto this number. Active-task count is not, on its own, a validated burnout predictor. Burnout is a chronic-stress syndrome shaped by workload and control, reward, community, fairness and values together (WHO ICD-11, 2019) — not a reading you can take from a task counter. So read a rising p90 as a prompt to check in, not as a forecast of who will burn out next month.
The thresholds we suggest — roughly 1–2 as comfortable, 3 as worth redistributing, 4+ as urgent — are Kavanah's starting bands, not measured limits. Calibrate them against your own team's baseline. The Resources tab on /portfolio surfaces the distribution as a horizontal bar with the p90 visible at a glance; when it climbs, redistribute, and use the availability requests and clean reassignment the system provides.
Skill growth — are people getting more capable over time?
The reinforcement loop produces a measurable signal: how many members crossed a proficiency threshold (Learning → Proficient, Proficient → Expert) on at least one skill in the last quarter.
A team where every active member is growing on at least one skill per quarter is a team that is learning. A team where most members are flat over a quarter is either not getting stretch assignments, not getting their skill tags reinforced (untagged tasks), or not getting the kinds of work that produce growth.
The action when growth flattens is to look at the assignment patterns. If everyone is doing what they already know, the team is optimizing for short-term throughput at the cost of long-term capability. Increasing stretch assignments is the usual lever, but treat the size of the increase and the time to see an effect as things you learn from your own team's history — there is no evidence for a specific 'X% in Y quarters' rule.
Idle ratio — the soft signal you usually miss
Idle ratio is the fraction of recent days where a member had no active task and no scheduled focus block. It is a soft signal — sometimes idle means resting, sometimes it means stuck, sometimes it means the system isn't routing work to them.
We suggest treating steady-state idle ratios under about 10% as unremarkable and investigating sustained spikes on any one member — the 10% figure is a Kavanah starting guideline to tune to your team, not a benchmark. A spike on one member (say, 30% idle over two weeks) is a coaching signal: ask what is going on. The default assumption should be 'I am routing them wrong' before 'they are slacking.' The former produces a system fix; the latter produces a defensive employee.
Kavanah surfaces the idle ratio on the Resources tab. The intent is not to police it. It is to make visible the case where the system is failing a member — assignments not being routed their way, capabilities under-used — before the member feels obligated to escalate.
Wire the people-metrics view
- 1
The capability-fit, throughput, active-load distribution, and idle-ratio views are all there. Confirm the data is populating.
- 2
Look at active-load p90
If it is at 3, plan a redistribution. If 4+, do not wait for next sprint.
- 3
Check stretch rate per member
Anyone at 0% stretch in the last quarter is on coasting trajectory. Set up a coaching conversation, not a metric review.
- 4
Keep throughput-per-member off public dashboards
It is a 1:1 input. Public throughput leaderboards corrupt the data within a quarter.
People metrics
- Capability fit (mean)
- Geometric mean of assignee skill strength × task skill tags, across recent assignments.
- Healthy signal: 0.6–0.8.
- Stretch rate per member
- Fraction of assignments that were a deliberate non-top match.
- Healthy signal: ~20% as a starting recommendation; calibrate to your work mix.
- Weighted throughput per member
- Estimated-hours of work shipped per sprint, by member.
- Healthy signal: Stable. Watch trends, not absolute levels.
- Active load p90
- p90 of in-progress task count across the team.
- Healthy signal: ≤ 3 as a Kavanah starting band. A prompt to check in, not a validated burnout predictor.
- Skill-threshold crossings per quarter
- Members who moved up a proficiency level on at least one skill.
- Healthy signal: Most active members per quarter.
- Idle ratio per member
- Fraction of recent working days with no active task and no scheduled focus block.
- Healthy signal: < 10% steady-state as a starting guideline; spikes worth a 1:1.
Key takeaways
- ·Capability fit + stretch rate together describe the team's growth trajectory.
- ·Active-load p90 is a practical load-health signal worth watching — a prompt to check in, not a validated burnout predictor.
- ·Skill-threshold crossings are the proof the reinforcement loop is producing learning.
- ·Throughput-per-member is a 1:1 input; never publish as a leaderboard.
People metrics describe the team. Work metrics describe what the team is producing. The next lesson is the work side of the dashboard.
Sources
- 1.The SPACE of Developer Productivity: There's more to it than you think
Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Brian Houck, Jenna Butler · ACM Queue 19(1), 20–48 · 2021
Productivity cannot be reduced to a single metric or measured on individuals — the basis for keeping throughput-per-member a 1:1 input, never a leaderboard.
- 2.Assessing the Impact of Planned Social Change (Campbell's Law)
Donald T. Campbell · Public Affairs Center, Dartmouth College · 1976
A quantitative indicator used for decision-making is subject to corruption pressure — why a public throughput leaderboard corrupts the data within a quarter.
- 3.The Cost of Interrupted Work: More Speed and Stress
Gloria Mark, Daniela Gudith, Ulrich Klocke · Proceedings of ACM CHI 2008 · 2008
Interrupted work is finished faster but at the cost of more stress, frustration and effort — the mechanism behind watching high concurrent load.
- 4.Burn-out an 'occupational phenomenon': International Classification of Diseases (ICD-11)
World Health Organization · who.int · 2019
Burnout is a syndrome resulting from chronic, unmanaged workplace stress — not a function of in-progress task count. Why p90 load is a prompt, not a predictor.
- 5.The Job Demands–Resources Model of Burnout
Evangelia Demerouti, Arnold Bakker, Friedhelm Nachreiner, Wilmar Schaufeli · Journal of Applied Psychology 86(3), 499–512 · 2001
Burnout is predicted by job demands and (lack of) resources together — the actual predictor structure, of which workload is only one part.
- 6.Six areas of worklife: a model of the organizational context of burnout
Michael P. Leiter, Christina Maslach · Journal of Health and Human Services Administration 21(4) · 1999
The established organizational drivers of burnout: workload, control, reward, community, fairness and values — workload is one of six.
- 7.State of the Global Workplace: 2026 Report
Gallup · gallup.com · 2026
Current engagement and stress baselines (20% engaged; 40% experienced a lot of stress the prior day) — real figures to reason with instead of an invented threshold.
- 8.70/20/10 model (learning and development) — origins and the Clardy (2018) critique
reference summary of Alan Clardy, Human Resource Development Review · en.wikipedia.org · 2018
The stretch-and-grow tradition rests on the 70-20-10 model, whose specific ratios are 'a fact in search of evidence' — why the stretch-rate number is framed as guidance, not a norm.