A pilot can report that burnout fell, sick days dropped and output held — and still have made things worse for a identifiable group of staff. Averages do that quietly. They blend the office team that gained a genuine Friday off with the frontline rota that absorbed the same work in four longer, harder days, and they present the blend as a success.
If you only publish the average, you will not see the fairness problem until it arrives as a grievance or a resignation cluster.
Why the average is the wrong unit
Most four-day-week pilots report headline movement first: wellbeing up, absence down, revenue broadly stable. Those headlines are useful, but they answer only one question — did the organisation, in aggregate, cope?
They do not answer the questions staff actually ask each other. Did the day off reach everyone, or only the roles that were easiest to schedule? Did part-time staff get a fair share of the benefit? Did carers gain time, or just a denser week? Did managers protect their teams’ day off, or did they protect the workload and let the day erode?
The research base itself points this way. In the UK 2022 trial, 60 percent of staff said it was easier to combine work with care responsibilities — a strong result, and also a reminder that the other 40 percent experienced something else. Follow the subgroup, not just the majority.
The three cuts to run on every people metric
Run every people measure — burnout, wellbeing, actual hours worked, after-hours activity, sick days, resignations and intention to leave — through at least three cuts before you interpret it.
By role and location. Separate frontline or coverage roles from office and remote roles, and managers from individual contributors. This is where a two-tier workforce first becomes visible. If office staff report recovery gains while coverage staff report higher intensity, you do not have a wellbeing success with a scheduling detail attached. You have an unequal intervention.
By caring status, where you can collect it lawfully. Pilots consistently find care is one of the biggest uses of the fifth day, and one of the biggest sources of value. But compressed designs can do the opposite for carers, because a longer working day collides with school and care pick-up times. If carers’ stress rises while non-carers’ falls, the model — not the people — needs adjusting.
By contract type. Compare full-time, part-time, fixed-term and hourly staff. The classic fault line is simple: full-time staff move from 40 to 32 hours on the same pay, an effective hourly uplift, while part-time staff see no change. Left unaddressed, that becomes a pay-equity problem that will outlast the pilot’s goodwill. Several organisations in the published pilots raised part-time hourly rates for exactly this reason.
Where sample sizes allow, add gender and age. Where they do not, say so rather than publishing a cut based on three people.
What to do when a subgroup is worse off
A bad equity cut is not automatically a reason to stop a pilot. It is a reason to redesign before the decision date.
Start by checking actual hours, not contracted hours. If one group’s hours did not really fall, the intervention did not reach them, whatever the policy document says. The dose matters: across pilots, larger real reductions in hours are associated with larger wellbeing gains, so a group with no real reduction should not be expected to show one.
Next, check workload, not effort. If a team’s output held only because intensity rose, the scorecard is recording a transfer — from staff recovery to organisational output — not an efficiency gain. That pattern shows up first in after-hours messages, skipped breaks and a day off that exists on the rota but not in practice.
Then fix the design for that group specifically: a different day-off pattern, staggered cover, a nine-day fortnight instead of a four-day week, additional staffing at peak times, or an equivalent benefit where the same model genuinely cannot work. Document what you changed and re-measure. An equity problem you found, named and fixed is evidence of a well-run pilot. An equity problem you averaged away is a trust problem waiting for a date.
Build the cuts in before you start
Equity cuts cannot be retrofitted if you never collected the baseline. Before launch, agree which subgroups you will report, collect baseline measures for each, and pre-commit to a rule: no group should end the pilot materially worse off on wellbeing or actual hours without a documented redesign response.
Publish the cuts, not just the averages, in your end-of-pilot report — including the uncomfortable ones. Staff can tell whether the day off is fairly shared long before they see your dashboard. A report that shows you looked, and shows what you did about what you found, is the one they will believe.
Sources
- Autonomy / University of Cambridge / Boston College, UK 2022 pilot results, reported 21 Feb 2023 — https://www.sciencedaily.com/releases/2023/02/230221113132.htm
- Fan, Schor, Kelly & Gu (2025), Nature Human Behaviour four-day-week study, reported 22 Jul 2025 — https://www.beckershospitalreview.com/workforce/4-day-workweek-boosts-employee-well-being-study/
- 4 Day Week Global, US and Ireland pilots, Dec 2022 — https://www.internationalworkplace.com/community-zone/news/four-day-week-pioneering-pilot-program-a-huge-success-new-research-reveals
Day5Group is a consultancy that assists in the transition to a 4 day work week. More about Day5Group · All blog articles