DORA Baseline Worksheet: Measure, Diagnose, Re-Measure
The team had moved to daily deploys and was proud of it. Then a bad release took four days to unwind, and nobody could say which of their numbers had predicted it.
The platform team had spent a quarter cutting its CI time and moving from a fortnightly release to deploying most days. Deployment frequency was the metric everyone quoted. When a schema change broke a downstream report, the fix took four days: there was no tested rollback, and the person who knew the migration was on leave. Afterwards the team realised it had been measuring the one number that was improving and none of the ones that weren't.
(Composite example, assembled from patterns that recur in delivery teams adopting DORA metrics. The organisation is not a real one.)
The four DORA metrics only tell you something when you read them together. This worksheet makes you write all four down, read the combination, choose one lever and measure again. The iteration is where the value is; a one-off benchmark mostly produces a slide.
Step 1: record your baseline
Fill in your current number for each metric, then find the band it falls in. The bands below come from the 2021 Accelerate State of DevOps report. DORA changes the thresholds between reports, and its 2025 research describes distributions rather than fixed tiers, so treat the band as orientation and always quote the year with it.
Metric | Your number | Elite (2021) | High (2021) | Medium (2021) | Low (2021) |
Deployment frequency (how often you deploy to production) | ______ | On demand, multiple per day | Once per week to once per month | Once per month to once every six months | Fewer than once every six months |
Lead time for changes (commit to running in production) | ______ | Less than one hour | One day to one week | One month to six months | More than six months |
Change failure rate (share of deploys causing a degradation) | ______ | 0–15% | 16–30% | 16–30% | 16–30% |
Failed deployment recovery time (formerly MTTR) | ______ | Less than one hour | Less than one day | One day to one week | More than six months |
Two notes on reading the table. In 2021, change failure rate separated elite teams from everyone else: high, medium and low shared the same band. And DORA renamed "mean time to restore" to failed deployment recovery time in 2023, narrowing it to recovery from a failed deployment rather than any incident. Use one definition of "deploy" and one of "failure" for every measurement, and write them down before you count.
Two metrics are about throughput (frequency, lead time) and two about stability (failure rate, recovery time). DORA's research has consistently found that high performers do well on both, so improving one at the expense of the other is a warning sign rather than a trade-off.
Step 2: read the combination, not the single numbers
If you see | It usually means | Where to look first |
High deployment frequency and a high change failure rate | The pipeline is outpacing the testing | Pre-merge validation, feature flags, canary releases |
Low frequency and a low failure rate | Safe but slow; work is being batched | Smaller changes, faster CI, fewer manual gates |
High frequency and a long recovery time | You ship fast but can't recover | Monitoring and alerting, runbooks, tested rollback |
All four in good shape | A healthy pipeline | Hold the line and work on the next-weakest |
(These pairings are ShiftQuality's diagnostic heuristics, not DORA findings. Use them to choose where to look, then confirm with your own data.)
My weakest dimension is: ______________________
Step 3: pick one lever for the weakest metric
Choose one, not four. A single change you can attribute is worth more than several you can't.
Weakest metric | Levers to consider |
Deployment frequency | Smaller changes · faster CI · automated tests · trunk-based development · fewer manual approvals |
Lead time for changes | Remove pipeline bottlenecks · reduce manual gates · automate manual tests · parallelise serial steps |
Change failure rate | Stronger automated tests · pre-merge validation · feature flags · canary deployments |
Failed deployment recovery time | Faster detection · runbooks for common failures · tested rollback · rehearse recovery (game days) |
Step 4: re-measure in about three months
Metric | Before | After | Moved? |
Deployment frequency | ______ | ______ | ☐ |
Lead time for changes | ______ | ______ | ☐ |
Change failure rate | ______ | ______ | ☐ |
Failed deployment recovery time | ______ | ______ | ☐ |
Illustrative progression, not a benchmark: weekly deploys, five-day lead time, 20% failure rate and four-hour recovery moving to three deploys a week, one-day lead time, 12% and two hours, after cutting CI from 30 to 10 minutes and adding tests for the most frequently failing change types.
How to use it
Who: the team that owns the pipeline, with its delivery lead. One team per sheet; aggregating across teams hides the signal.
What you need: timestamps you already emit (commits, deploys, incidents). No new tooling for the first pass.
Done when: all four numbers are written down against written definitions, one weakest dimension is named, one lever is chosen, and a re-measure date is in the calendar.
Where this changes
Four rules keep the numbers honest. Track against your own history rather than a leaderboard: moving from medium to high matters more than coasting at elite. Never use the metrics in individual performance reviews, because they are easy to game by splitting commits or avoiding risky work. Treat them as a prompt for a team conversation, not surveillance. And remember what they don't measure: DORA says nothing about whether you are building the right thing, so pair it with product and quality measures.
To see your own levels instantly, the DORA Metrics Calculator does Step 1 for you and names the metric that is setting your ceiling. For the reasoning behind each metric, read DORA metrics: a framework for delivery performance, and for a measurement walkthrough, the DORA four.
Final takeaway
The baseline is not the point; the second measurement is. Write down all four numbers, fix the weakest one on purpose, and measure again. A team that does that twice understands its delivery better than one that has quoted its deployment frequency for a year.
Sources
DORA, Accelerate State of DevOps Report 2021, software delivery performance table, p. 9. https://dora.dev/publications/pdf/state-of-devops-2021.pdf
DORA, A history of DORA's software delivery metrics, on the 2023 rename to failed deployment recovery time. https://dora.dev/insights/dora-metrics-history/
DORA, DORA's software delivery performance metrics, the current metric set and definitions. https://dora.dev/guides/dora-metrics-four-keys/


