Deployment frequency
Elite2.8 / day
deployments per day
From weekly releases to multiple deploys a day, without adding engineers.
Example dashboard — demo data.Synthetic figures built to illustrate the instrumentation. No employer data.
These four boards are worked examples by Alvaro Garcia, built to demonstrate business acumen: what to measure, who to measure it for, and what to conclude from the reading. The numbers are invented; the judgment is the point.
DORA
The four DORA metrics
Each metric is tagged with its published performance band. The bands are computed from the values rather than written next to them, so a label here can never contradict its number.
2.8 / day
deployments per day
From weekly releases to multiple deploys a day, without adding engineers.
18 h
hours, commit to production
Under a day. Trunk-based development and a faster pipeline did this jointly.
4.9%
% of deployments needing remediation
Failure rate fell while deploy frequency quadrupled — the trade-off was not paid.
1.4 h
hours to restore
Just outside the Elite threshold of one hour. Rollback automation is the gap.
Unit: deployments per day, and % of deployments failing
Deploys per day rose 4.7× while the failure rate more than halved. The conventional trade-off between speed and stability was not paid here — smaller, more frequent changes are individually less risky.
| Month | Deploys/day | Change failure rate | Lead time | Time to restore |
|---|---|---|---|---|
| Sep | 0.6 | 12.1% | 72h | 5.4h |
| Oct | 0.7 | 11.5% | 66h | 5.1h |
| Nov | 0.8 | 10.8% | 61h | 4.8h |
| Dec | 0.7 | 12.6% | 68h | 5.6h |
| Jan | 1.0 | 9.8% | 52h | 4.2h |
| Feb | 1.2 | 9.1% | 45h | 3.8h |
| Mar | 1.4 | 8.4% | 39h | 3.4h |
| Apr | 1.6 | 7.8% | 34h | 2.9h |
| May | 1.9 | 7.1% | 29h | 2.5h |
| Jun | 2.1 | 6.4% | 25h | 2.1h |
| Jul | 2.4 | 5.8% | 21h | 1.7h |
| Aug | 2.8 | 4.9% | 18h | 1.4h |
Unit: minutes per run, and % of builds passing
Pipeline duration halved across the year. Roughly two-thirds of the lead-time improvement traces to this rather than to process change — it is usually the cheapest lever and the least funded.
| Month | Pipeline duration | Build success | Test coverage |
|---|---|---|---|
| Sep | 28.4 min | 88.2% | 61.2% |
| Oct | 27.1 min | 89.1% | 62.8% |
| Nov | 26.2 min | 89.7% | 64.1% |
| Dec | 27.8 min | 87.4% | 63.9% |
| Jan | 23.6 min | 90.8% | 66.4% |
| Feb | 21.9 min | 91.6% | 67.8% |
| Mar | 20.4 min | 92.3% | 69.1% |
| Apr | 18.8 min | 93.1% | 70.3% |
| May | 17.2 min | 93.8% | 71.6% |
| Jun | 15.9 min | 94.4% | 72.8% |
| Jul | 14.6 min | 95.1% | 74.1% |
| Aug | 13.2 min | 95.8% | 75.4% |
Secondary delivery signals
Pipeline health and security debt. These are the levers; the four above are the outcome.
13.2 min
minutes per full run · target ≤ 15 min
Halved across the year. Under fifteen minutes is where engineers stop context-switching.
95.8%
% of builds passing · target ≥ 95.0%
Flaky-test quarantine removed most of the noise below 92%.
75.4%
% of lines covered · target ≥ 80.0%
Rising 1.2 points a month. Useful as a direction, dangerous as a goal.
0
count, critical severity · target 0
Zero for three consecutive months; five highs remain in the remediation queue.
Unit: open finding count at month close
Critical findings reached zero three months ago and stayed there. The five remaining highs are the queue that matters; the low count is noise that should never drive a decision.
| Month | Critical | High | Medium | Low |
|---|---|---|---|---|
| Sep | 6 | 24 | 68 | 142 |
| Oct | 5 | 22 | 64 | 138 |
| Nov | 4 | 19 | 61 | 134 |
| Dec | 7 | 26 | 71 | 147 |
| Jan | 3 | 17 | 57 | 129 |
| Feb | 2 | 15 | 54 | 125 |
| Mar | 2 | 13 | 51 | 121 |
| Apr | 1 | 11 | 48 | 118 |
| May | 1 | 9 | 45 | 114 |
| Jun | 0 | 8 | 42 | 111 |
| Jul | 0 | 6 | 39 | 108 |
| Aug | 0 | 5 | 36 | 104 |
Unit: % of days green
Preview environments are always the least stable and that is acceptable — the number to defend is production, which has not been below 99% since December.
| Month | Production | Staging | Preview |
|---|---|---|---|
| Sep | 98.1% | 94.2% | 89.1% |
| Oct | 98.4% | 94.8% | 89.7% |
| Nov | 98.6% | 95.1% | 90.2% |
| Dec | 97.2% | 92.8% | 87.8% |
| Jan | 98.9% | 95.8% | 91.4% |
| Feb | 99.1% | 96.3% | 92.1% |
| Mar | 99.3% | 96.8% | 92.8% |
| Apr | 99.4% | 97.1% | 93.4% |
| May | 99.5% | 97.4% | 93.9% |
| Jun | 99.6% | 97.8% | 94.5% |
| Jul | 99.7% | 98.1% | 95.1% |
| Aug | 99.8% | 98.4% | 95.6% |
How to read this
The numbers above are instrumentation. This is the part that is actually the job — what the pattern means, what it does not mean, and what I would do about it.
Deploy frequency went from 0.6 to 2.8 per day while change failure rate fell from 12.1% to 4.9%. The conventional reading is that speed costs stability; here it bought it, because smaller and more frequent changes are individually less risky. That is the argument to make when someone proposes slowing releases down to be safer.
At 1.4 hours the team sits in the High band, just past the one-hour Elite threshold. Every other DORA metric is Elite. The gap is not detection — MTTA is already under an hour — it is that rollback still requires a human decision. Automating the rollback trigger for failed canaries is the single change that moves this.
Coverage at 75.4% against an 80% target is the metric on this board most likely to be gamed. Coverage that rises while change failure rate also rises means tests are being written to touch lines rather than to catch defects. Here both moved the right way, so the number is currently telling the truth — keep checking that pairing rather than the coverage figure alone.
Lead time fell 75% and pipeline duration fell 54% over the same period. Roughly two-thirds of the lead-time improvement traces to the pipeline getting faster rather than to process change. This is the cheapest remaining lever in most organizations and the one least often funded.
Change failure rate, build success and every environment degraded in December, then recovered. A dashboard where the holiday release freeze is invisible is a dashboard that is smoothing its data. Leaving the spike in place is what makes the rest of the trend believable.