top of page

The Cynefin Framework for Engineering Decisions

Shawn West
3 days ago
7 min read

The proposal to roll out AI test generation across a 300-person quality engineering organisation was a model of its kind. It had a vendor comparison, a twelve-month plan, a training schedule, and a business case that projected a 40% reduction in test-writing effort, based on the vendor's case studies and a two-week trial with one team. The steering committee approved it on that number.


Nine months later, the committee asked whether it was working, and nobody could answer. Some teams used the tool daily; others had quietly stopped. The teams that used it most were producing more tests and spending more time on flaky-test triage. Two teams had found an unexpected use, generating test data rather than tests, which nobody had planned for and the business case didn't measure. The 40% figure couldn't be confirmed or refuted, because the work had changed shape underneath it.


This example is a composite, built from patterns common in enterprise AI tooling rollouts. The figures are illustrative.


Nobody made a careless decision. The proposal was exactly what a well-run organisation asks for. The problem is that it was the right kind of plan for a different kind of problem.


Why the analyse-then-plan approach is attractive


Most engineering leaders are promoted for being good at complicated problems: work where the answer can be found in advance by someone with enough expertise and analysis. Capacity planning, a database migration, choosing between two architectures against known requirements. For those, the right move is to gather data, analyse, decide and execute a plan. Governance processes (business cases, stage gates, projected returns) are built for this kind of work, and they reward people who can produce a confident number up front.


That approach fails, predictably, when it is applied to a problem whose answer cannot be known in advance. Dave Snowden's Cynefin framework, which he developed in 1999 while at IBM and set out for managers with Mary Boone in Harvard Business Review in 2007, exists to make that distinction.


What Cynefin distinguishes


Cynefin (pronounced kuh-NEV-in) sorts situations by the relationship between cause and effect:


Domain

Cause and effect

How to act

Engineering example

Clear (earlier called Simple, then Obvious)

Obvious to anyone

Sense → categorise → respond: apply the known practice

Rotating an expiring certificate with a runbook

Complicated

Discoverable by analysis or expertise

Sense → analyse → respond: get the experts in

Diagnosing a memory leak; sizing a cluster

Complex

Only clear in hindsight; the system adapts as you act

Probe → sense → respond: run experiments, amplify what works

Changing how 300 people write tests; adopting AI tools; team culture

Chaotic

No discernible relationship right now

Act → sense → respond: stabilise first

A production outage with unknown cause

Confusion

You don't know which domain you're in

Break the situation into parts and place each

Most new initiatives, at the start


The critical boundary for engineering leaders is between complicated and complex. Both are hard, and both involve uncertainty. The difference is whether more analysis would reduce that uncertainty.


Why AI adoption is a complex problem


An AI tool rollout looks complicated. There are vendors to compare, features to test and costs to model. But the outcome depends on how hundreds of people change their work in response to the tool, and on how the work then changes in response to them. The test-generation case shows each sign of a complex system:


  • Behaviour changed the thing being measured. More generated tests produced more flaky tests, which changed where time went.

  • Uses emerged that nobody designed. Test data generation came from two teams experimenting, not from the plan.

  • Context mattered more than the tool. The same tool helped one team and was abandoned by another with a different codebase and different habits.

  • The forecast couldn't be checked. A projected percentage assumes a stable baseline, and the baseline moved.


The mechanism is: complex problem → treated as complicated (analyse, forecast, plan) → a confident number commits the organisation to one path → the system adapts in ways the plan didn't anticipate → nobody is measuring the emerging patterns → the initiative can't be judged, so it drifts. This is the same pattern behind the pilot-to-production gap: the model worked in the pilot, and the organisation around it was the unplanned variable.


The discovery move: two questions that place a problem


Before writing a plan, ask two questions about the problem, preferably with the people who'll do the work in the room.


Two questions that place a problem as complicated (analyse and plan) or complex (run safe-to-fail probes)

1. Would two competent experts, given time and the same data, agree on the answer? If yes, the problem is complicated: find the expert and do the analysis. If they'd disagree, and only agree in hindsight about what happened, it's complex.


2. Will people's behaviour change in response to what we do, and change the outcome? If yes, treat it as complex, however technical it looks. Tool adoption, process changes, new quality gates, reorganisations and incentive changes all pass this test.


Most initiatives are mixed, which is what the Confusion domain is for. The test-generation rollout had complicated parts (security review of the vendor, licensing cost, SSO integration) that deserved analysis and a plan, and a complex core (how teams would actually use it) that deserved experiments. The mistake was planning the whole thing one way.


What to do instead in the complex part: safe-to-fail probes


In the complex domain, Snowden's guidance is to run several small, parallel safe-to-fail probes: experiments designed so that failure is contained and tolerable, approaching the problem from different angles. The aim is not for each probe to succeed but to learn how the system responds, then amplify patterns that help and dampen ones that don't.


The test-generation rollout, redone as probes, might have looked like this:


Probe

Team

Amplify if…

Dampen if…

Cost cap / time box

AI writes unit tests for new code

Payments (greenfield service)

Review time per test falls and escaped defects don't rise

Flaky-test rate climbs over two sprints

1 team, 6 weeks

AI writes tests for legacy code with no coverage

Claims (legacy monolith)

Characterisation tests catch a real regression

Engineers rewrite most generated tests

1 team, 6 weeks

AI generates test data only

Data platform

Test-environment setup time falls

Data privacy review flags generated records

1 team, 6 weeks

No tool; the team writes test patterns by hand

Mobile

(Control: the baseline the others are compared against)

—

1 team, 6 weeks


Each probe has signals the team can observe in weeks, not a forecast for the year. The steering committee's question changes from "is it delivering 40%?" to "which patterns should we amplify next quarter?", and that question can be answered.


How to tell which domain your current initiative is being run in


Look at the governance artefacts for something your organisation is doing now.


The forecast test. Does the approval depend on a single projected number, such as a percentage saving or a payback date, for an outcome that depends on how people change their work? If so, it's being governed as complicated.


The signal test. List what is reviewed at each checkpoint. If every item is a milestone (trained, deployed, licensed) and none is an observed pattern (what people are actually doing differently, and what it's producing), nobody can see emergence.


The kill test. Ask what would cause the initiative to be stopped or changed. If nobody can name a signal, there is no dampening mechanism, and the initiative will continue on momentum.


The single-path test. Count the parallel approaches being tried. One path, chosen up front, is a complicated-domain bet on a complex problem.


Where the framework is misused


  • As a sorting hat for avoiding analysis. Calling something complex doesn't excuse skipping the complicated parts. Security, cost and integration still need experts.

  • As a permanent label. Problems move between domains. Once a probe's pattern is stable and understood, it becomes complicated, and then often clear, and should be standardised.

  • As permission for endless experimentation. Probes need time boxes, cost caps and decisions at the end. An experiment with no amplify or dampen criteria isn't safe-to-fail; it's unmanaged.

  • In a crisis. A production outage is chaotic, not complex. Stabilise first; experiment later, in the post-incident review.


Running a decision through it


  1. List the parts of the initiative. Vendor security, cost, integration, usage, measurement, training.

  2. Place each part using the two questions. Expect a mix.

  3. Plan the clear and complicated parts the usual way, with owners and dates.

  4. Design three to five probes for the complex parts, each with amplify and dampen signals, a cost cap and a time box. Include a control.

  5. Review patterns, not milestones, at each checkpoint, and decide what to amplify, dampen or stop.

  6. Move stable patterns out of the complex domain into standards and playbooks.


Final takeaway


The most common way strong engineering leaders get a decision wrong is by applying the method that works for complicated problems to a complex one: analysing, forecasting and committing to a plan when the outcome depends on how people and systems adapt. Cynefin's value is a fast, shared way to notice that, and a different move for the complex part: small, parallel, safe-to-fail probes with clear signals.


Your next action: take one initiative you're running now and apply the four governance tests above. If it fails the forecast and kill tests, split it into parts, keep the plan for the complicated parts, and propose three probes with amplify and dampen signals for the rest at the next steering meeting. To spot when an AI programme has slipped into demonstration rather than adoption, see pilot theater: real adoption vs demo.


Sources


  • Snowden, D. J., Boone, M. E., "A Leader's Framework for Decision Making", Harvard Business Review, November 2007. https://hbr.org/2007/11/a-leaders-framework-for-decision-making

  • The Cynefin Co, "Safe-fail probes". https://thecynefin.co/safe-fail-probes/

  • Cynefin wiki, "Safe to fail probes". https://cynefin.io/wiki/Safe_to_fail_probes

  • "Cynefin framework", Wikipedia: origin (1999, IBM Global Services) and domain naming history (Simple, then Obvious, now Clear). https://en.wikipedia.org/wiki/Cynefin_framework



Part of the Engineering Leadership guide — ShiftQuality's complete map to leading engineers.

bottom of page