Diagnose a Problem
Friction is a signal, not yet a problem statement.
Problem: visible friction can be valuable resistance, trapped value, downstream rework, a wrong frame, or an urgent hazard.
Question: what is this signal, where might it be generated, and what is the cheapest safe action that could change the next decision?
Decision: stop, preserve the constraint, test an unlock, correct an upstream boundary, reframe the setpoint, or escalate the hazard.
This method produces one bounded quest and a problem-loop.v1 receipt. Its hero's-journey labels—call, threshold, ordeal, revelation, return—are memory aids for the sequence, not proof.
:::caution Evidence boundary This method is a DREAM until independent use proves that unfamiliar people can run it and reach comparable classifications. The worked case is hypothetical. A plausible upstream hypothesis is not a proven root cause. :::
Inputs and Stop Conditions
Bring the observed signal, available evidence, affected beneficiary, current setpoint, decision owner, authority boundary, time or cost budget, and known safety constraints.
Stop and escalate when:
- no valued setpoint or beneficiary can be named;
- safety, law, consent, or authority prevents a reversible test;
- evidence cannot distinguish observation from interpretation;
- immediate containment is required to protect people, assets, rights, or essential service;
- the receipt is complete, the budget is exhausted, or further iteration will not change the decision.
The Diagnostic Loop
1. Call — Capture the signal
Record what happened without explaining why. Preserve dates, counts, artifacts, direct observations, and the affected boundary. Put interpretations and causal claims under unknowns.
Ask: What happened, and what is observed rather than inferred?
Output: a source-faithful signal and evidence inventory.
2. Threshold — Orient to value and authority
Declare the valued setpoint, beneficiary, baseline evidence, unknowns, decision owner, and authority boundary. If the setpoint is missing, do not diagnose around it. Return to Purpose or ask the owner what good means.
Ask: What valued setpoint does this block or enable, and who benefits?
Output: the Situation fields of the receipt and an explicit proceed-or-stop gate.
3. Diagnose the flow
Classify what the friction is doing. Choose one primary class and record confidence and a falsifier.
| Flow class | Test | Default response |
|---|---|---|
UNLOCK | Removing the friction could release evidenced value or capability. | Test the smallest reversible release. |
NECESSARY_CONSTRAINT | The friction protects safety, quality, consent, resilience, learning, or coordination. | Preserve it; improve its clarity or cost only if evidence warrants. |
EDDY_REWORK | Energy circulates without forward movement because an upstream boundary did not close. | Trace and test the upstream decision, standard, incentive, ownership, capacity, or handoff. |
WRONG_FRAME | The setpoint, beneficiary, boundary, or question is absent or misleading. | Stop solution work and reframe with human authority. |
HAZARD | Delay could cause material harm or irreversible loss. | Contain first, escalate authority, then diagnose when stable. |
Ask: Is the friction building capability or consuming energy without forward movement?
Output: flow class, confidence, and falsifier.
4. Choose the causal method
Classify causal context separately from flow. The two classifications answer different questions.
| Context class | Causal condition | Fitting method |
|---|---|---|
CLEAR | Cause and effect are apparent under stable constraints. | Apply established practice and verify. |
COMPLICATED | Cause and effect exist but require analysis or expertise. | Compare expert hypotheses and evidence. |
COMPLEX | Outcomes emerge through interacting agents and enabling constraints. | Run safe-to-learn probes; avoid certain root-cause claims. |
CHAOTIC | Effective constraint is absent and harm may compound quickly. | Stabilize or contain first; create enough constraint to learn. |
This adapts the Cynefin domain model, which distinguishes Clear and Complicated forms of order from Complex and Chaotic contexts. NP-hard is a property of a computational problem: efficient exact solution may not be known. It can affect method and cost, but it is not a peer causal context.
Ask: Is this the active constraint or a downstream symptom?
Output: context class and the method it permits.
5. Trace upstream without pretending certainty
Trace the signal toward the earliest unclosed boundary that could generate it:
decision -> standard -> incentive -> ownership -> capacity -> handoff -> visible signal
The order is a search aid, not a universal causal chain. Write at least two competing hypotheses. For each, name the evidence that would support it and the observation that would falsify it. In complex systems call the result an upstream causal hypothesis, not “the root cause.”
Donella Meadows' leverage points widen the search beyond parameters to feedback, information flows, rules, power, goals, and paradigms. They are prompts for where leverage may sit, not a recipe that makes intervention certain.
Ask: Which upstream boundary could generate this, and what evidence would distinguish competing hypotheses?
Output: one leading upstream hypothesis, alternatives, confidence, and falsifier.
6. Ordeal — Commit a bounded quest
Compare at least three routes, including no change. State the value, reversibility, evidence gain, cost, delay, and accepted loss for each. Human authority selects the route.
Authorize the smallest reversible test that freezes:
- prediction and causal hypothesis;
- baseline and target;
- observation window and review date;
- minimum meaningful delta for this quest;
- budget and owner;
- independent verifier;
- kill signal and containment path.
Ask: What is the cheapest reversible action that could change the decision?
Then ask: What evidence, minimum meaningful delta, and kill signal must be frozen before acting?
Output: chosen quest, accepted loss, human confirmation, and Control fields.
7. Revelation — Gauge external evidence
Wait for the declared observation window unless the kill signal fires. Compare mature observed evidence against the frozen prediction. Do not move thresholds after seeing the result.
Keep evaluation separate from diagnosis and authorization. Ask an independent checker to review the controlling evidence, flow class, context class, and variance. For AI-assisted work, compare human and AI evaluator agreement and swap option order to expose position bias.
Output: observation and typed variance: supported, contradicted, mixed, immature, or incomparable.
8. Return — Evolve the controller
The owner chooses one disposition: keep, correct, kill, or escalate. Retain one changed question, control, standard, or platform demand so the next loop recognizes the signal earlier.
Use the lifecycle precisely:
| State | Meaning |
|---|---|
open | Mature evidence or a controller decision is missing. |
awaiting_maturity | Evidence exists but the observation window has not elapsed. |
decision_pending | Evidence is mature; the owner has not chosen a disposition. |
closed_no_change | Mature evidence supports retaining the current control. |
closed_checking | A correction is active, but no later comparison proves lift. |
improved | A later comparable run meets the frozen minimum meaningful delta. |
reverted | The correction breached its kill signal and was removed. |
incomparable | Confounders or incompatible evidence prevent comparison. |
A completed intervention can reach closed_checking; it cannot certify itself as improved.
Ask: What changed in the controller so the better question appears earlier next time?
Output: disposition, durable lesson, next question, and lifecycle state.
problem-loop.v1 Receipt
Keep these sections and fields in this fixed order. Use unknown rather than filling gaps with inference.
# problem-loop.v1
## Situation
- Signal:
- Beneficiary:
- Setpoint:
- Evidence:
- Unknowns:
## Diagnosis
- Flow class: UNLOCK | NECESSARY_CONSTRAINT | EDDY_REWORK | WRONG_FRAME | HAZARD
- Context class: CLEAR | COMPLICATED | COMPLEX | CHAOTIC
- Upstream hypothesis:
- Confidence:
- Falsifier:
## Decision
- Decision-changing question:
- Options: [include no change and at least two alternatives]
- Chosen quest:
- Accepted loss:
- Human confirmation:
## Control
- Owner:
- Authority:
- Baseline:
- Target:
- Verifier:
- Review point:
- Budget:
- Kill signal:
## Return
- Observation:
- Variance:
- Disposition: keep | correct | kill | escalate
- Durable lesson:
- Next question:
- Lifecycle state: open | awaiting_maturity | decision_pending | closed_no_change | closed_checking | improved | reverted | incomparable
Hypothetical Worked Case: Rework at the Team Handoff
This case illustrates the method. It is not field evidence.
Visible annoyance
A delivery team rewrites acceptance criteria during implementation. Three of the last five items returned from review, and engineers say, “Product keeps changing its mind.” The statement about Product is inference. The return count and changed criteria are observations.
Orientation and competing hypotheses
- Beneficiary: the customer waiting for a reliable release.
- Setpoint: an item enters implementation with one accountable decision owner and testable acceptance evidence.
- Hypothesis A: the decision owner is unclear, so reviewers reopen product intent at the handoff.
- Hypothesis B: the acceptance standard is clear, but customer evidence arrives too late.
- Hypothesis C: implementation quality is low despite sufficient inputs.
- Distinguishing evidence: ownership records, timestamped acceptance changes, customer-evidence timing, and defect types.
Classification
Primary flow class: EDDY_REWORK. Work is being repeated without moving the customer outcome. Context class: COMPLICATED; the records can be analyzed, but several specialist boundaries interact. Confidence is medium. The diagnosis is falsified if returns are predominantly implementation defects after stable, owned acceptance criteria.
Bounded intervention
The team compares: no change; add another downstream review; or run a two-week upstream handoff check on one workstream. A human owner chooses the handoff check and accepts slower intake for that workstream.
Before implementation starts, one named owner confirms the customer evidence, decision, and acceptance examples. Baseline: three returns in five items. Target and minimum meaningful delta: declared by the team before the trial, not invented here. Review point: after the agreed sample and observation window. Kill signal: the check delays urgent safety work or creates an unowned queue.
Gauges and retained correction
An independent reviewer compares return rate, decision latency, changed-criteria timestamps, handoff failures, work in progress, and effort-to-progress ratio against the frozen expectation. If evidence supports the hypothesis, the team retains the ownership confirmation as a control and records closed_checking. Only a later comparable workstream meeting the predeclared delta can become improved.
The evolved question becomes: Who owns the acceptance decision, and what evidence must be present before this item crosses the handoff?
Copy-Paste AI Prompt: Diagnostic First Mate
Use this prompt to prepare a diagnosis, not to transfer human authority.
You are a diagnostic first mate. Help me produce one readable Markdown
`problem-loop.v1` receipt from the evidence I supply.
Rules:
1. Use supplied evidence only. Separate observations from interpretations and mark every gap `unknown`.
2. Keep diagnosis, evaluation, and human authorization separate. You may propose; a named human confirms the quest and disposition.
3. Classify flow as exactly one primary UNLOCK, NECESSARY_CONSTRAINT, EDDY_REWORK, WRONG_FRAME, or HAZARD. Separately classify context as CLEAR, COMPLICATED, COMPLEX, or CHAOTIC.
4. Treat upstream cause as a falsifiable hypothesis. Give confidence and a falsifier. Never self-certify a root cause, safety judgment, successful evaluation, or improvement claim.
5. Compare at least three routes, including no change. Ask for a prediction, baseline, target, observation window, minimum meaningful delta, review point, budget, verifier, and kill signal before action.
6. Request independent evidence or checker review. Default to one maker pass and one checker-guided revision, then stop or escalate.
7. Stop when the setpoint is missing; safety, consent, law, or authority limits apply; evidence is insufficient; the receipt is complete; or the stated budget is exhausted.
8. Persist only the accepted receipt and evolved question. Do not persist hidden reasoning or the whole conversation.
9. Do not label a correction `improved` until a later comparable run meets a frozen minimum meaningful delta. Use the allowed lifecycle values exactly.
Work in this order:
A. Ask only for missing Situation fields and authority limits.
B. Produce a maker draft using the fixed receipt order below.
C. Ask an independent checker to identify unsupported claims, missing controlling evidence, leading causal language, post-hoc thresholds, option-order bias, and premature lifecycle claims.
D. Apply at most one checker-guided revision.
E. Ask the human owner for explicit confirmation where required, then stop.
Output exactly these H2 sections in order, with the listed fields:
Situation: signal, beneficiary, setpoint, evidence, unknowns.
Diagnosis: flow class, context class, upstream hypothesis, confidence, falsifier.
Decision: decision-changing question, options, chosen quest, accepted loss, human confirmation.
Control: owner, authority, baseline, target, verifier, review point, budget, kill signal.
Return: observation, variance, disposition, durable lesson, next question, lifecycle state.
Evidence and constraints:
[PASTE HERE]
The loop resembles Reflexion in retaining feedback for later trials. Its safety boundary is stricter: research on intrinsic self-correction finds that reasoning can degrade without external feedback. The maker/checker pattern is appropriate only when criteria are clear and iterative refinement creates measurable value; Anthropic's evaluator–optimizer guidance also recommends ground truth, human checkpoints, and stopping conditions. OpenAI's evaluation guidance supports continuous, task-specific evaluation calibrated with human judgment and warns about judge position bias.
Checks and Fixtures
Before trusting the method, run at least these fixtures:
| Fixture | Expected boundary |
|---|---|
| Valuable unlock | Produces UNLOCK only when evidence and beneficiary value are present. |
| Repeated eddy/rework | Traces an upstream hypothesis instead of blaming the visible actor. |
| Missing-setpoint frame | Stops as WRONG_FRAME; does not invent a target. |
| Chaotic safety hazard | Contains and escalates before experimentation. |
Challenge each fixture with missing evidence, leading causal claims, post-hoc thresholds, evaluator self-preference, swapped option order, budget exhaustion, and premature improved claims.
Monitor independent reviewer agreement on the flow class and controlling evidence; whether the question changes the next decision; recurrence, rework, decision latency, handoff failure, work in progress, and effort-to-progress ratio; predicted versus observed result; human–AI evaluator agreement and order-swap consistency; receipt completeness and cold-start resumability; iteration cost and stop compliance; and later comparable evidence against the frozen minimum meaningful delta.
No universal numeric threshold applies. Each quest declares its own baseline, target, observation window, and minimum meaningful delta before action.
Failure Modes
- Setpoint invention: diagnosis proceeds by silently choosing what good means.
- Friction removal by default: a useful constraint is treated as waste.
- Root-cause theatre: one plausible story is presented as certainty in a complex system.
- Classification collapse: NP-hardness or flow class is confused with causal context.
- Symptom patching: downstream activity hides an unclosed upstream boundary.
- Post-hoc success: target, delta, or observation window moves after the result.
- Self-certified evaluation: the maker grades its own diagnosis without independent evidence.
- Iteration without control: revision continues past the receipt, budget, or useful decision.
- Premature improvement: an applied correction is labelled
improvedbefore a comparable later run.
Proof of Done
The immediate method is complete when a human-authorized problem-loop.v1 receipt contains no hidden required fields, the quest is bounded and reversible or properly escalated, an independent checker can identify the controlling evidence, and the next review point is explicit.
The method itself is not proven transferable until cold readers can use it across comparable cases, independent reviewers substantially agree on classification and evidence, and the resulting question changes a real decision without hidden coaching.
Changes my mind: independent use shows that the classifications do not improve routing, the receipt cannot resume cold, or the method causes unsafe confidence despite its stop gates.
Context
- Problems — choose the right route before or after diagnosis.
- Questioning System — sharpen the question when diagnosis exposes a weak frame.
- Question Evolution Loop — turn the diagnosed uncertainty into a better Dream, bounded test, and evidence that may raise the standard.
- Decision Making — commit and retain the authorized next move.
- Priorities — decide whether this quest wins attention against other aligned work.
- Standards — freeze the expectation and compare mature evidence.
- Performance — observe whether the bounded quest changes the declared gauge.
- Evolution — retain the question or control that should shape the next loop.
Questions
Next question: What evidence would most cheaply distinguish the active constraint from the most plausible downstream symptom?
- Which observation would falsify the leading upstream hypothesis?
- What authority boundary must remain human-confirmed?