How robust are your causal pathways?
Crossing the Rubicon from a mass of claims to conclusions you can defend
Causal Pathways Coffee Break · 16 June 2026 · Steve Powell & Gabriele Caldas Cabral, Causal Map Ltd
A coffee break, so we will take this gently and step by step.
Four things in the next half hour:
What causal mapping is, in plain terms
Two real examples: Love Alliance, and Jewlya Lynn’s work
The nine-step workflow, one step at a time
The real question: how do you make it rigorous?
A way to make sense of what people say causes what.
Collect
Interviews, focus groups, reports: wherever people explain why things changed.
Code
Code each claim:
“The training raised her confidence” becomes a link from one factor to another, with the verbatim quote.
Query
Combine the links from many sources and you have a causal map.
A network of what people believe drives what, built from many sources. Every arrow is backed by quotes you can read.
AI lets a single project run to tens or hundreds of thousands of causal claims.
So we have to try even harder to make sure our findings and conclusions are rigorous.
The gap between “we have many claims” and “we have conclusions we can defend” is easy to underestimate. That gap is what today is about.
Causal mapping is an evidence broker. It gathers and organises causal claims so that the methods you already use can make the judgement: contribution analysis, process tracing, Outcome Harvesting, realist evaluation, QuIP, Most Significant Change. Most real evaluations combine several.
Twenty people, or twenty thousand, saying X influenced Y does not on its own prove that it did.
The app assembles the claims so that you can make the judgement. It does not make the judgement for you. That is the preparatory step for almost any approach, especially theory-based ones like contribution analysis.
From a mass of raw claims to a smaller set you can vouch for.
Claims
What sources said, each with its quote. Evidence, not yet weighed.
The crossing
Your quality-assurance work: code, check, bundle, trace, judge.
Conclusions
A defensible answer, with every step on show.
The warranting is always yours. We provide the structures that make it easier, more transparent and more auditable.
A five-year partnership for the health and rights of key populations affected by HIV, across ten African countries.
Southern Hemisphere led the end-term evaluation; we supported the causal-mapping strand.
AI did the low-level coding across 176 documents: over 22,000 causal claims, 13,756 kept for the maps, and a quote on every link.
Code wide and cheap first, judge later.
Maps at three scales, the ten countries, the regions and a global picture, and cut by theme.
When a worry surfaced that the programme’s own advocacy might be provoking a backlash, we put the question to the data and traced it through the sources, quote by quote. The evidence pointed the other way.
A ten-year, systems-change retrospective, some of you saw her present it at April’s Coffee Break.
She coded every single claim separately, each with its own quote.
Coterminal claims, the ones saying the same X influenced Y, were gathered into bundles.
Each bundle was scored against a five-level rubric: over 350 causes and 150 connections, every one verified against the evidence.
The final map was displayed in Kumu, where each bundle shows as a single link. One clean arrow on the screen can stand for many separately coded claims underneath.
A map like this in Kumu. Each link looks simple, but behind it is a bundle of individually coded claims, each with its source and quote. We will now follow that same path, step by step.
Collect
1 Start from the question
2 Gather the data
Code
3 Manage the codebook
4 Code each claim
5 Check and tag
Query
6 From claims to bundles
7 From bundles to pathways
8 Value and alternatives
9 Holistic judgement
A few wide, cheap passes to capture the evidence, then steadily narrower judgement. We will now walk all nine.
Write down what you want to be able to say at the end, and to whom.
Good questions for the method:
It will not give you effect sizes, and on its own it does not prove X causes Y. Sketch the map that would answer your question; that sketch is your target.
The question decides the data.
Narrative material works best: ask people what changed and why, and you get causal claims to code. QuIP-style “stories of change” are gathered in exactly this way.
If you will want to compare women and men, or staff and clients, or early and late, those groups have to be in the data and recorded, so the comparisons are possible later.
How tightly are the labels fixed in advance?
Free
Start from nothing and let the source’s own terms emerge. Finds more, leaves more to tidy.
Fixed
Start from a set codebook, such as a theory of change. Cleaner, but misses links.
Most projects are somewhere between, and you revise as you go.
QA · Are the labels consistent and at the right grain? Whose world view do they encode?
Each causal claim becomes one row: a link from one factor to another, with its source and a verbatim quote. By hand or with AI, the unit is the same.
A coded link means: there is evidence that this source claims X influenced Y.
Not that X really did influence Y. Twenty people saying so is not proof. It is evidence you can now weigh. Crossing from claims to conclusions is your job, and the back half of this workflow is about doing it well.
1
verbatim quote for every single link, no exceptions.
Without it you are no longer showing your working, and you cannot justify the conclusions you draw. The app does not enforce this, so you have to.
QA · Two coding targets: precision (are the links right?) and recall (did you miss any?). Tune the instruction on a small, varied sample until both hold.
Some links will be wrong, so check them. Tag a doubtful or surprising claim (#doubtful, #surprising) so you can filter it in or out later.
QA · Conviction codes how sure the source sounds, not how strong the link is. And unmarked means not mentioned, not medium.
The core quality-assurance move, and the answer to question one: how do you interrogate a mass of evidence?
A bundle is every claim that says the same X influenced Y. Weigh each one: how many sources, how convincing, do they agree or pull apart?
You can record the verdict by collapsing a bundle into a single assessed link carrying your quality score. The app will not let you, by hand or with AI, until you have written your criteria into a rubric first, on purpose. Jewlya used a five-level scale.
QA · Write the rubric before you judge. Plausibility, uniqueness and triangulation are a usable set of criteria.
The five-level scale from Jewlya Lynn’s retrospective. A finding had to reach at least Level 3, multiple sources from distinct perspectives, to be included. The criteria are written down first, then every bundle is judged against them.
1000
raw claims
30
bundles
25
assessed links: a much cleaner basis for argument
The transitivity trap, and the answer to question three: how do you compare alternative explanations across a pathway?
A pig farmer says:
the cash grant gave me more cash
A wheat farmer says:
more cash let me buy more seed
So cash grants lead to more seed?
No. One source said A to B, another said B to C. Stitched into one story A to B to C that nobody actually told.
The single most important pitfall of any causal diagram.
QA · A pathway is only as warranted as the within-source stories that run end to end.
Path tracing shows every link on a route between two factors, across all sources. Easy to misread as a story someone actually told.
The conservative move: keep only sources whose own account runs all the way from A to C. Every link is then part of at least one complete story told by one person, and you can review the evidence source by source.
How much did it matter, and compared to what?
Judging value and relative contribution is central evaluation territory, covered extensively by John Mayne and others. We lean on QuIP for value.
The discipline: compare the influence you care about against rival explanations on the same map, not in isolation. Count the sources whose narratives actually run from your driver to your outcome.
Finally you draw the conclusion. This is the answer to question two: which weak or doubtful claims still hold up when you look at everything together.
Behind a single tidy map there may be hundreds of quotes. Does the overall claim hold? Do the links in every pathway really belong to the same context?
The AI vignette can take exactly these questions: is each link part of a coherent, complete story from source factor to target factor? It does only what a patient reader could, so treat its draft as a starting point and edit it.
The whole field takes one question seriously: how do you assess the strength of evidence behind a causal claim?
Built-in tests
Process tracing weighs each link with hoop and smoking-gun tests.
Built-in story
Contribution analysis builds and tests a contribution story.
Written rubric
Where no test is built in, a rubric agreed in advance does the same job.
Our bundle rubric in Step 6 is exactly this device, like the CLARISSA and Jewlya Lynn seafood-retrospective rubrics. A workable set of rubric criteria: plausibility, uniqueness, triangulation. None of these removes the final judgement; they make it transparent.
At no point does the causal mapping move on its own from claims to facts.
What it provides
Tags, columns, the assessed-link switch, source tracing, vignettes: structures that make warranting easier and auditable.
What it does not provide
An engine that turns “twenty people said so” into “therefore it is so”.
Not in a statistical sense. It is a disciplined way to assemble evidence, weigh it transparently, and reach conclusions you can defend.
We use this every day in our consultancy at Causal Map Ltd, and it keeps evolving. If you want to go on this journey with us, get in touch.
Companion working papers in the Causal Map Garden: “A workflow for causal coding” and “Quality assurance at each step”. App: app.causalmap.app