How robust are your causal pathways?

Crossing the Rubicon from a mass of claims to conclusions you can defend

Welcome

A coffee break, so we will take this gently and step by step.

Four things in the next half hour:

What causal mapping is, in plain terms

Two real examples: Love Alliance, and Jewlya Lynn’s work

The nine-step workflow, one step at a time

The real question: how do you make it rigorous?

What is causal mapping?

A way to make sense of what people say causes what.

Collect

Interviews, focus groups, reports: wherever people explain why things changed.

Code

Code each claim:

“The training raised her confidence” becomes a link from one factor to another, with the verbatim quote.

Query

Combine the links from many sources and you have a causal map.

Why this matters even more now

AI lets a single project run to tens or hundreds of thousands of causal claims.

So we have to try even harder to make sure our findings and conclusions are rigorous.

The gap between “we have many claims” and “we have conclusions we can defend” is easy to underestimate. That gap is what today is about.

Causal mapping is not causal inference.

Twenty people, or twenty thousand, saying X influenced Y does not on its own prove that it did.

The app assembles the claims so that you can make the judgement. It does not make the judgement for you. That is the preparatory step for almost any approach, especially theory-based ones like contribution analysis.

Crossing the Rubicon

From a mass of raw claims to a smaller set you can vouch for.

Claims

What sources said, each with its quote. Evidence, not yet weighed.

The crossing

Your quality-assurance work: code, check, bundle, trace, judge.

Conclusions

A defensible answer, with every step on show.

The warranting is always yours. We provide the structures that make it easier, more transparent and more auditable.

Two real examples

Love Alliance: causal mapping at scale

A five-year partnership for the health and rights of key populations affected by HIV, across ten African countries.

Southern Hemisphere led the end-term evaluation; we supported the causal-mapping strand.

AI did the low-level coding across 176 documents: over 22,000 causal claims, 13,756 kept for the maps, and a quote on every link.

Code wide and cheap first, judge later.

Maps at three scales, the ten countries, the regions and a global picture, and cut by theme.

When a worry surfaced that the programme’s own advocacy might be provoking a backlash, we put the question to the data and traced it through the sources, quote by quote. The evidence pointed the other way.

Jewlya’s retrospective

A ten-year, systems-change retrospective, some of you saw her present it at April’s Coffee Break.

She coded every single claim separately, each with its own quote.

Coterminal claims, the ones saying the same X influenced Y, were gathered into bundles.

Each bundle was scored against a five-level rubric: over 350 causes and 150 connections, every one verified against the evidence.

The final map was displayed in Kumu, where each bundle shows as a single link. One clean arrow on the screen can stand for many separately coded claims underneath.

Nine steps, three tasks

Collect

1 Start from the question

2 Gather the data

Code

3 Manage the codebook

4 Code each claim

5 Check and tag

Query

6 From claims to bundles

7 From bundles to pathways

8 Value and alternatives

9 Holistic judgement

A few wide, cheap passes to capture the evidence, then steadily narrower judgement. We will now walk all nine.

Collect

Step 1: start from the question

Write down what you want to be able to say at the end, and to whom.

Good questions for the method:

  • Which factors matter most
  • What influences or follows from a factor
  • How different groups see things
  • How well the evidence supports a pathway or theory of change

It will not give you effect sizes, and on its own it does not prove X causes Y. Sketch the map that would answer your question; that sketch is your target.

Step 2: gather the data

The question decides the data.

Narrative material works best: ask people what changed and why, and you get causal claims to code. QuIP-style “stories of change” are gathered in exactly this way.

If you will want to compare women and men, or staff and clients, or early and late, those groups have to be in the data and recorded, so the comparisons are possible later.

Code

Step 3: the codebook

How tightly are the labels fixed in advance?

Free

Start from nothing and let the source’s own terms emerge. Finds more, leaves more to tidy.

Fixed

Start from a set codebook, such as a theory of change. Cleaner, but misses links.

Most projects are somewhere between, and you revise as you go.

QA · Are the labels consistent and at the right grain? Whose world view do they encode?

We code claims, not facts

A coded link means: there is evidence that this source claims X influenced Y.

Not that X really did influence Y. Twenty people saying so is not proof. It is evidence you can now weigh. Crossing from claims to conclusions is your job, and the back half of this workflow is about doing it well.

The one rule you never break

1

verbatim quote for every single link, no exceptions.

Without it you are no longer showing your working, and you cannot justify the conclusions you draw. The app does not enforce this, so you have to.

QA · Two coding targets: precision (are the links right?) and recall (did you miss any?). Tune the instruction on a small, varied sample until both hold.

Query

Step 6: from claims to bundles

The core quality-assurance move, and the answer to question one: how do you interrogate a mass of evidence?

A bundle is every claim that says the same X influenced Y. Weigh each one: how many sources, how convincing, do they agree or pull apart?

You can record the verdict by collapsing a bundle into a single assessed link carrying your quality score. The app will not let you, by hand or with AI, until you have written your criteria into a rubric first, on purpose. Jewlya used a five-level scale.

QA · Write the rubric before you judge. Plausibility, uniqueness and triangulation are a usable set of criteria.

From a mass of claims to a set you vouch for

1000

raw claims

30

bundles

25

assessed links: a much cleaner basis for argument

Step 7: from bundles to pathways

The transitivity trap, and the answer to question three: how do you compare alternative explanations across a pathway?

A pig farmer says:

the cash grant gave me more cash

A wheat farmer says:

more cash let me buy more seed

So cash grants lead to more seed?

No. One source said A to B, another said B to C. Stitched into one story A to B to C that nobody actually told.

The single most important pitfall of any causal diagram.

QA · A pathway is only as warranted as the within-source stories that run end to end.

Path tracing alone will not save you

Path tracing shows every link on a route between two factors, across all sources. Easy to misread as a story someone actually told.

Step 8: value and alternatives

How much did it matter, and compared to what?

Judging value and relative contribution is central evaluation territory, covered extensively by John Mayne and others. We lean on QuIP for value.

The discipline: compare the influence you care about against rival explanations on the same map, not in isolation. Count the sources whose narratives actually run from your driver to your outcome.

Step 9: holistic judgement

Finally you draw the conclusion. This is the answer to question two: which weak or doubtful claims still hold up when you look at everything together.

Behind a single tidy map there may be hundreds of quotes. Does the overall claim hold? Do the links in every pathway really belong to the same context?

The AI vignette can take exactly these questions: is each link part of a coherent, complete story from source factor to target factor? It does only what a patient reader could, so treat its draft as a starting point and edit it.

So how do you make it rigorous?

It overlaps with what you already do

The whole field takes one question seriously: how do you assess the strength of evidence behind a causal claim?

Built-in tests

Process tracing weighs each link with hoop and smoking-gun tests.

Built-in story

Contribution analysis builds and tests a contribution story.

Written rubric

Where no test is built in, a rubric agreed in advance does the same job.

Our bundle rubric in Step 6 is exactly this device, like the CLARISSA and Jewlya Lynn seafood-retrospective rubrics. A workable set of rubric criteria: plausibility, uniqueness, triangulation. None of these removes the final judgement; they make it transparent.

Causal mapping organises the evidence, you warrant it

At no point does the causal mapping move on its own from claims to facts.

What it provides

Tags, columns, the assessed-link switch, source tracing, vignettes: structures that make warranting easier and auditable.

What it does not provide

An engine that turns “twenty people said so” into “therefore it is so”.

None of this is causal inference

Not in a statistical sense. It is a disciplined way to assemble evidence, weigh it transparently, and reach conclusions you can defend.

We use this every day in our consultancy at Causal Map Ltd, and it keeps evolving. If you want to go on this journey with us, get in touch.

Companion working papers in the Causal Map Garden: “A workflow for causal coding” and “Quality assurance at each step”. App: app.causalmap.app

Home