Consider how a research finding travels. It might begin life in a mixed-methods evaluation of a livelihoods programme where a household survey conducted across two districts, with a modest sample, suggests that women took a greater part in household financial decisions over the course of the programme. The evaluation team reports this with appropriate care, noting the size of the sample, the reliance on self-reported measures, and the absence of a comparison group.

Six months later, the same finding appears in the implementing organisation’s annual report as evidence that the programme increased women’s economic empowerment. A year on, it is cited in a donor strategy as part of the rationale for extending the model to three further countries. By then the caveats have fallen away. No one set out to mislead; each person who passed the finding along softened it a little, and at no point was anyone charged with asking whether the original claim could bear the weight now placed upon it.

This example is a composite, drawn from patterns rather than any single report, yet I suspect that anyone who has worked with evidence in humanitarian and development settings will recognise its shape. It speaks to a structural gap in how our sector produces and relies upon knowledge, and one that we discuss far less often than we should.

A different standard of scrutiny

Academic research, whatever its shortcomings, has a checkpoint built into its process. Before an article appears in a journal, reviewers with relevant expertise are asked whether the methods are appropriate, the analysis is sound, and the conclusions follow from the evidence. Peer review is slow and uneven, and those who have been through it rarely describe it with affection, but it exists, and its existence shapes how researchers write.

Much of the evidence that actually informs humanitarian and development decisions, however, never passes through a journal. It takes the form of evaluations, baseline studies, needs assessments, learning reviews, policy briefs, and research reports commissioned by governments, NGOs, donors, UN agencies, and think tanks, and it is this body of work that guides where funding is directed, which programmes continue, and how policy is framed. Some larger agencies have developed thoughtful quality assurance systems for their evaluations, yet a considerable share of commissioned research sits outside any such arrangement. A report may be read by a programme manager, a technical adviser, a communications colleague, and perhaps a reference group, each of whom brings something of value. What is frequently absent is a reader whose specific responsibility is to ask whether the methodology supports the claims that rest upon it.

None of this, in my experience, is a matter of carelessness. Research timelines in our sector are compressed, and review is often the first thing to be sacrificed when a deadline slips, particularly when it was never budgeted in the terms of reference. Those who commission research are usually experts in their thematic field rather than in research methods, and it would be unreasonable to expect a protection specialist, for instance, to judge whether a sampling strategy can sustain a generalised claim or whether a qualitative analysis has been read too generously. Where internal review does take place, it is often carried out by colleagues with some stake in the findings, whether through the success of the programme, the standing of the organisation, or a working relationship with the research team. Nor would academic peer review serve us well if simply transplanted, since it moves too slowly for operational decisions and its reviewers are not always familiar with the realities of collecting data in crisis settings or without university frameworks.

The gap is also widening. As localisation shifts greater responsibility towards national and local organisations, and as donors ask for more evidence of results, more research is being produced by more organisations, often without dedicated research support. Increasingly, that research is also prepared with the assistance of AI tools, which can be genuinely useful for drafting and synthesis, but which have a particular capacity to lend weak reasoning an air of assurance. A report can now be fluent, well organised, and polished while its conclusions travel some distance beyond what its data can support. Polish was once a rough proxy for care, and it can no longer be relied upon in that way.

What independent review might offer

The consequences of unexamined evidence are seldom visible at publication, which is perhaps why they are so easily discounted. They emerge later, when programmes are scaled on uncertain foundations, when donors grow quietly less confident in the evidence base, or when methodological weaknesses come to light after a report has already been widely cited and the damage to an organisation’s reputation is harder to contain. There is a further cost that I think deserves more attention than it receives. The people who gave their time as research participants, frequently in difficult and precarious circumstances, did so in the expectation that their contributions would be used well, and poorly founded findings fall short of that expectation in a way that is an ethical failing as much as a methodological one.

I do not believe the remedy lies in importing the apparatus of academic publishing into the humanitarian sector, but perhaps in borrowing its intent. What seems valuable is review that is independent, proportionate, and undertaken before publication, while there is still an opportunity to strengthen the work rather than merely to judge it. Such review asks a fairly settled set of questions: whether the methods fit the research questions, the evidence supports the findings, the conclusions and recommendations follow from those findings, limitations are acknowledged honestly and carried through to the summary, and the report will be understood by the audience for whom it is intended. It is most useful when it is constructive in spirit. In the great majority of outputs I have reviewed, there is a sound core of work, and the task is less one of finding fault than of bringing the claims into closer alignment with the evidence. It also depends on independence, in the simple sense that the reviewer should have no stake in the findings and no relationship with the programme under examination.

Much can be done without great expense. Commissioners might build independent review into the terms of reference from the outset, so that it is protected when timelines tighten, and might ask for a clear account of methods and an honest discussion of limitations, checking that those limitations survive into the executive summary. It can be illuminating to ask a research team, before sign-off, what evidence would have led them to a different conclusion. It is also worth being clear about where evidence ends and advocacy begins: each has its place, but both are weakened when they are confused. Those who produce research can do something similar by inviting a critical friend from outside the project to read a draft purely for the fit between its claims and its evidence, setting aside questions of style and messaging.

No single study can hope to capture every truth about a complex humanitarian or development situation, and that has never been the standard to which research should be held. The more modest, and in some ways more demanding, standard is that research should be true to the claims it makes and to the people whose time and experience made it possible. At present, a good deal of the evidence on which our sector relies reaches decision-makers without anyone having asked whether it meets that standard. Closing the gap does not require a new layer of bureaucracy so much as a change in how we regard review, from an occasional luxury to an ordinary and budgeted part of producing evidence that deserves to be trusted.