A key challenge of the Information Age is making the information that already exists useful. Too much information can be as much of a problem as too little, especially when what exists is not accessible, digestible, or actionable.
At the Scaling Community of Practice (SCoP), the problem is particularly acute because we draw on the experience of more than 5,000 members but our secretariat is fewer than ten people, none of them full time. We explored whether artificial intelligence (AI), specifically large language models (LLMs), can help us distil what our community knows without sacrificing accuracy.
We are still early in gathering lessons, but three recent efforts are far enough along to share.
For readers unfamiliar with SCoP, a brief description is provided below.
About the Scaling Community of Practice
The Scaling Community of Practice (SCoP) connects researchers and practitioners who are interested in exchanging views and experience on how to achieve sustainable impact at the scale of the problem. The SCoP produces research, policy analysis and tools, issues newsletters with contributions from its members about their scaling experience, organizes webinars and forums, provides scaling advice and advocates for a more effective focus on scaling promising innovations and interventions by all development actors, including governments, civil society actors, academia and international funders.
1. Asking AI to Review the Literature
The SCoP recently examined whether and how country coordination platforms (”country platforms”) can support pathways to sustainable impact at scale by bringing together national and international stakeholders. As usual, we searched the literature ourselves. This time we also asked ChatGPT to summarize the available knowledge supported by relevant references. The results appear in an annex to “Country Platforms and Scaling: An Exploration of Key Issues.”
Positives
The AI response was consistent with the findings of our traditional review. However, it offered a more comprehensive inventory of country platform experience, in far less time, than we could have assembled with the resources available. It may have mitigated some of our own biases by guiding us to relevant references we had not previously encountered, thereby reducing the “echo chamber” effect and allowing us to feature new examples. Its responses were well structured around the main factors to be taken into consideration. As essentially pattern recognition software, the AI identified linkages that we had alluded to but not yet clearly articulated.
Challenges
To get the information you need, you have to ask the right questions. ChatGPT’s suggested follow-ups were generally, but not always, useful.
Unlike econometric tools, AI gives you a result with no test of significance and no measure of its own confidence. In fact, AI often overstates its confidence in its own conclusions. Ground-truthing against your own literature review is essential.
AI gives you the high-level implications, but you still have to get into the weeds of individual sources to judge their reasoning and limitations.
Conclusion
AI is only the beginning of the research process, perhaps 10 percent of the insights in the final paper came from AI. Our own ideas had to drive the rest. And for transparency, share the questions you asked and the answers you got, so readers can check your interpretation against the original response.
2. Scoring institutional performance across 28 case studies
Our flagship Mainstreaming Scaling in Funder Organizations initiative produced 28 case studies of individual funders, written by more than 15 authors. We wanted to compare them semi-quantitatively against our Mainstreaming Tracking Tool, but there were too many for one person to score systematically, and different reviewers rated the same case differently. So we used NotebookLM (now Gemini Notebook) to apply the assessment tool across all 28 case studies.
Positives
By providing NotebookLM guidance documents on assessment and through the systematic use of consistent prompts, we were able to score each case study against our assessment tool. This simply would not have been possible without the use of AI.
At the time, NotebookLM was unusual in restricting its answers to uploaded documents, although its broader training still shaped responses. Other tools, including Claude Cowork, now do the same; however, we still like that NotebookLM quotes the passage behind each claim, hyperlinked to the source.
Challenges
Unsurprisingly, given that human reviewers were inconsistent, the AI sometimes scored the same case differently when re-prompted. We cannot say whether it was the more consistent judge; we did not compute formal agreement statistics. Using the same prompts and reference documents, and starting a fresh session for each case, limited its variation.
The main scoring challenge was leniency. Reviewers often thought the AI scored too generously, a common failure in LLM-as-a-judge applications, attributed to leniency bias, verbosity bias, and sycophancy. The cause is likely structural. Source-grounded systems answer by retrieving the passages most relevant to the question. Ask “does this case meet criterion X?” and you get passages that discuss X, which read as evidence in favour. Absence of evidence is hard to retrieve and easy to miss. NotebookLM often re-served the same quotes as evidence for meeting several criteria, inflating apparent support.
None of this was set-and-forget. Scoring 28 cases took several hours, since each case needed a fresh session and every result had to be checked.
Conclusions
Rapid, AI-based judgement was adequate for low-stakes, heuristic assessment, but significantly more training would likely have been necessary to get consistent scoring. Investment in such training would have negated the purpose of using the AI to accelerate the scoring process.
3. Building a searchable database of member experience
In its 11 years, the SCoP published 36 newsletters covering almost 450 member projects. A newsletter gives each project a paragraph and a hyperlink, leaving a member no practical way to find the few examples relevant to them across all newsletters. Summarizing projects by hand was impractical for a secretariat our size, so we used Claude to extract the information and build an interactive dashboard, designing prompts to limit hallucination and spot-checking the results.
Positives
AI was able to systematically extract narrative information from text and provide it in a structured, dashboard format with high accuracy. The database lets members learn from one another’s work, find partners, and see where clusters of scaling activity sit and where the gaps are. It would not have existed without AI-assisted extraction.
AI (Claude code) was also used to build the dashboard itself. Humans described in detail what they wanted the dashboard to look like and provided feedback, but no human coding was involved in the development of the dashboard.
Challenges
This experience raised several ethical questions, which we considered carefully:
Reducing hallucination and respecting authorship pull against each other. We asked Claude to extract direct quotes to stay anchored to the source, then to paraphrase them to avoid plagiarizing our members. We remain unsure where the balance lies.
Hallucinations are still possible, and our quality assurance was limited. We accept the risk because the database is a tool for exploration, not a basis for action. Users have to go to the underlying documents to get enough information for decision making and will identify errors then.
Consent is unresolved. Recent newsletters ask members for consent to be included in the database. We included earlier submissions on the grounds that members had already agreed to have their work disseminated through our network, inclusion benefits members, and that the database uses only public information. We announced it with an opt-out in the newsletter where the projects were first submitted.
What We Would Tell a Colleague
AI changed what a ten-person secretariat could achieve. What it gave us in all three examples was more coverage: more literature than we could have read, more case studies than one person could score, more member projects than we could have summarized by hand. What it did not give us was judgment: which references mattered, whether a funder had genuinely mainstreamed scaling or merely described it well, whether a project summary was faithful to what the member wrote, what is the significance of a result. AI can be used to increase output, but critical human reflection is still necessary. When working with colleagues, we discuss, verify, and bring our own judgement to the conversation. The same can be done when working with AI. LLMs can be a partner and accelerator, but not a final source of truth.




