Applying a Realist Evaluation Framework to Public Sector AI and Climate Governance
Artificial Intelligence is firmly on the agenda of governments worldwide. From the Republic of Korea to Spain, and India to Brazil, IT leaders and ministers are actively defining policies to apply AI to modernize public sector organizations, vastly improving the effectiveness and efficiency of administration.
The ambition is staggering. As highlighted during the recent OECD global report launch for Governing with Artificial Intelligence, the UK estimates that AI could automate 84% of repetitive public service transactions, saving the equivalent of 1,200 person-years of work annually. In Brazil, the government has introduced a strategic plan to avoid climate disaster in the Amazon using AI-enabled environmental monitoring systems.
Yet, despite these ambitious modernization programs, governments face a growing crisis: the rise of failed implementations. Millions are spent on AI solutions that do not scale or never progress past the pilot stage.
Why are these transformative initiatives failing to deliver genuine public value? The answer lies not in the technology, but in inappropriate evaluation methods.
The OECD Framework and the Evaluation Trap
To guide leaders in their AI-enabled modernization programs, the OECD (2025) proposed the Framework for Trustworthy AI in Government. It is built upon three broad pillars:
- Enablers: The foundational building blocks required to facilitate AI adoption, including data, digital infrastructure, skills/talent, and governance mechanisms (such as AI accelerators and innovation hubs).
- Guardrails: The binding and non-binding instruments needed to guide trustworthy use, encompassing risk management, transparency, accountability, and robust oversight.
- Engagement: Approaches to shape user-centered adoption by actively involving citizens, civil servants, and stakeholders to ensure AI is responsive to societal needs.
However, even with these pillars in place, the scaling problem persists. During the OECD launch event, speakers widely acknowledged the “challenges that we’re all facing in actually scaling AI projects… from a pilot to large-scale change.” The core issue is that standard evaluations are fundamentally flawed for complex digital transformation. Traditional evaluations treat success as a binary on/off switch: Was the software deployed? Yes or No? Did the users log in? Yes or No? But human adoption, trust, and institutional change in complex environments—whether in a sprawling public agency or the Amazon basin—do not operate in binary terms.
The Solution: A Realist Approach to Digital Transformation
As the Spanish Minister emphasized at the OECD event, we must look deeply into specific use cases to understand successes and failures, reminding us that “this is AI for humans… and that’s also talking about transparency because that is a key issue now.” Relying on use cases makes sense because it provides insights into the unique contextual conditions of an application domain.
To solve the scaling problem, a novel digital transformation evaluation framework has been developed by researchers from Reykjavik University alongside industry practitioners (Hjaltalin et al., forthcoming).
This framework moves beyond asking “Did it work?” by adopting and adapting a Context + MechanismMechanism refers to a specific intervention introduced in a particular social program. For example, the police install CCTV cameras on streets with a high rate of burglary. The functioning of mechanism in a program is best understood through metaphors. One is that of a ‘trigger’: a pistol fires a bullet triggered by the ignition of gunpowder within the pistol’s barrel. Another is the ‘dimmer-switch’ metaphor: a dimmer-switch operates on a continuum, where at maximum the light is very bright and at minimum it is pitch black. My understanding aligns closer to the latter metaphor and is used as such in my research. = Outcome (CMO) realist evaluation approach:
- Context (C): It examines the particular contextual conditions (e.g., governance structures, common standards, digital rights) that enable or constrain the initiative.
- Mechanism (M): It assesses the human drivers—such as trust and behavioral shifts among civil servants or citizens.
- Outcome (O): It evaluates to what extent the AI initiative actually produced public value outcomes under those particular circumstances.
Moving from Binary to Continuum
The true power of this framework is that it treats a particular mechanism as producing outcomes on a continuum, rather than a binary on/off switch.
Take Brazil’s AI climate governance plan as an example. Instead of a standard evaluation that simply asks, “Do the environmental enforcement agents use the AI risk predictions? (Yes/No),” the CMO framework evaluates the mechanism of trust on a continuum.
Perhaps the agents in a specific region are currently at “skeptical but experimenting,” which produces the outcome of “partial verification of AI data.” By mapping this continuum and identifying how specific contextual conditions constrain these mechanisms, policymakers don’t just see a “failed” pilot. They see exactly where the human friction is. They can then intervene—perhaps through targeted training or greater stakeholder engagement—moving the mechanism further along the continuum toward active reliance and co-creation.
Bridging the Gap Between Pilot and Public Value
Deploying technology is only ten percent of the battle. An incredibly sophisticated AI model that environmental agents ignore on the ground produces zero real-world value.
To overcome the scaling problem, governments and enterprises need to stop treating digital transformation as a simple IT deployment. Leaders must adopt analytical guardrails that measure context, mechanism, and outcomes to ensure that when an AI initiative scales, it genuinely empowers users and delivers measurable public value.
How Stirna Consulting Can Help
At Stirna Consulting, we understand that AI implementation is deeply sociotechnical. Led by our CEO and lead analyst, Hjaltalin—co-author of the forthcoming realist evaluation framework mentioned above—we don’t just help build systems. We provide the proprietary analytical methodologies and cross-functional change management necessary to ensure your digital transformation doesn’t stall at the pilot phase.
Whether you are deploying enterprise CRM software or navigating complex public sector IT programs, we ensure your technology empowers your team rather than dictating to them. Reach out to Stirna Consulting today for practical advice on digital transformation and clear next steps.
Sources
Hjaltalin, I. T., Candi, M., Sigurðarson, H. Th. (2026). Agile Governance in Public Sector Digital Transformation: a Program Evaluation [Manuscript submitted for publication]. Department of Business and Economics, Reykjavik University.
Hjaltalin, I. T., Bjartmarz, Th. (2026). Artificial Intelligence in Environmental Governance: Reconfiguring Public Value, Legitimacy, and Operational Capacity in EU Multi-Level Networks [Manuscript in preparation]. Department of Strategy & Innovation, Copenhagen Business School.

Leave a Reply