ERAM INTELLIGENCE
Decisioning platforms

Simulate before you deploy: testing decisioning strategy safely

Some decisions cannot be A/B tested into existence. Simulation is what turns "we think this will work" into something a risk function can authorise.

There is a category of decision that cannot be A/B tested into existence. Changing the contact policy for a national debt collection programme is one. Changing a credit approval threshold is another. The population is real, the consequences are material, and "we tried it and it went badly" is not an acceptable answer.

This is where simulation earns its place, and it is the part of decisioning platform work that gets least attention relative to its importance.

What a decisioning simulation actually does

In a mature decisioning platform such as Pega Customer Decision Hub, a strategy is not a single model. It is a chain: eligibility rules determine who can receive an action, applicability and suitability narrow it further, adaptive and predictive models score the remaining candidates, arbitration ranks them against business priorities, and constraints govern how often anyone can be contacted.

A simulation runs a defined population through that entire chain without dispatching anything, and reports what would have happened.

That answers questions no amount of model validation can:

Each of those is a business question rather than a modelling one, and each is invisible from a model performance metric.

A model can be individually excellent and still produce a strategy that is operationally unworkable.

The default value problem

A specific example from recent work on a national campaign programme for a UK government department.

Adaptive models in production decisioning need to score customers before they have accumulated evidence. Every such model therefore has behaviour for the cold-start case, and that behaviour is governed by a default. In Pega's adaptive decision manager this surfaces through the uncertainty-adjusted score used in arbitration.

The default is a configuration choice, and it is consequential. Set it too high and untested actions dominate arbitration, flooding the population with propositions the system knows nothing about. Set it too low and new actions never accumulate enough exposure to learn, so they stay permanently unproven. Neither failure is visible in a model quality report. Both are immediately visible in a simulation.

Diagnosing that behaviour meant working through adaptive model test loops and analysing default value handling directly, then documenting the intended configuration so that the decision was recorded rather than rediscovered by the next person.

Synthetic data as a first-class tool

Simulation needs a population. In regulated and public sector environments, using real customer records for strategy development carries governance overhead that can stall a programme for months.

Synthetic populations solve a narrower problem than people assume. They are poor for validating model accuracy, because the relationships they contain are the ones you put there. They are excellent for validating strategy mechanics, because mechanics do not depend on the data being real.

If you need to know whether the arbitration logic ranks correctly, whether contact policies suppress as intended, whether a customer with a particular profile becomes eligible for the right treatment, a synthetic population answers it without touching a single real record. That distinction is worth being precise about with a risk function, because it turns a governance blocker into a scoped and defensible tool.

Getting models into the platform

One practical note. Where models are built outside the decisioning platform, in Python, they have to be brought back in as deployable artefacts.

Two interchange formats dominate. PMML is older, XML-based, and widely supported. ONNX is the better choice for most production deployments: broader support for modern model types, better handling of preprocessing pipelines, and more active development behind it. PMML remains a reasonable fallback where a platform's ONNX support is limited, but it should be a fallback rather than a default.

Getting this decision right early matters more than it appears, because it determines whether the data science team can iterate independently or has to queue behind a platform release.

Why this is a governance argument, not just an engineering one

Gartner has forecast that more than forty per cent of agentic AI projects will be cancelled by the end of 2027, citing unclear business value and inadequate risk controls among the principal causes.1

Simulation addresses both. It converts "we think this strategy will work" into a quantified volume and outcome projection that a business owner can sign off. And it gives a risk function a mechanism to inspect a strategy's behaviour before it touches a customer, which is precisely the control that autonomous decisioning otherwise lacks.

For any organisation moving toward more automated decisioning, I would treat simulation capability as a prerequisite rather than a refinement. It is the difference between a system you have authorised and a system you are hoping about.

References & notes

  1. Gartner press release, June 2025, forecasting cancellation of more than 40 per cent of agentic AI projects by end-2027 owing to escalating costs, unclear business value and inadequate risk controls.
  2. Pega Customer Decision Hub architecture references: Adaptive Decision Manager, Prediction Studio, Next-Best-Action Designer and decisioning simulation workflow, Pega product documentation. docs.pega.com
  3. ONNX (Open Neural Network Exchange) is an open standard for machine learning interoperability. onnx.ai

Eram Intelligence advises on AI decisioning, production machine learning and generative AI for enterprise and government.

Start a conversation     All insights