Prove the result before integration.
A retrospective backtest comes first. Production authority follows only after the result survives your data, policy and review process.
Authority is earned in stages.
We begin offline, observe live decisions without changing them, and add action only where the evidence supports it.
Retrospective backtest
Offline evaluation
You provide 12 months of historical applications after identifiers are hashed inside your environment. We return the flagged cohort, the loss attached to it and a precision curve matched to your review capacity.
Shadow decisioning
No action taken
We score live applications beside your existing process. Your policy still decides the outcome while both teams compare reason codes against later repayment and fraud labels.
Step-up and review
Controlled authority
After the shadow results are signed off, the score can route selected cases to extra verification or an analyst queue. Decline authority remains outside the first production stage.
Policy decisioning
Measured rollout
Decline rules are considered only after thresholds, exception paths, monitoring and challenge procedures are documented for the institution and its supervisory obligations.
What the backtest needs.
The initial data request is narrow enough to prepare under an NDA and detailed enough to connect a score with a real outcome.
Application history
A 12-month extract with application identifiers, declared identity fields, device and channel attributes, and the decision recorded at origination.
Observed outcomes
First-payment default, confirmed fraud, charge-off, account closure and investigation outcomes where those labels exist.
Local tokenisation
Direct identifiers are transformed before transfer. The field map and hashing procedure are agreed before a file is prepared.
Review capacity
Your current manual-review volume sets the operating point for the precision curve. We do not optimise against a review queue your team cannot work.
Commercial terms are scoped to the institution.
Before paid work begins, a written proposal defines the pricing unit, volume assumptions, selected deployment and services included.
Custom, volume-based pricing
Production pricing is tailored to expected application volume, selected modules, deployment model and support scope. The order form records the pricing unit, volume tiers and any minimum commitment.
Deployment and service scope
A written proposal identifies the environments, implementation work, reporting, retention and support included in the quote so infrastructure and service costs are not hidden inside one rate.
Evaluation terms are separate
Any paid pilot is defined in its own statement of work, including the evaluation cohort, success measure, data handling, duration and fees. Production terms follow only if both parties choose to proceed.
No integration commitment
A backtest does not commit either party to a pilot, production integration or consortium membership. Any next stage requires its own scope and agreement.
Accuracy is not the operating metric.
The useful question is how much confirmed loss appears inside the share of applications your team can review.
A backtest reports precision at an agreed review volume, the loss linked to the flagged cohort, and the threshold at which the queue stops paying for itself. We also show misses and thin-file cases. A single headline detection rate would hide those trade-offs.
The production score is then monitored against the same definitions. If the label changes, the benchmark changes with it. That discipline matters more than a flattering percentage taken from a different lender's book.
Put your own loss file behind the claim.
The first call covers the cohort, available outcome labels, secure transfer path and the decision your team wants to test.
Book a backtest