From gut feel and a bureau score to an engine
By 2017 Netgíró was lending at national scale, and every credit decision rested on two things: an external bureau score and human judgement. Both are useful; neither learns. The bet we made that year — the founding bet of Moberg's data practice — was that the platform's own behavioural data could decide better: who to approve, what limit to set, and when to change it.
The first scorecard project combined Netgíró's behavioural history with years of bureau-score history, and concluded with something rarer than a model: a pilot of several competing models in production. That habit — never trust one model, make them compete — defined everything that followed.
How a scorecard actually gets built
The method we established then is the one we still teach. Segment the portfolio into homogeneous groups — new clients against clients with behavioural history, product by product. Define the target, analysing delinquency thresholds and time horizons until “default” means something precise. Engineer features across data sources and horizons, and test each one's predictive power. Then train competing model families — logistic regression, decision trees, ensembles — and let the evidence pick.
The evidence picked something interesting. A random forest won the early pilots, but the model that has held production longest is a logistic regression over complex engineered features — repayment and delinquency patterns crossed with bureau-score history. Explainable to a regulator, defensible to a board, and the best sales-risk balance of anything tested. The simplest method that wins.
Five generations, no rewrite
2018 · First blood
Initial scorecard development & pilot
Behavioural data meets bureau history, with several alternative models piloted head-to-head in production. The initial rollout restructured lending toward provably better risks — and cut risk losses by nearly 2×.
2019 · Listening
Pilot evaluation & recalibration
Results judged on risk and on sales — and on feedback from the customer-service team, whose case-by-case observations shaped the next model as much as the statistics did.
2020 · Stress test
Fresh model & anti-COVID measures
A new balanced model shipped just as the world changed, and limit-setting tightened sharply through the outbreak. The engine's job flipped from growth to protection — same platform, new policy.
2021 · Reopening
Controlled, monitored release
Step-by-step relaxation of the COVID rules, each step watched closely, until lending was back to pre-pandemic levels — proof that a decision engine lets you reopen with a dial rather than a switch.
2022 · Beyond one score
Trial runs & the multiple-scores architecture
A controlled trial testing higher limits for good-score clients fed a new architecture: one stable, explainable core score covering everyone, blended with alternative scores for sub-segments — shopping behaviour, mobile device, collections history — each improving decisions where it applies.
Underneath the generations sits the operating rhythm: champion–challenger A/B testing with a 10–20% champion group, a roughly one-year lifecycle, and evaluation on three axes — risk performance (entries to 60 and 90 days past due, expected loss), sales performance (volumes, attrition, limit volatility), and structured feedback from customer service and collections. Models earn production; they are never granted it.
Auditable ML, years before “MLOps”
Every month-end run of the models is centralised and traceable. Tasks are created automatically with explicit dependencies, an execution server pulls the exact code version from Git, results land in the warehouse carrying a reference to the precise commit that produced them, and every run — successful or failed — is logged.
A historical figure can be reproduced years later, from the exact code that computed it. Auditors noticed; so did we. This discipline became the seed of everything we now offer as ML & AI Ops.
One engine, a whole practice
Accounting-grade risk
IFRS 9 expected credit loss
The scorecard's probabilities feed audited monthly ECL — staging, PD curves, exposure and loss-given-default, with daily balance reconciliation and a second generation built on audit feedback. Even seasonality is modelled: holiday instalment behaviour can move segment ECL several-fold.
Planning
Budget projections & scenarios
The same definitions that book provisions project the portfolio forward — balances, repayments, product migrations, revenues. Business users compose their own scenarios in a spreadsheet and get full projections back, no developer required.
Growth
Segmentation, CLTV & beyond
Customer segmentation across risk, value and habits; 24-month customer lifetime value, fully automated in Power BI; repayment-schedule design; process mining over onboarding and checkout flows. The decision engine became a decision culture.
Team
- —Andriy Zhubryd
- —Ihor Protsiv
Key objectives
Replace an external bureau score and human judgement with decisions that learn from the platform's own behavioural data.
Set approvals, limits and terms in seconds — explainably enough to defend to a regulator and to a board.
Build an operating layer that carries new model generations without ever forcing a rewrite.
Services
- —Credit risk modelling
- —Feature engineering
- —Model governance & MLOps
- —IFRS 9 expected credit loss
- —Budget projection & scenarios
- —Segmentation & customer lifetime value
Technology
Numbers
~2×
Reduction in risk losses after the first rollout
5
Model generations on one operating layer
10–20%
Champion group in always-on A/B testing
100%
Of runs versioned, logged and reproducible



