Proving Ground a gated model release template

Release report. Static, offline, read-only.

Overview

Champion
c2c40600b7290013 (glm)
Previous champion
none
Candidate
none
Stage
stable
Latest decision
rollback seq 5, 2026-10-08T01:07:12Z: Drift alert from synthetic scenario 'young_driver_surge'; restoring previous champion per runbook.
Audit events
hold: 1, promote: 2, reject: 1, rollback: 1
Audit chain
verified (5 entries)
Candidates gated
3 (2 passed, 1 rejected)

Model comparison (same holdout, 95% bootstrap CIs)

modelcandidatedeviancedeviance CIskill vs constantGiniGini CI
Constant baselinec7e902deae1.1881[1.117, 1.231]00[0, 0]
Poisson GLMc2c40600b70.804[0.7641, 0.8432]0.3230.551[0.523, 0.583]
LightGBM5fc4383bbf0.80227[0.7645, 0.8349]0.3250.551[0.523, 0.58]

Lowest deviance: LightGBM. Differences smaller than the CI width are not evidence of a real difference; see the paired-bootstrap G3 check for the champion/challenger decision. Holdout is one random grouped split, not a time split.

Release timeline

Release timelineAudit events in sequence order: #1 promote, #2 reject, #3 promote, #4 hold, #5 rollback#1 promote✔promote#1#2 reject✖reject#2#3 promote✔promote#3#4 hold⚠hold#4#5 rollback⚠rollback#5
#time (UTC)eventactorreason and evidencecandidatechampion before → afterflags
12026-10-08T01:07:02Z promotedemoBootstrap promotion: no champion existed, so G3 was skipped and shadow/canary had nothing to compare against.
evidence (1)
c2c40600b7none → c2c40600b7synthetic
22026-10-08T01:07:03Z rejectproving-groundGate(s) G2, G3, G4, G5 failed; candidate rejected.
evidence (1)
  • gate_report: 60528a9499f3939d/gate_report.json failed: G2:deviance_skill, G2:gini, G2:poisson_deviance, G3:poisson_deviance, G3:gini, G4:decile_calibration_error, G5:region=other, G5:bonus_malus=55-70, G5:exposure=0.25-0.75, G5:exposure=0.75-1
60528a9499c2c40600b7 → c2c40600b7synthetic
32026-10-08T01:07:06Z promotedemoPassed gates, shadow and canary; promoted to champion.
evidence (3)
5fc4383bbfc2c40600b7 → 5fc4383bbfsynthetic
42026-10-08T01:07:11Z holdproving-groundDrift alert in window 10 (drivers: DrivAge, prediction). Rollback to previous champion c2c40600b7290013 recommended.
evidence (1)
5fc4383bbf5fc4383bbf → 5fc4383bbfsynthetic
52026-10-08T01:07:12Z rollbackdemoDrift alert from synthetic scenario 'young_driver_surge'; restoring previous champion per runbook.
evidence (2)
  • drift_summary: drift/incident/timeline.json alert; see window files
  • health_check: served champion c2c40600b7290013 via GET /health; expected c2c40600b7290013: OK
5fc4383bbf5fc4383bbf → c2c40600b7synthetic

Hashes shortened to 10 characters; full values in the JSON evidence.

Gate report

Candidate 5fc4383bbfa34776 (gbm) pass
data version
synthetic-7c8426d485-n40000
git
bcbdf51a7d5d
seed
20260101
gate config
7a4a29aea97dd405
lockfile
325e385df3aac6ed
evidence
gate_report.json · gate_report.md

Candidate vs champion (holdout, exposure-weighted)

metriccandidate95% CIchampion
deviance_skill0.32472[0.2924, 0.3482]0.32327
gini0.5508[0.5229, 0.5803]0.5507
poisson_deviance0.80227[0.7645, 0.8349]0.804

All checks

gatecheckstatusmeasuredthreshold95% CIchampiondetail
G1Data validation pass
schema[train] pass––––24001 rows valid
schema[val] pass––––4000 rows valid
schema[holdout] pass––––6000 rows valid
row_floor[train] pass2400110000––
row_floor[val] pass40001000––
row_floor[holdout] pass60001000––
null_rate pass00.0––
split_leakage pass0 shared ids0––no id may appear in more than one split
G2Metric floors pass
deviance_skill pass0.3247min 0.005[0.2924, 0.3482]–
gini pass0.5508min 0.1[0.5229, 0.5803]–
poisson_deviance pass0.8023max 0.9[0.7645, 0.8349]–
G3No regression vs champion pass
poisson_deviance pass0.8023relative worsening <= 0.02[-0.01198, 0.005411]0.804relative worsening -0.2143% (95% CI of worsening shown)
gini pass0.5508relative worsening <= 0.1[-0.01276, 0.01367]0.5507relative worsening -0.0170% (95% CI of worsening shown)
G4Calibration pass
decile_calibration_error pass0.08826max 0.35––
G5Slice performance warn
bonus_malus=55-70 pass0.9403|gap| <= 0.45 vs 1.0[0.8587, 1.01]–n=2884, gap 6.0%
bonus_malus=70-90 insufficientinsufficient datan >= 2000––n=746: slice too small to judge; reported, not passed silently
bonus_malus=90-110 insufficientinsufficient datan >= 2000––n=104: slice too small to judge; reported, not passed silently
bonus_malus=<55 pass0.9577|gap| <= 0.45 vs 1.0[0.8426, 1.098]–n=2241, gap 4.2%
bonus_malus=>=110 insufficientinsufficient datan >= 2000––n=25: slice too small to judge; reported, not passed silently
driver_age=25-35 insufficientinsufficient datan >= 2000––n=1011: slice too small to judge; reported, not passed silently
driver_age=35-45 insufficientinsufficient datan >= 2000––n=1481: slice too small to judge; reported, not passed silently
driver_age=45-55 insufficientinsufficient datan >= 2000––n=1477: slice too small to judge; reported, not passed silently
driver_age=55-65 insufficientinsufficient datan >= 2000––n=969: slice too small to judge; reported, not passed silently
driver_age=65-75 insufficientinsufficient datan >= 2000––n=420: slice too small to judge; reported, not passed silently
driver_age=<25 insufficientinsufficient datan >= 2000––n=513: slice too small to judge; reported, not passed silently
driver_age=>=75 insufficientinsufficient datan >= 2000––n=129: slice too small to judge; reported, not passed silently
exposure=0.25-0.75 pass1.027|gap| <= 0.45 vs 1.0[0.9592, 1.1]–n=3200, gap 2.7%
exposure=0.75-1 pass0.9596|gap| <= 0.45 vs 1.0[0.8747, 1.032]–n=2083, gap 4.0%
exposure=<0.25 insufficientinsufficient datan >= 2000––n=715: slice too small to judge; reported, not passed silently
exposure=>=1 insufficientinsufficient datan >= 2000––n=2: slice too small to judge; reported, not passed silently
region=R21 insufficientinsufficient datan >= 2000––n=299: slice too small to judge; reported, not passed silently
region=R31 insufficientinsufficient datan >= 2000––n=257: slice too small to judge; reported, not passed silently
region=R42 insufficientinsufficient datan >= 2000––n=281: slice too small to judge; reported, not passed silently
region=R52 insufficientinsufficient datan >= 2000––n=293: slice too small to judge; reported, not passed silently
region=R72 insufficientinsufficient datan >= 2000––n=293: slice too small to judge; reported, not passed silently
region=R73 insufficientinsufficient datan >= 2000––n=292: slice too small to judge; reported, not passed silently
region=R94 insufficientinsufficient datan >= 2000––n=259: slice too small to judge; reported, not passed silently
region=other pass0.9725|gap| <= 0.45 vs 1.0[0.9157, 1.02]–n=4026, gap 2.7%
vehicle_age=1-3 insufficientinsufficient datan >= 2000––n=1094: slice too small to judge; reported, not passed silently
vehicle_age=10-15 insufficientinsufficient datan >= 2000––n=933: slice too small to judge; reported, not passed silently
vehicle_age=3-6 insufficientinsufficient datan >= 2000––n=1781: slice too small to judge; reported, not passed silently
vehicle_age=6-10 insufficientinsufficient datan >= 2000––n=1590: slice too small to judge; reported, not passed silently
vehicle_age=<1 insufficientinsufficient datan >= 2000––n=202: slice too small to judge; reported, not passed silently
vehicle_age=>=15 insufficientinsufficient datan >= 2000––n=400: slice too small to judge; reported, not passed silently
G6Performance budget pass
p95_latency_ms pass3.479max 60––single-row predict(), n=100
artifact_mb pass0.3max 50––
G7Reproducibility pass
same_hash pass5fc4383bbfa347765fc4383bbfa34776––identical hash on retrain with same seed
G8Model card pass
card_exists pass––––
card_matches_candidate pass––––card must name this candidate hash (regenerated per candidate)
required_sections passall present9 sections––
Candidate 60528a9499f3939d (gbm_degraded) fail failed: G2, G3, G4, G5
data version
synthetic-7c8426d485-n40000
git
bcbdf51a7d5d
seed
20260101
gate config
7a4a29aea97dd405
lockfile
325e385df3aac6ed
evidence
gate_report.json · gate_report.md

Candidate vs champion (holdout, exposure-weighted)

metriccandidate95% CIchampion
deviance_skill-0.19907[-0.2195, -0.165]0.32327
gini-0.16149[-0.1901, -0.1206]0.5507
poisson_deviance1.4246[1.303, 1.499]0.804

All checks

gatecheckstatusmeasuredthreshold95% CIchampiondetail
G1Data validation pass
schema[train] pass––––24001 rows valid
schema[val] pass––––4000 rows valid
schema[holdout] pass––––6000 rows valid
row_floor[train] pass2400110000––
row_floor[val] pass40001000––
row_floor[holdout] pass60001000––
null_rate pass00.0––
split_leakage pass0 shared ids0––no id may appear in more than one split
G2Metric floors fail
deviance_skill fail-0.1991min 0.005[-0.2195, -0.165]–
gini fail-0.1615min 0.1[-0.1901, -0.1206]–
poisson_deviance fail1.425max 0.9[1.303, 1.499]–
G3No regression vs champion fail
poisson_deviance fail1.425relative worsening <= 0.02[0.6452, 0.8515]0.804relative worsening +77.1867% (95% CI of worsening shown)
gini fail-0.1615relative worsening <= 0.1[1.19, 1.39]0.5507relative worsening +129.3237% (95% CI of worsening shown)
G4Calibration fail
decile_calibration_error fail0.6431max 0.35––
G5Slice performance fail
driver_age=25-35 insufficientinsufficient datan >= 2000––n=1011: slice too small to judge; reported, not passed silently
driver_age=35-45 insufficientinsufficient datan >= 2000––n=1481: slice too small to judge; reported, not passed silently
driver_age=45-55 insufficientinsufficient datan >= 2000––n=1477: slice too small to judge; reported, not passed silently
driver_age=55-65 insufficientinsufficient datan >= 2000––n=969: slice too small to judge; reported, not passed silently
driver_age=65-75 insufficientinsufficient datan >= 2000––n=420: slice too small to judge; reported, not passed silently
driver_age=<25 insufficientinsufficient datan >= 2000––n=513: slice too small to judge; reported, not passed silently
driver_age=>=75 insufficientinsufficient datan >= 2000––n=129: slice too small to judge; reported, not passed silently
vehicle_age=1-3 insufficientinsufficient datan >= 2000––n=1094: slice too small to judge; reported, not passed silently
vehicle_age=10-15 insufficientinsufficient datan >= 2000––n=933: slice too small to judge; reported, not passed silently
vehicle_age=3-6 insufficientinsufficient datan >= 2000––n=1781: slice too small to judge; reported, not passed silently
vehicle_age=6-10 insufficientinsufficient datan >= 2000––n=1590: slice too small to judge; reported, not passed silently
vehicle_age=<1 insufficientinsufficient datan >= 2000––n=202: slice too small to judge; reported, not passed silently
vehicle_age=>=15 insufficientinsufficient datan >= 2000––n=400: slice too small to judge; reported, not passed silently
region=R21 insufficientinsufficient datan >= 2000––n=299: slice too small to judge; reported, not passed silently
region=R31 insufficientinsufficient datan >= 2000––n=257: slice too small to judge; reported, not passed silently
region=R42 insufficientinsufficient datan >= 2000––n=281: slice too small to judge; reported, not passed silently
region=R52 insufficientinsufficient datan >= 2000––n=293: slice too small to judge; reported, not passed silently
region=R72 insufficientinsufficient datan >= 2000––n=293: slice too small to judge; reported, not passed silently
region=R73 insufficientinsufficient datan >= 2000––n=292: slice too small to judge; reported, not passed silently
region=R94 insufficientinsufficient datan >= 2000––n=259: slice too small to judge; reported, not passed silently
region=other fail2.285|gap| <= 0.45 vs 1.0[2.124, 2.48]–n=4026, gap 128.5%
bonus_malus=55-70 fail2.172|gap| <= 0.45 vs 1.0[1.957, 2.345]–n=2884, gap 117.2%
bonus_malus=70-90 insufficientinsufficient datan >= 2000––n=746: slice too small to judge; reported, not passed silently
bonus_malus=90-110 insufficientinsufficient datan >= 2000––n=104: slice too small to judge; reported, not passed silently
bonus_malus=<55 pass0.9957|gap| <= 0.45 vs 1.0[0.8687, 1.165]–n=2241, gap 0.4%
bonus_malus=>=110 insufficientinsufficient datan >= 2000––n=25: slice too small to judge; reported, not passed silently
exposure=0.25-0.75 fail2.467|gap| <= 0.45 vs 1.0[2.214, 2.706]–n=3200, gap 146.7%
exposure=0.75-1 fail2.237|gap| <= 0.45 vs 1.0[2.094, 2.415]–n=2083, gap 123.7%
exposure=<0.25 insufficientinsufficient datan >= 2000––n=715: slice too small to judge; reported, not passed silently
exposure=>=1 insufficientinsufficient datan >= 2000––n=2: slice too small to judge; reported, not passed silently
G6Performance budget pass
p95_latency_ms pass3.734max 60––single-row predict(), n=100
artifact_mb pass0.193max 50––
G7Reproducibility pass
same_hash pass60528a9499f3939d60528a9499f3939d––identical hash on retrain with same seed
G8Model card pass
card_exists pass––––
card_matches_candidate pass––––card must name this candidate hash (regenerated per candidate)
required_sections passall present9 sections––
Candidate c2c40600b7290013 (glm) pass
data version
synthetic-7c8426d485-n40000
git
bcbdf51a7d5d
seed
20260101
gate config
7a4a29aea97dd405
lockfile
325e385df3aac6ed
evidence
gate_report.json · gate_report.md

Candidate vs champion (holdout, exposure-weighted)

metriccandidate95% CIchampion
deviance_skill0.32327[0.2906, 0.3472]no champion at the time
gini0.5507[0.5228, 0.583]no champion at the time
poisson_deviance0.804[0.7641, 0.8432]no champion at the time

All checks

gatecheckstatusmeasuredthreshold95% CIchampiondetail
G1Data validation pass
schema[train] pass––––24001 rows valid
schema[val] pass––––4000 rows valid
schema[holdout] pass––––6000 rows valid
row_floor[train] pass2400110000––
row_floor[val] pass40001000––
row_floor[holdout] pass60001000––
null_rate pass00.0––
split_leakage pass0 shared ids0––no id may appear in more than one split
G2Metric floors pass
deviance_skill pass0.3233min 0.005[0.2906, 0.3472]–
gini pass0.5507min 0.1[0.5228, 0.583]–
poisson_deviance pass0.804max 0.9[0.7641, 0.8432]–
G3No regression vs champion skipNo champion exists: first promotion. G3 skipped and recorded as skipped.
regression skip––––No champion exists: first promotion. G3 skipped and recorded as skipped.
G4Calibration pass
decile_calibration_error pass0.05228max 0.35––
G5Slice performance warn
driver_age=25-35 insufficientinsufficient datan >= 2000––n=1011: slice too small to judge; reported, not passed silently
driver_age=35-45 insufficientinsufficient datan >= 2000––n=1481: slice too small to judge; reported, not passed silently
driver_age=45-55 insufficientinsufficient datan >= 2000––n=1477: slice too small to judge; reported, not passed silently
driver_age=55-65 insufficientinsufficient datan >= 2000––n=969: slice too small to judge; reported, not passed silently
driver_age=65-75 insufficientinsufficient datan >= 2000––n=420: slice too small to judge; reported, not passed silently
driver_age=<25 insufficientinsufficient datan >= 2000––n=513: slice too small to judge; reported, not passed silently
driver_age=>=75 insufficientinsufficient datan >= 2000––n=129: slice too small to judge; reported, not passed silently
vehicle_age=1-3 insufficientinsufficient datan >= 2000––n=1094: slice too small to judge; reported, not passed silently
vehicle_age=10-15 insufficientinsufficient datan >= 2000––n=933: slice too small to judge; reported, not passed silently
vehicle_age=3-6 insufficientinsufficient datan >= 2000––n=1781: slice too small to judge; reported, not passed silently
vehicle_age=6-10 insufficientinsufficient datan >= 2000––n=1590: slice too small to judge; reported, not passed silently
vehicle_age=<1 insufficientinsufficient datan >= 2000––n=202: slice too small to judge; reported, not passed silently
vehicle_age=>=15 insufficientinsufficient datan >= 2000––n=400: slice too small to judge; reported, not passed silently
region=R21 insufficientinsufficient datan >= 2000––n=299: slice too small to judge; reported, not passed silently
region=R31 insufficientinsufficient datan >= 2000––n=257: slice too small to judge; reported, not passed silently
region=R42 insufficientinsufficient datan >= 2000––n=281: slice too small to judge; reported, not passed silently
region=R52 insufficientinsufficient datan >= 2000––n=293: slice too small to judge; reported, not passed silently
region=R72 insufficientinsufficient datan >= 2000––n=293: slice too small to judge; reported, not passed silently
region=R73 insufficientinsufficient datan >= 2000––n=292: slice too small to judge; reported, not passed silently
region=R94 insufficientinsufficient datan >= 2000––n=259: slice too small to judge; reported, not passed silently
region=other pass0.9656|gap| <= 0.45 vs 1.0[0.9093, 1.009]–n=4026, gap 3.4%
bonus_malus=55-70 pass0.9613|gap| <= 0.45 vs 1.0[0.8824, 1.035]–n=2884, gap 3.9%
bonus_malus=70-90 insufficientinsufficient datan >= 2000––n=746: slice too small to judge; reported, not passed silently
bonus_malus=90-110 insufficientinsufficient datan >= 2000––n=104: slice too small to judge; reported, not passed silently
bonus_malus=<55 pass0.9587|gap| <= 0.45 vs 1.0[0.8409, 1.102]–n=2241, gap 4.1%
bonus_malus=>=110 insufficientinsufficient datan >= 2000––n=25: slice too small to judge; reported, not passed silently
exposure=0.25-0.75 pass1.01|gap| <= 0.45 vs 1.0[0.9436, 1.09]–n=3200, gap 1.0%
exposure=0.75-1 pass0.9559|gap| <= 0.45 vs 1.0[0.8736, 1.027]–n=2083, gap 4.4%
exposure=<0.25 insufficientinsufficient datan >= 2000––n=715: slice too small to judge; reported, not passed silently
exposure=>=1 insufficientinsufficient datan >= 2000––n=2: slice too small to judge; reported, not passed silently
G6Performance budget pass
p95_latency_ms pass5.269max 60––single-row predict(), n=100
artifact_mb pass0.005max 50––
G7Reproducibility pass
same_hash passc2c40600b7290013c2c40600b7290013––identical hash on retrain with same seed
G8Model card pass
card_exists pass––––
card_matches_candidate pass––––card must name this candidate hash (regenerated per candidate)
required_sections passall present9 sections––

Calibration and slices

Candidate 5fc4383bbf (gbm) pass
ae_ratio
0.9946
decile_calibration_error
0.08826
Calibration by risk decile, candidate 5fc4383bbfdecile 1: predicted 0.0794, observed 0.0674; decile 2: predicted 0.0994, observed 0.0995; decile 3: predicted 0.1150, observed 0.0729; decile 4: predicted 0.1303, observed 0.1399; decile 5: predicted 0.1482, observed 0.1614; decile 6: predicted 0.1702, observed 0.0997; decile 7: predicted 0.2010, observed 0.2155; decile 8: predicted 0.2560, observed 0.2635; decile 9: predicted 0.3993, observed 0.3825; decile 10: predicted 1.4297, observed 1.50970.0000.0000.4080.4080.8150.8151.2231.2231.6311.631decile 1: predicted 0.0794, observed 0.0674decile 2: predicted 0.0994, observed 0.0995decile 3: predicted 0.1150, observed 0.0729decile 4: predicted 0.1303, observed 0.1399decile 5: predicted 0.1482, observed 0.1614decile 6: predicted 0.1702, observed 0.0997decile 7: predicted 0.2010, observed 0.2155decile 8: predicted 0.2560, observed 0.2635decile 9: predicted 0.3993, observed 0.3825decile 10: predicted 1.4297, observed 1.5097predicted frequencyobserved frequency
Data table for the chart
decileexposurepredictedobserved
1370.720.079360.06744
2371.990.099430.09947
3370.490.1150.07288
4371.730.13030.1399
5371.710.14820.1614
6371.070.17020.09971
7371.290.2010.2155
8371.890.2560.2635
9371.240.39930.3825
10371.591.431.51

Slices (actual / expected claims; 1.0 is calibrated)

slicelevelnA/E95% CIflag
bonus_malus55-7028840.94[0.859, 1.01]
bonus_malus70-907461.09– insufficient data
bonus_malus90-1101041.12– insufficient data
bonus_malus<5522410.958[0.843, 1.1]
bonus_malus>=110250.83– insufficient data
driver_age25-3510111.02– insufficient data
driver_age35-4514810.828– insufficient data
driver_age45-5514770.983– insufficient data
driver_age55-659691.07– insufficient data
driver_age65-754200.713– insufficient data
driver_age<255131.06– insufficient data
driver_age>=751290.522– insufficient data
exposure0.25-0.7532001.03[0.959, 1.1]
exposure0.75-120830.96[0.875, 1.03]
exposure<0.257151– insufficient data
exposure>=122.57– insufficient data
regionR212990.912– insufficient data
regionR312571.17– insufficient data
regionR422810.892– insufficient data
regionR522930.836– insufficient data
regionR722930.945– insufficient data
regionR732921.3– insufficient data
regionR942591.25– insufficient data
regionother40260.973[0.916, 1.02]
vehicle_age1-310940.97– insufficient data
vehicle_age10-159331.05– insufficient data
vehicle_age3-617811.04– insufficient data
vehicle_age6-1015900.938– insufficient data
vehicle_age<12021.13– insufficient data
vehicle_age>=154000.905– insufficient data
Candidate 60528a9499 (gbm_degraded) fail
ae_ratio
2.359
decile_calibration_error
0.6431
Calibration by risk decile, candidate 60528a9499decile 1: predicted 0.0883, observed 0.3288; decile 2: predicted 0.0973, observed 0.3394; decile 3: predicted 0.1024, observed 0.4012; decile 4: predicted 0.1077, observed 0.4629; decile 5: predicted 0.1140, observed 0.4101; decile 6: predicted 0.1220, observed 0.3387; decile 7: predicted 0.1307, observed 0.2422; decile 8: predicted 0.1422, observed 0.2184; decile 9: predicted 0.1603, observed 0.1588; decile 10: predicted 0.2122, observed 0.11290.0000.0000.1250.1250.2500.2500.3750.3750.5000.500decile 1: predicted 0.0883, observed 0.3288decile 2: predicted 0.0973, observed 0.3394decile 3: predicted 0.1024, observed 0.4012decile 4: predicted 0.1077, observed 0.4629decile 5: predicted 0.1140, observed 0.4101decile 6: predicted 0.1220, observed 0.3387decile 7: predicted 0.1307, observed 0.2422decile 8: predicted 0.1422, observed 0.2184decile 9: predicted 0.1603, observed 0.1588decile 10: predicted 0.2122, observed 0.1129predicted frequencyobserved frequency
Data table for the chart
decileexposurepredictedobserved
1371.090.088340.3288
2371.260.097330.3394
3371.430.10240.4012
4371.560.10770.4629
5370.670.1140.4101
63720.1220.3387
7371.530.13070.2422
8370.80.14220.2184
9371.440.16030.1588
10371.940.21220.1129

Slices (actual / expected claims; 1.0 is calibrated)

slicelevelnA/E95% CIflag
bonus_malus55-7028842.17[1.96, 2.34]
bonus_malus70-907466.36– insufficient data
bonus_malus90-11010412.9– insufficient data
bonus_malus<5522410.996[0.869, 1.17]
bonus_malus>=1102519.7– insufficient data
driver_age25-3510113.21– insufficient data
driver_age35-4514811.19– insufficient data
driver_age45-5514771.31– insufficient data
driver_age55-659690.919– insufficient data
driver_age65-754200.464– insufficient data
driver_age<2551314.9– insufficient data
driver_age>=751290.249– insufficient data
exposure0.25-0.7532002.47[2.21, 2.71]
exposure0.75-120832.24[2.09, 2.41]
exposure<0.257152.59– insufficient data
exposure>=123.93– insufficient data
regionR212992.18– insufficient data
regionR312572.8– insufficient data
regionR422812.34– insufficient data
regionR522932.21– insufficient data
regionR722931.85– insufficient data
regionR732923.06– insufficient data
regionR942593.2– insufficient data
regionother40262.28[2.12, 2.48]
vehicle_age1-310942.13– insufficient data
vehicle_age10-159332.49– insufficient data
vehicle_age3-617812.41– insufficient data
vehicle_age6-1015902.49– insufficient data
vehicle_age<12022.65– insufficient data
vehicle_age>=154001.84– insufficient data
Candidate c2c40600b7 (glm) pass
ae_ratio
0.985
decile_calibration_error
0.05228
Calibration by risk decile, candidate c2c40600b7decile 1: predicted 0.0653, observed 0.0809; decile 2: predicted 0.0885, observed 0.0700; decile 3: predicted 0.1060, observed 0.0862; decile 4: predicted 0.1228, observed 0.1399; decile 5: predicted 0.1406, observed 0.1429; decile 6: predicted 0.1637, observed 0.1345; decile 7: predicted 0.1968, observed 0.1938; decile 8: predicted 0.2529, observed 0.2665; decile 9: predicted 0.4027, observed 0.3715; decile 10: predicted 1.5197, observed 1.52670.0000.0000.4120.4120.8240.8241.2371.2371.6491.649decile 1: predicted 0.0653, observed 0.0809decile 2: predicted 0.0885, observed 0.0700decile 3: predicted 0.1060, observed 0.0862decile 4: predicted 0.1228, observed 0.1399decile 5: predicted 0.1406, observed 0.1429decile 6: predicted 0.1637, observed 0.1345decile 7: predicted 0.1968, observed 0.1938decile 8: predicted 0.2529, observed 0.2665decile 9: predicted 0.4027, observed 0.3715decile 10: predicted 1.5197, observed 1.5267predicted frequencyobserved frequency
Data table for the chart
decileexposurepredictedobserved
1370.650.065290.08094
2371.680.088520.06995
3371.290.1060.08619
4371.790.12280.1399
5370.880.14060.1429
6371.70.16370.1345
7371.450.19680.1938
8371.440.25290.2665
9371.460.40270.3715
10371.381.521.527

Slices (actual / expected claims; 1.0 is calibrated)

slicelevelnA/E95% CIflag
bonus_malus55-7028840.961[0.882, 1.03]
bonus_malus70-907461.07– insufficient data
bonus_malus90-1101041.14– insufficient data
bonus_malus<5522410.959[0.841, 1.1]
bonus_malus>=110250.595– insufficient data
driver_age25-3510111– insufficient data
driver_age35-4514810.881– insufficient data
driver_age45-5514770.982– insufficient data
driver_age55-659691.21– insufficient data
driver_age65-754200.74– insufficient data
driver_age<255131– insufficient data
driver_age>=751290.54– insufficient data
exposure0.25-0.7532001.01[0.944, 1.09]
exposure0.75-120830.956[0.874, 1.03]
exposure<0.257151.02– insufficient data
exposure>=123.27– insufficient data
regionR212990.887– insufficient data
regionR312571.19– insufficient data
regionR422810.891– insufficient data
regionR522930.811– insufficient data
regionR722930.91– insufficient data
regionR732921.3– insufficient data
regionR942591.22– insufficient data
regionother40260.966[0.909, 1.01]
vehicle_age1-310940.871– insufficient data
vehicle_age10-159331.08– insufficient data
vehicle_age3-617811.05– insufficient data
vehicle_age6-1015900.947– insufficient data
vehicle_age<12021.13– insufficient data
vehicle_age>=154000.92– insufficient data

Shadow and canary

All traffic below is simulated (held-out policies replayed in random order).

Candidate 5fc4383bbfa34776

Shadow pass

requests
3600 (minimum 3000)
agreement within tolerance
0.694
rank correlation (Spearman)
0.916
mean latency delta (candidate − champion)
-0.0189 ms
exception rate
0
poisson_deviance vs champion (labelled rows)
candidate 0.8339 vs champion 0.8434 (n=3600)
evidence
shadow_summary.json
checkstatusvaluethreshold
spearman pass0.9158>= 0.85
exception_rate pass0<= 0.01
latency_delta_ms pass-0.01893<= 30
poisson_deviance_vs_champion pass-0.01126relative worsening <= 0.05

Canary pass

traffic
10% canary: 372 canary / 3228 control requests
error rate
0
p95 latency
0.0365 ms
prediction PSI vs shadow baseline
0.0259
poisson_deviance vs control
candidate 0.7661 vs champion 0.7797
evidence
canary_summary.json
guardrailstatusvaluethreshold
error_rate pass0<= 0.02
p95_latency_ms pass0.03651<= 50
prediction_psi_vs_shadow pass0.02587<= 0.2
poisson_deviance_vs_control pass-0.01753relative worsening <= 0.1

Drift

Drift below is injected on purpose into simulated traffic (synthetic: true).

Run incident

Drift status per signal and window, run incidentRows are signals, columns are time windows; each cell is ok, watch or alert.678910111213AreaArea / window 6: ok✔Area / window 7: ok✔Area / window 8: ok✔Area / window 9: ok✔Area / window 10: ok✔Area / window 11: ok✔Area / window 12: ok✔Area / window 13: ok✔BonusMalusBonusMalus / window 6: ok✔BonusMalus / window 7: ok✔BonusMalus / window 8: ok✔BonusMalus / window 9: ok✔BonusMalus / window 10: watch⚠wBonusMalus / window 11: alert✖ABonusMalus / window 12: alert✖ABonusMalus / window 13: alert✖ADensityDensity / window 6: ok✔Density / window 7: ok✔Density / window 8: ok✔Density / window 9: ok✔Density / window 10: ok✔Density / window 11: ok✔Density / window 12: ok✔Density / window 13: ok✔DrivAgeDrivAge / window 6: ok✔DrivAge / window 7: ok✔DrivAge / window 8: ok✔DrivAge / window 9: watch⚠wDrivAge / window 10: alert✖ADrivAge / window 11: alert✖ADrivAge / window 12: alert✖ADrivAge / window 13: alert✖ARegionRegion / window 6: ok✔Region / window 7: ok✔Region / window 8: ok✔Region / window 9: ok✔Region / window 10: ok✔Region / window 11: ok✔Region / window 12: ok✔Region / window 13: ok✔VehAgeVehAge / window 6: ok✔VehAge / window 7: ok✔VehAge / window 8: ok✔VehAge / window 9: ok✔VehAge / window 10: ok✔VehAge / window 11: ok✔VehAge / window 12: ok✔VehAge / window 13: ok✔VehBrandVehBrand / window 6: ok✔VehBrand / window 7: ok✔VehBrand / window 8: ok✔VehBrand / window 9: ok✔VehBrand / window 10: ok✔VehBrand / window 11: ok✔VehBrand / window 12: ok✔VehBrand / window 13: ok✔VehGasVehGas / window 6: ok✔VehGas / window 7: ok✔VehGas / window 8: ok✔VehGas / window 9: ok✔VehGas / window 10: ok✔VehGas / window 11: ok✔VehGas / window 12: ok✔VehGas / window 13: ok✔VehPowerVehPower / window 6: ok✔VehPower / window 7: ok✔VehPower / window 8: ok✔VehPower / window 9: ok✔VehPower / window 10: ok✔VehPower / window 11: ok✔VehPower / window 12: ok✔VehPower / window 13: ok✔predictionprediction / window 6: ok✔prediction / window 7: ok✔prediction / window 8: ok✔prediction / window 9: watch⚠wprediction / window 10: alert✖Aprediction / window 11: alert✖Aprediction / window 12: alert✖Aprediction / window 13: alert✖Avalidation failuresvalidation failures / window 6: ok✔validation failures / window 7: ok✔validation failures / window 8: ok✔validation failures / window 9: ok✔validation failures / window 10: ok✔validation failures / window 11: ok✔validation failures / window 12: ok✔validation failures / window 13: ok✔labels (A/E)labels (A/E) / window 6: ok✔labels (A/E) / window 7: ok✔labels (A/E) / window 8: ok✔labels (A/E) / window 9: ok✔labels (A/E) / window 10: ok✔labels (A/E) / window 11: ok✔labels (A/E) / window 12: ok✔labels (A/E) / window 13: ok✔WINDOWWINDOW / window 6: ok✔WINDOW / window 7: ok✔WINDOW / window 8: ok✔WINDOW / window 9: watch⚠wWINDOW / window 10: alert✖AWINDOW / window 11: alert✖AWINDOW / window 12: alert✖AWINDOW / window 13: alert✖A
windowlevelrowsstart (UTC)scenarioprediction PSIvalidation failure ratelabel drift (rel. change)evidence
6 ok12002026-10-08T01:07:06Znone (synthetic)0.013400.044 (n.s.)json · Evidently report
7 ok12002026-10-08T01:07:06Znone (synthetic)0.0085900.004 (n.s.)json · Evidently report
8 ok12002026-10-08T01:07:06Znone (synthetic)0.005300.102 (n.s.)json · Evidently report
9 watch12002026-10-08T01:07:06Zyoung_driver_surge (synthetic)0.18900.065 (n.s.)json · Evidently report
10 alert12002026-10-08T01:07:06Zyoung_driver_surge (synthetic)0.76600.071 (n.s.)json · Evidently report
11 alert12002026-10-08T01:07:07Zyoung_driver_surge (synthetic)1.8800.039 (n.s.)json · Evidently report
12 alert12002026-10-08T01:07:12Zyoung_driver_surge (synthetic)1.9300.001 (n.s.)json · Evidently report
13 alert12002026-10-08T01:07:12Zyoung_driver_surge (synthetic)1.8300.087json · Evidently report

Per-feature detail at the worst window (13)

featurelevelPSIKS statisticnew-category rateout-of-range rate
Area ok0.00715–0–
BonusMalus alert0.4980.266–0.00833
Density ok0.03140.0433–0
DrivAge alert2.090.652–0
Region ok0.0463–0–
VehAge ok0.0210.0228–0
VehBrand ok0.0202–0–
VehGas ok9e-06–0–
VehPower ok0.02280.0577–0

Scenario matrix: declared vs observed (synthetic)

scenariowhat changesdeclared expectationwindow levelsresult
covariate_shiftDrivAge shifts upward by ~2 standard deviationsalert on feature:DrivAgeok → ok → alert → alert → alert → alert detected
new_categoryunseen levels appear in Areaalert on feature:Areaok → ok → alert → alert → alert → alert detected
schema_breakDrivAge is sent as a stringalert on validationok → ok → alert → alert → alert → alert detected
label_shiftdelayed labels show a higher base ratealert on labelok → ok → alert → alert → alert → alert detected
silent_decayVehicle age drifts up slowly (older vehicles gradually over-represented)alert on feature:VehAge, watch before alertok → ok → ok → ok → ok → ok → watch → watch → alert → alert detected
young_driver_surgeA younger cohort enters the book: young drivers are over-represented in trafficalert on feature:DrivAgeok → ok → ok → alert → alert → alert → alert detected
new_vehicle_brandUnseen VehBrand values appearalert on feature:VehBrandok → ok → alert → alert → alert → alert detected
bonus_malus_schema_breakBonusMalus arrives as a stringalert on validationok → ok → alert → alert → alert → alert detected
claims_frequency_shiftDelayed labels show a higher base claim ratealert on labelok → ok → alert → watch → alert → alert detected

Every declared expectation was met in this run. See docs/MONITORING.md for sensitivity limits.

Incident story

Synthetic drift on simulated traffic, driven by the audit log and drift reports.

  1. Healthy champion
    Window 6, champion 5fc4383bbf, level ok. A/E 1.04, decile calibration error 0.115, deviance 0.8893 (n=1200)
  2. Drift begins (synthetic)
    Scenario young_driver_surge starts at window 9 with intensity 0.09; level watch.
  3. Monitor reaches watch
    Window 9 (intensity 0.09): watch. A/E 0.935, decile calibration error 0.0841, deviance 0.9643 (n=1200)
  4. Monitor reaches alert
    Window 10 (intensity 0.3, 2026-10-08T01:07:06Z): alert. A/E 0.929, decile calibration error 0.162, deviance 1.175 (n=1200)
  5. Rollback recommended
    Audit #4 at 2026-10-08T01:07:11Z: Drift alert in window 10 (drivers: DrivAge, prediction). Rollback to previous champion c2c40600b7290013 recommended.
  6. Rollback executed
    Audit #5 at 2026-10-08T01:07:12Z: 5fc4383bbf → c2c40600b7. Drift alert from synthetic scenario 'young_driver_surge'; restoring previous champion per runbook.
  7. Restored state
    Window 12: serving c2c40600b7, level alert. A/E 0.999, decile calibration error 0.141, deviance 1.416 (n=1200). The input drift is still present after rollback: rolling back restores the previous known-good model, it does not fix the data.
Actual / expected claims per windowA/E for the serving champion by window; shaded bands mark before, during and after.0.8870.9290.9711.011.06678910111213A/E (champion in service) @ 6: 1.044A/E (champion in service) @ 7: 1.004A/E (champion in service) @ 8: 0.8984A/E (champion in service) @ 9: 0.9355A/E (champion in service) @ 10: 0.9286A/E (champion in service) @ 11: 0.9612A/E (champion in service) @ 12: 0.9985A/E (champion in service) @ 13: 0.9134window (bands: before | drift | after rollback)A/E━ A/E (champion in service)
Data table for the chart
windowphaselevelserving championA/Edecile errorrows
6before ok5fc4383bbf1.040.1151200
7before ok5fc4383bbf10.06231200
8during ok5fc4383bbf0.8980.1591200
9during watch5fc4383bbf0.9350.08411200
10during alert5fc4383bbf0.9290.1621200
11during alert5fc4383bbf0.9610.1411200
12after alertc2c40600b70.9990.1411200
13after alertc2c40600b70.9130.1371200

Calibration: Before (healthy), window 6

Calibration by decile: Before (healthy)decile 1: predicted 0.0798, observed 0.1355; decile 2: predicted 0.1003, observed 0.0946; decile 3: predicted 0.1160, observed 0.2031; decile 4: predicted 0.1310, observed 0.1751; decile 5: predicted 0.1476, observed 0.1228; decile 6: predicted 0.1714, observed 0.1214; decile 7: predicted 0.2032, observed 0.2026; decile 8: predicted 0.2497, observed 0.2445; decile 9: predicted 0.3841, observed 0.4457; decile 10: predicted 1.4470, observed 1.41760.0000.0000.3910.3910.7810.7811.1721.1721.5631.563decile 1: predicted 0.0798, observed 0.1355decile 2: predicted 0.1003, observed 0.0946decile 3: predicted 0.1160, observed 0.2031decile 4: predicted 0.1310, observed 0.1751decile 5: predicted 0.1476, observed 0.1228decile 6: predicted 0.1714, observed 0.1214decile 7: predicted 0.2032, observed 0.2026decile 8: predicted 0.2497, observed 0.2445decile 9: predicted 0.3841, observed 0.4457decile 10: predicted 1.4470, observed 1.4176predicted frequencyobserved frequency

Calibration: At first alert, window 10

Calibration by decile: At first alertdecile 1: predicted 0.0871, observed 0.1312; decile 2: predicted 0.1195, observed 0.1180; decile 3: predicted 0.1550, observed 0.1048; decile 4: predicted 0.2173, observed 0.2211; decile 5: predicted 0.3533, observed 0.2775; decile 6: predicted 0.5987, observed 0.4536; decile 7: predicted 0.9376, observed 0.7194; decile 8: predicted 1.2148, observed 1.2013; decile 9: predicted 1.5810, observed 1.8442; decile 10: predicted 2.6535, observed 2.28050.0000.0000.7160.7161.4331.4332.1492.1492.8662.866decile 1: predicted 0.0871, observed 0.1312decile 2: predicted 0.1195, observed 0.1180decile 3: predicted 0.1550, observed 0.1048decile 4: predicted 0.2173, observed 0.2211decile 5: predicted 0.3533, observed 0.2775decile 6: predicted 0.5987, observed 0.4536decile 7: predicted 0.9376, observed 0.7194decile 8: predicted 1.2148, observed 1.2013decile 9: predicted 1.5810, observed 1.8442decile 10: predicted 2.6535, observed 2.2805predicted frequencyobserved frequency

Calibration: After rollback, window 12

Calibration by decile: After rollbackdecile 1: predicted 0.0997, observed 0.0936; decile 2: predicted 0.1944, observed 0.3080; decile 3: predicted 0.3619, observed 0.4235; decile 4: predicted 0.6810, observed 0.4821; decile 5: predicted 0.9011, observed 1.0101; decile 6: predicted 1.0618, observed 1.1298; decile 7: predicted 1.2806, observed 1.4542; decile 8: predicted 1.5594, observed 1.8198; decile 9: predicted 1.8724, observed 1.6512; decile 10: predicted 3.2433, observed 2.86650.0000.0000.8760.8761.7511.7512.6272.6273.5033.503decile 1: predicted 0.0997, observed 0.0936decile 2: predicted 0.1944, observed 0.3080decile 3: predicted 0.3619, observed 0.4235decile 4: predicted 0.6810, observed 0.4821decile 5: predicted 0.9011, observed 1.0101decile 6: predicted 1.0618, observed 1.1298decile 7: predicted 1.2806, observed 1.4542decile 8: predicted 1.5594, observed 1.8198decile 9: predicted 1.8724, observed 1.6512decile 10: predicted 3.2433, observed 2.8665predicted frequencyobserved frequency

Drift is injected into simulated traffic; nothing here is real-world behavior.

Model card

Card for candidate c2c40600b7290013 (the model most recently in service among those with a card). Generated, not hand-written.

Model card: claim frequency glm / candidate c2c40600b7290013

Synthetic data. This candidate was trained on a generated stand-in dataset, not real insurance data. Numbers below describe the stand-in only.

Intended use

Educational and portfolio demonstration of a gated model release process for a motor third-party-liability claim-frequency model (expected claims per exposure-year). It exists to exercise gates, shadow/canary, drift monitoring and rollback.

Out-of-scope use

Not for pricing, underwriting, reserving, or any decision about real people. No claim of regulatory compliance, actuarial sign-off or fairness certification. Severity (claim cost) is not modeled, so nothing here estimates expected loss.

Data

  • Source: Generated by examples/claims_frequency/data.py (labeled synthetic).
  • Data version: synthetic-7c8426d485-n40000
  • Rows per split: {'train': 24001, 'val': 4000, 'holdout': 6000, 'stream': 5999}
  • Split: policy- and risk-profile-grouped, claim-stratified; NO temporal split (dataset has no dates)
  • Known gaps: the dataset has no timestamps. A time-based split is impossible, so temporal generalization was not evaluated. Simulated production traffic replays held-out policies in random order and is not real production behavior.
  • One row per policy; severity data not used; exposure is capped at 1 year, claim counts at 4 (row counts per cleaning rule are in the data-quality report).

Training procedure

  • Family: glm; seed 20260101; params: {"alpha": 0.0001, "recalibrate": true}
  • Target: claim rate (claims per exposure-year), Poisson loss, exposure as sample weight (equivalent to an offset). Tuning used grouped cross-validation inside the training split only; the holdout is evaluated once per candidate.
  • Optional scale-factor recalibration is fitted on the validation split and is part of the candidate (and its hash).

Metrics

Evaluated on the holdout, exposure-weighted, 95% bootstrap CIs (policies resampled).

metricvalue95% CI
poisson_deviance0.80400[0.76413, 0.84317]
deviance_skill0.32327[0.29062, 0.34725]
gini0.55070[0.52279, 0.58302]

Slices and calibration

Calibration (exposure-weighted decile error, relative to portfolio frequency): 0.05228; overall actual/expected: 0.98497.

slicelevelnA/E95% CIflag
driver_age25-3510111.004insufficient data
driver_age35-4514810.881insufficient data
driver_age45-5514770.982insufficient data
driver_age55-659691.208insufficient data
driver_age65-754200.740insufficient data
driver_age<255131.001insufficient data
driver_age>=751290.540insufficient data
vehicle_age1-310940.871insufficient data
vehicle_age10-159331.080insufficient data
vehicle_age3-617811.051insufficient data
vehicle_age6-1015900.947insufficient data
vehicle_age<12021.130insufficient data
vehicle_age>=154000.920insufficient data
regionR212990.887insufficient data
regionR312571.190insufficient data
regionR422810.891insufficient data
regionR522930.811insufficient data
regionR722930.910insufficient data
regionR732921.301insufficient data
regionR942591.223insufficient data
regionother40260.966[0.909, 1.009]
bonus_malus55-7028840.961[0.882, 1.035]
bonus_malus70-907461.066insufficient data
bonus_malus90-1101041.144insufficient data
bonus_malus<5522410.959[0.841, 1.102]
bonus_malus>=110250.595insufficient data
exposure0.25-0.7532001.010[0.944, 1.090]
exposure0.75-120830.956[0.874, 1.027]
exposure<0.257151.024insufficient data
exposure>=123.266insufficient data

Explainability

Describes the fitted model, not causal effects. Computed on a validation sample, not the holdout.

Permutation importance (increase in exposure-weighted deviance when a feature is shuffled):

featuredeviance increase
DrivAge0.424056
BonusMalus0.085794
Density0.008884
VehAge0.008623
VehGas0.007170
VehPower0.004453
Exposure0.004240
Area0.002342
VehBrand0.002023
Region-0.000998

Largest GLM coefficients (rate ratio = exp(coef)):

termcoefrate ratio
bins__DrivAge_21.76905.865
bins__DrivAge_11.71305.546
bins__BonusMalus_81.48614.420
bins__BonusMalus_1-1.08050.339
bins__BonusMalus_70.90582.474
bins__BonusMalus_2-0.82060.440
bins__DrivAge_10-0.67620.509
bins__DrivAge_11-0.61750.539
bins__DrivAge_9-0.61050.543
bins__BonusMalus_3-0.56260.570
bins__DrivAge_13-0.51960.595
bins__DrivAge_12-0.47420.622
bins__DrivAge_14-0.45210.636
bins__DrivAge_30.44741.564
bins__DrivAge_40.41701.517

Partial dependence, DrivAge: 18.0: 1.1604, 29.0: 0.3273, 36.0: 0.1676, 42.0: 0.1562, 47.0: 0.1709, 53.0: 0.1662, 60.0: 0.1064, 75.0: 0.1244

Partial dependence, BonusMalus: 50.0: 0.1718, 51.0: 0.2228, 53.0: 0.2228, 56.0: 0.2228, 58.0: 0.2228, 62.0: 0.2883, 68.0: 0.2883, 89.0: 0.5429

Partial dependence, VehAge: 0.0: 0.3219, 2.0: 0.3605, 3.0: 0.3080, 5.0: 0.2779, 6.0: 0.2779, 8.0: 0.2717, 11.0: 0.2273, 20.0: 0.2331

Partial dependence, Density: 5.0: 0.1943, 39.27: 0.2374, 91.0: 0.2589, 170.0: 0.2764, 331.0: 0.2963, 672.44: 0.3192, 1539.73: 0.3482, 9043.6: 0.4195

Limitations

  • No timestamps: temporal generalization, seasonality and trend were not evaluated.
  • Simulated traffic and drift: shadow/canary/drift results come from replaying held-out policies and from synthetic perturbations; they are demonstrations of the machinery, not evidence about real production.
  • Frequency only; no severity. One country/product/period; extrapolation elsewhere is untested.
  • Small slices are flagged insufficient data and not judged. Bootstrap CIs resample policies and ignore policy-to-policy dependence.
  • The audit log is a local file: tamper-evident, not tamper-proof.
Sensitive features and proxy risk

DrivAge and Region (and, indirectly, Density/Area) can act as proxies for protected characteristics (age, and in some jurisdictions ethnicity, religion or socio-economic status). They are kept for realism and reported by slice; this card does not claim the model is fair, compliant, or fit for pricing or underwriting.

Provenance

fieldvalue
candidate hashc2c40600b7290013
git SHAbcbdf51a7d5d
data versionsynthetic-7c8426d485-n40000
seed20260101
lockfile hash325e385df3aac6ed
gate config hash7a4a29aea97dd405
Card for candidate 5fc4383bbf

Model card: claim frequency gbm / candidate 5fc4383bbfa34776

Synthetic data. This candidate was trained on a generated stand-in dataset, not real insurance data. Numbers below describe the stand-in only.

Intended use

Educational and portfolio demonstration of a gated model release process for a motor third-party-liability claim-frequency model (expected claims per exposure-year). It exists to exercise gates, shadow/canary, drift monitoring and rollback.

Out-of-scope use

Not for pricing, underwriting, reserving, or any decision about real people. No claim of regulatory compliance, actuarial sign-off or fairness certification. Severity (claim cost) is not modeled, so nothing here estimates expected loss.

Data

  • Source: Generated by examples/claims_frequency/data.py (labeled synthetic).
  • Data version: synthetic-7c8426d485-n40000
  • Rows per split: {'train': 24001, 'val': 4000, 'holdout': 6000, 'stream': 5999}
  • Split: policy- and risk-profile-grouped, claim-stratified; NO temporal split (dataset has no dates)
  • Known gaps: the dataset has no timestamps. A time-based split is impossible, so temporal generalization was not evaluated. Simulated production traffic replays held-out policies in random order and is not real production behavior.
  • One row per policy; severity data not used; exposure is capped at 1 year, claim counts at 4 (row counts per cleaning rule are in the data-quality report).

Training procedure

  • Family: gbm; seed 20260101; params: {"n_estimators": 600, "learning_rate": 0.05, "num_leaves": 16, "min_child_samples": 400, "subsample": 0.8, "subsample_freq": 1, "colsample_bytree": 0.8, "reg_lambda": 5.0, "num_threads": 4, "recalibrate": true}
  • Target: claim rate (claims per exposure-year), Poisson loss, exposure as sample weight (equivalent to an offset). Tuning used grouped cross-validation inside the training split only; the holdout is evaluated once per candidate.
  • Optional scale-factor recalibration is fitted on the validation split and is part of the candidate (and its hash).

Metrics

Evaluated on the holdout, exposure-weighted, 95% bootstrap CIs (policies resampled).

metricvalue95% CI
deviance_skill0.32472[0.29245, 0.34820]
gini0.55080[0.52287, 0.58032]
poisson_deviance0.80227[0.76450, 0.83489]
Three-model comparison (same holdout)
modelcandidatedevianceCIskill vs constantGiniCI
Constant baselinec7e902deae28e7081.18807[1.11657, 1.23138]0.000000.00000[0.00000, 0.00000]
Poisson GLMc2c40600b72900130.80400[0.76413, 0.84317]0.323270.55070[0.52279, 0.58302]
LightGBM5fc4383bbfa347760.80227[0.76450, 0.83489]0.324720.55080[0.52287, 0.58032]

Lowest deviance: LightGBM. Differences smaller than the CI width are not evidence of a real difference; see the paired-bootstrap G3 check for the champion/challenger decision. Holdout is one random grouped split, not a time split.

Slices and calibration

Calibration (exposure-weighted decile error, relative to portfolio frequency): 0.08826; overall actual/expected: 0.99463.

slicelevelnA/E95% CIflag
bonus_malus55-7028840.940[0.859, 1.010]
bonus_malus70-907461.094insufficient data
bonus_malus90-1101041.119insufficient data
bonus_malus<5522410.958[0.843, 1.098]
bonus_malus>=110250.830insufficient data
driver_age25-3510111.017insufficient data
driver_age35-4514810.828insufficient data
driver_age45-5514770.983insufficient data
driver_age55-659691.067insufficient data
driver_age65-754200.713insufficient data
driver_age<255131.057insufficient data
driver_age>=751290.522insufficient data
exposure0.25-0.7532001.027[0.959, 1.100]
exposure0.75-120830.960[0.875, 1.032]
exposure<0.257151.003insufficient data
exposure>=122.575insufficient data
regionR212990.912insufficient data
regionR312571.168insufficient data
regionR422810.892insufficient data
regionR522930.836insufficient data
regionR722930.945insufficient data
regionR732921.304insufficient data
regionR942591.247insufficient data
regionother40260.973[0.916, 1.020]
vehicle_age1-310940.970insufficient data
vehicle_age10-159331.050insufficient data
vehicle_age3-617811.036insufficient data
vehicle_age6-1015900.938insufficient data
vehicle_age<12021.128insufficient data
vehicle_age>=154000.905insufficient data

Explainability

Describes the fitted model, not causal effects. Computed on a validation sample, not the holdout.

Permutation importance (increase in exposure-weighted deviance when a feature is shuffled):

featuredeviance increase
DrivAge0.409138
BonusMalus0.081809
VehGas0.013334
Density0.006920
VehPower0.004433
VehAge0.004273
VehBrand0.001130
Area0.000184
Exposure0.000113
Region-0.001787

TreeSHAP via LightGBM pred_contrib (log-rate scale). Mean |SHAP| by feature:

featuremean abs SHAP
DrivAge0.4602
BonusMalus0.1996
Density0.1336
VehAge0.1002
VehPower0.0645
VehGas0.0496
Region0.0484
VehBrand0.0295
Area0.0131
Exposure0.0109

Partial dependence, DrivAge: 18.0: 1.1539, 29.0: 0.3415, 36.0: 0.1755, 42.0: 0.1715, 47.0: 0.1719, 53.0: 0.1696, 60.0: 0.1245, 75.0: 0.1275

Partial dependence, BonusMalus: 50.0: 0.2017, 51.0: 0.2017, 53.0: 0.2198, 56.0: 0.2389, 58.0: 0.2435, 62.0: 0.2706, 68.0: 0.3267, 89.0: 0.5546

Partial dependence, VehAge: 0.0: 0.3342, 2.0: 0.3332, 3.0: 0.3258, 5.0: 0.2925, 6.0: 0.2885, 8.0: 0.2785, 11.0: 0.2561, 20.0: 0.2509

Partial dependence, Density: 5.0: 0.2205, 39.27: 0.2462, 91.0: 0.2488, 170.0: 0.2833, 331.0: 0.2999, 672.44: 0.3231, 1539.73: 0.3477, 9043.6: 0.3729

Limitations

  • No timestamps: temporal generalization, seasonality and trend were not evaluated.
  • Simulated traffic and drift: shadow/canary/drift results come from replaying held-out policies and from synthetic perturbations; they are demonstrations of the machinery, not evidence about real production.
  • Frequency only; no severity. One country/product/period; extrapolation elsewhere is untested.
  • Small slices are flagged insufficient data and not judged. Bootstrap CIs resample policies and ignore policy-to-policy dependence.
  • The audit log is a local file: tamper-evident, not tamper-proof.
Sensitive features and proxy risk

DrivAge and Region (and, indirectly, Density/Area) can act as proxies for protected characteristics (age, and in some jurisdictions ethnicity, religion or socio-economic status). They are kept for realism and reported by slice; this card does not claim the model is fair, compliant, or fit for pricing or underwriting.

Provenance

fieldvalue
candidate hash5fc4383bbfa34776
git SHAbcbdf51a7d5d
data versionsynthetic-7c8426d485-n40000
seed20260101
lockfile hash325e385df3aac6ed
gate config hash7a4a29aea97dd405
Card for candidate 60528a9499

Model card: claim frequency gbm_degraded / candidate 60528a9499f3939d

Synthetic data. This candidate was trained on a generated stand-in dataset, not real insurance data. Numbers below describe the stand-in only.

Intended use

Educational and portfolio demonstration of a gated model release process for a motor third-party-liability claim-frequency model (expected claims per exposure-year). It exists to exercise gates, shadow/canary, drift monitoring and rollback.

Out-of-scope use

Not for pricing, underwriting, reserving, or any decision about real people. No claim of regulatory compliance, actuarial sign-off or fairness certification. Severity (claim cost) is not modeled, so nothing here estimates expected loss.

Data

  • Source: Generated by examples/claims_frequency/data.py (labeled synthetic).
  • Data version: synthetic-7c8426d485-n40000
  • Rows per split: {'train': 24001, 'val': 4000, 'holdout': 6000, 'stream': 5999}
  • Split: policy- and risk-profile-grouped, claim-stratified; NO temporal split (dataset has no dates)
  • Known gaps: the dataset has no timestamps. A time-based split is impossible, so temporal generalization was not evaluated. Simulated production traffic replays held-out policies in random order and is not real production behavior.
  • One row per policy; severity data not used; exposure is capped at 1 year, claim counts at 4 (row counts per cleaning rule are in the data-quality report).

Training procedure

  • Family: gbm_degraded; seed 20260101; params: {"n_estimators": 100, "learning_rate": 0.05, "num_leaves": 16, "min_child_samples": 400, "num_threads": 4, "recalibrate": false, "corrupt_preprocessing": true}
  • Target: claim rate (claims per exposure-year), Poisson loss, exposure as sample weight (equivalent to an offset). Tuning used grouped cross-validation inside the training split only; the holdout is evaluated once per candidate.
  • Optional scale-factor recalibration is fitted on the validation split and is part of the candidate (and its hash).

Metrics

Evaluated on the holdout, exposure-weighted, 95% bootstrap CIs (policies resampled).

metricvalue95% CI
poisson_deviance1.42458[1.30299, 1.49901]
deviance_skill-0.19907[-0.21953, -0.16500]
gini-0.16149[-0.19006, -0.12056]

Slices and calibration

Calibration (exposure-weighted decile error, relative to portfolio frequency): 0.64306; overall actual/expected: 2.35865.

slicelevelnA/E95% CIflag
driver_age25-3510113.212insufficient data
driver_age35-4514811.194insufficient data
driver_age45-5514771.307insufficient data
driver_age55-659690.919insufficient data
driver_age65-754200.464insufficient data
driver_age<2551314.948insufficient data
driver_age>=751290.249insufficient data
vehicle_age1-310942.134insufficient data
vehicle_age10-159332.489insufficient data
vehicle_age3-617812.409insufficient data
vehicle_age6-1015902.492insufficient data
vehicle_age<12022.653insufficient data
vehicle_age>=154001.838insufficient data
regionR212992.179insufficient data
regionR312572.797insufficient data
regionR422812.340insufficient data
regionR522932.209insufficient data
regionR722931.855insufficient data
regionR732923.062insufficient data
regionR942593.204insufficient data
regionother40262.285[2.124, 2.480]
bonus_malus55-7028842.172[1.957, 2.345]
bonus_malus70-907466.359insufficient data
bonus_malus90-11010412.880insufficient data
bonus_malus<5522410.996[0.869, 1.165]
bonus_malus>=1102519.684insufficient data
exposure0.25-0.7532002.467[2.214, 2.706]
exposure0.75-120832.237[2.094, 2.415]
exposure<0.257152.587insufficient data
exposure>=123.928insufficient data

Explainability

Describes the fitted model, not causal effects. Computed on a validation sample, not the holdout.

Permutation importance (increase in exposure-weighted deviance when a feature is shuffled):

featuredeviance increase
Density0.003261
VehBrand0.001447
VehAge0.001337
Region0.000633
Area0.000491
Exposure0.000043
VehGas0.000000
VehPower-0.000070
BonusMalus-0.024224
DrivAge-0.031138

TreeSHAP via LightGBM pred_contrib (log-rate scale). Mean |SHAP| by feature:

featuremean abs SHAP
DrivAge0.3813
BonusMalus0.1830
Density0.0948
VehAge0.0763
VehPower0.0292
Region0.0291
VehGas0.0209
VehBrand0.0105
Area0.0038
Exposure0.0029

Partial dependence, DrivAge: 18.0: 0.1156, 29.0: 0.1156, 36.0: 0.1156, 42.0: 0.1156, 47.0: 0.1156, 53.0: 0.1228, 60.0: 0.1313, 75.0: 0.1852

Partial dependence, BonusMalus: 50.0: 0.1457, 51.0: 0.1457, 53.0: 0.1456, 56.0: 0.1190, 58.0: 0.1175, 62.0: 0.1172, 68.0: 0.1172, 89.0: 0.1172

Partial dependence, VehAge: 0.0: 0.1461, 2.0: 0.1461, 3.0: 0.1456, 5.0: 0.1215, 6.0: 0.1209, 8.0: 0.1190, 11.0: 0.1159, 20.0: 0.1159

Partial dependence, Density: 5.0: 0.1057, 39.27: 0.1131, 91.0: 0.1131, 170.0: 0.1312, 331.0: 0.1333, 672.44: 0.1386, 1539.73: 0.1407, 9043.6: 0.1407

Limitations

  • No timestamps: temporal generalization, seasonality and trend were not evaluated.
  • Simulated traffic and drift: shadow/canary/drift results come from replaying held-out policies and from synthetic perturbations; they are demonstrations of the machinery, not evidence about real production.
  • Frequency only; no severity. One country/product/period; extrapolation elsewhere is untested.
  • Small slices are flagged insufficient data and not judged. Bootstrap CIs resample policies and ignore policy-to-policy dependence.
  • The audit log is a local file: tamper-evident, not tamper-proof.
Sensitive features and proxy risk

DrivAge and Region (and, indirectly, Density/Area) can act as proxies for protected characteristics (age, and in some jurisdictions ethnicity, religion or socio-economic status). They are kept for realism and reported by slice; this card does not claim the model is fair, compliant, or fit for pricing or underwriting.

Provenance

fieldvalue
candidate hash60528a9499f3939d
git SHAbcbdf51a7d5d
data versionsynthetic-7c8426d485-n40000
seed20260101
lockfile hash325e385df3aac6ed
gate config hash7a4a29aea97dd405

About and limitations

  • No timestamps in the dataset. The public frequency dataset has no date column, so a time-based split is impossible. Temporal generalization was not evaluated.
  • Traffic is simulated. "Production" traffic replays held-out policies in random order. Shadow and canary results demonstrate the machinery, not real production behavior.
  • Drift is synthetic. Every scenario is injected on purpose and tagged synthetic: true.
  • The audit log is a local file. It is tamper-evident, not tamper-proof: anyone who can rewrite the whole file can recompute the chain. Anchor the head hash somewhere you trust if that matters.
  • This page is read-only. It is generated from JSON reports and the audit log, never writes to them, makes no network requests and loads no external resources.
  • No compliance or fairness claim. Driver age and region can proxy for protected characteristics; no actuarial sign-off, pricing or underwriting advice is implied.
  • Verification in the browser. The audit chain is verified when this page is built. The button below recomputes it in your browser with Web Crypto; if your browser does not provide Web Crypto for file:// pages, the button is hidden and the build-time result is the only one shown.