Overview
- Champion
c2c40600b7290013(glm)- Previous champion
- none
- Candidate
- none
- Stage
- stable
- Latest decision
- rollback seq 5, 2026-10-08T01:07:12Z: Drift alert from synthetic scenario 'young_driver_surge'; restoring previous champion per runbook.
- Audit events
- hold: 1, promote: 2, reject: 1, rollback: 1
- Audit chain
- verified (5 entries)
- Candidates gated
- 3 (2 passed, 1 rejected)
Model comparison (same holdout, 95% bootstrap CIs)
| model | candidate | deviance | deviance CI | skill vs constant | Gini | Gini CI |
|---|---|---|---|---|---|---|
| Constant baseline | c7e902deae | 1.1881 | [1.117, 1.231] | 0 | 0 | [0, 0] |
| Poisson GLM | c2c40600b7 | 0.804 | [0.7641, 0.8432] | 0.323 | 0.551 | [0.523, 0.583] |
| LightGBM | 5fc4383bbf | 0.80227 | [0.7645, 0.8349] | 0.325 | 0.551 | [0.523, 0.58] |
Lowest deviance: LightGBM. Differences smaller than the CI width are not evidence of a real difference; see the paired-bootstrap G3 check for the champion/challenger decision. Holdout is one random grouped split, not a time split.
Release timeline
| # | time (UTC) | event | actor | reason and evidence | candidate | champion before → after | flags |
|---|---|---|---|---|---|---|---|
| 1 | 2026-10-08T01:07:02Z | promote | demo | Bootstrap promotion: no champion existed, so G3 was skipped and shadow/canary had nothing to compare against.evidence (1)
| c2c40600b7 | none → c2c40600b7 | synthetic |
| 2 | 2026-10-08T01:07:03Z | reject | proving-ground | Gate(s) G2, G3, G4, G5 failed; candidate rejected.evidence (1)
| 60528a9499 | c2c40600b7 → c2c40600b7 | synthetic |
| 3 | 2026-10-08T01:07:06Z | promote | demo | Passed gates, shadow and canary; promoted to champion.evidence (3)
| 5fc4383bbf | c2c40600b7 → 5fc4383bbf | synthetic |
| 4 | 2026-10-08T01:07:11Z | hold | proving-ground | Drift alert in window 10 (drivers: DrivAge, prediction). Rollback to previous champion c2c40600b7290013 recommended.evidence (1)
| 5fc4383bbf | 5fc4383bbf → 5fc4383bbf | synthetic |
| 5 | 2026-10-08T01:07:12Z | rollback | demo | Drift alert from synthetic scenario 'young_driver_surge'; restoring previous champion per runbook.evidence (2)
| 5fc4383bbf | 5fc4383bbf → c2c40600b7 | synthetic |
Hashes shortened to 10 characters; full values in the JSON evidence.
Gate report
Candidate 5fc4383bbfa34776 (gbm) pass
- data version
synthetic-7c8426d485-n40000- git
bcbdf51a7d5d- seed
- 20260101
- gate config
7a4a29aea97dd405- lockfile
325e385df3aac6ed- evidence
- gate_report.json · gate_report.md
Candidate vs champion (holdout, exposure-weighted)
| metric | candidate | 95% CI | champion |
|---|---|---|---|
| deviance_skill | 0.32472 | [0.2924, 0.3482] | 0.32327 |
| gini | 0.5508 | [0.5229, 0.5803] | 0.5507 |
| poisson_deviance | 0.80227 | [0.7645, 0.8349] | 0.804 |
All checks
| gate | check | status | measured | threshold | 95% CI | champion | detail |
|---|---|---|---|---|---|---|---|
| G1 | Data validation | pass | |||||
| schema[train] | pass | – | – | – | – | 24001 rows valid | |
| schema[val] | pass | – | – | – | – | 4000 rows valid | |
| schema[holdout] | pass | – | – | – | – | 6000 rows valid | |
| row_floor[train] | pass | 24001 | 10000 | – | – | ||
| row_floor[val] | pass | 4000 | 1000 | – | – | ||
| row_floor[holdout] | pass | 6000 | 1000 | – | – | ||
| null_rate | pass | 0 | 0.0 | – | – | ||
| split_leakage | pass | 0 shared ids | 0 | – | – | no id may appear in more than one split | |
| G2 | Metric floors | pass | |||||
| deviance_skill | pass | 0.3247 | min 0.005 | [0.2924, 0.3482] | – | ||
| gini | pass | 0.5508 | min 0.1 | [0.5229, 0.5803] | – | ||
| poisson_deviance | pass | 0.8023 | max 0.9 | [0.7645, 0.8349] | – | ||
| G3 | No regression vs champion | pass | |||||
| poisson_deviance | pass | 0.8023 | relative worsening <= 0.02 | [-0.01198, 0.005411] | 0.804 | relative worsening -0.2143% (95% CI of worsening shown) | |
| gini | pass | 0.5508 | relative worsening <= 0.1 | [-0.01276, 0.01367] | 0.5507 | relative worsening -0.0170% (95% CI of worsening shown) | |
| G4 | Calibration | pass | |||||
| decile_calibration_error | pass | 0.08826 | max 0.35 | – | – | ||
| G5 | Slice performance | warn | |||||
| bonus_malus=55-70 | pass | 0.9403 | |gap| <= 0.45 vs 1.0 | [0.8587, 1.01] | – | n=2884, gap 6.0% | |
| bonus_malus=70-90 | insufficient | insufficient data | n >= 2000 | – | – | n=746: slice too small to judge; reported, not passed silently | |
| bonus_malus=90-110 | insufficient | insufficient data | n >= 2000 | – | – | n=104: slice too small to judge; reported, not passed silently | |
| bonus_malus=<55 | pass | 0.9577 | |gap| <= 0.45 vs 1.0 | [0.8426, 1.098] | – | n=2241, gap 4.2% | |
| bonus_malus=>=110 | insufficient | insufficient data | n >= 2000 | – | – | n=25: slice too small to judge; reported, not passed silently | |
| driver_age=25-35 | insufficient | insufficient data | n >= 2000 | – | – | n=1011: slice too small to judge; reported, not passed silently | |
| driver_age=35-45 | insufficient | insufficient data | n >= 2000 | – | – | n=1481: slice too small to judge; reported, not passed silently | |
| driver_age=45-55 | insufficient | insufficient data | n >= 2000 | – | – | n=1477: slice too small to judge; reported, not passed silently | |
| driver_age=55-65 | insufficient | insufficient data | n >= 2000 | – | – | n=969: slice too small to judge; reported, not passed silently | |
| driver_age=65-75 | insufficient | insufficient data | n >= 2000 | – | – | n=420: slice too small to judge; reported, not passed silently | |
| driver_age=<25 | insufficient | insufficient data | n >= 2000 | – | – | n=513: slice too small to judge; reported, not passed silently | |
| driver_age=>=75 | insufficient | insufficient data | n >= 2000 | – | – | n=129: slice too small to judge; reported, not passed silently | |
| exposure=0.25-0.75 | pass | 1.027 | |gap| <= 0.45 vs 1.0 | [0.9592, 1.1] | – | n=3200, gap 2.7% | |
| exposure=0.75-1 | pass | 0.9596 | |gap| <= 0.45 vs 1.0 | [0.8747, 1.032] | – | n=2083, gap 4.0% | |
| exposure=<0.25 | insufficient | insufficient data | n >= 2000 | – | – | n=715: slice too small to judge; reported, not passed silently | |
| exposure=>=1 | insufficient | insufficient data | n >= 2000 | – | – | n=2: slice too small to judge; reported, not passed silently | |
| region=R21 | insufficient | insufficient data | n >= 2000 | – | – | n=299: slice too small to judge; reported, not passed silently | |
| region=R31 | insufficient | insufficient data | n >= 2000 | – | – | n=257: slice too small to judge; reported, not passed silently | |
| region=R42 | insufficient | insufficient data | n >= 2000 | – | – | n=281: slice too small to judge; reported, not passed silently | |
| region=R52 | insufficient | insufficient data | n >= 2000 | – | – | n=293: slice too small to judge; reported, not passed silently | |
| region=R72 | insufficient | insufficient data | n >= 2000 | – | – | n=293: slice too small to judge; reported, not passed silently | |
| region=R73 | insufficient | insufficient data | n >= 2000 | – | – | n=292: slice too small to judge; reported, not passed silently | |
| region=R94 | insufficient | insufficient data | n >= 2000 | – | – | n=259: slice too small to judge; reported, not passed silently | |
| region=other | pass | 0.9725 | |gap| <= 0.45 vs 1.0 | [0.9157, 1.02] | – | n=4026, gap 2.7% | |
| vehicle_age=1-3 | insufficient | insufficient data | n >= 2000 | – | – | n=1094: slice too small to judge; reported, not passed silently | |
| vehicle_age=10-15 | insufficient | insufficient data | n >= 2000 | – | – | n=933: slice too small to judge; reported, not passed silently | |
| vehicle_age=3-6 | insufficient | insufficient data | n >= 2000 | – | – | n=1781: slice too small to judge; reported, not passed silently | |
| vehicle_age=6-10 | insufficient | insufficient data | n >= 2000 | – | – | n=1590: slice too small to judge; reported, not passed silently | |
| vehicle_age=<1 | insufficient | insufficient data | n >= 2000 | – | – | n=202: slice too small to judge; reported, not passed silently | |
| vehicle_age=>=15 | insufficient | insufficient data | n >= 2000 | – | – | n=400: slice too small to judge; reported, not passed silently | |
| G6 | Performance budget | pass | |||||
| p95_latency_ms | pass | 3.479 | max 60 | – | – | single-row predict(), n=100 | |
| artifact_mb | pass | 0.3 | max 50 | – | – | ||
| G7 | Reproducibility | pass | |||||
| same_hash | pass | 5fc4383bbfa34776 | 5fc4383bbfa34776 | – | – | identical hash on retrain with same seed | |
| G8 | Model card | pass | |||||
| card_exists | pass | – | – | – | – | ||
| card_matches_candidate | pass | – | – | – | – | card must name this candidate hash (regenerated per candidate) | |
| required_sections | pass | all present | 9 sections | – | – |
Candidate 60528a9499f3939d (gbm_degraded) fail failed: G2, G3, G4, G5
- data version
synthetic-7c8426d485-n40000- git
bcbdf51a7d5d- seed
- 20260101
- gate config
7a4a29aea97dd405- lockfile
325e385df3aac6ed- evidence
- gate_report.json · gate_report.md
Candidate vs champion (holdout, exposure-weighted)
| metric | candidate | 95% CI | champion |
|---|---|---|---|
| deviance_skill | -0.19907 | [-0.2195, -0.165] | 0.32327 |
| gini | -0.16149 | [-0.1901, -0.1206] | 0.5507 |
| poisson_deviance | 1.4246 | [1.303, 1.499] | 0.804 |
All checks
| gate | check | status | measured | threshold | 95% CI | champion | detail |
|---|---|---|---|---|---|---|---|
| G1 | Data validation | pass | |||||
| schema[train] | pass | – | – | – | – | 24001 rows valid | |
| schema[val] | pass | – | – | – | – | 4000 rows valid | |
| schema[holdout] | pass | – | – | – | – | 6000 rows valid | |
| row_floor[train] | pass | 24001 | 10000 | – | – | ||
| row_floor[val] | pass | 4000 | 1000 | – | – | ||
| row_floor[holdout] | pass | 6000 | 1000 | – | – | ||
| null_rate | pass | 0 | 0.0 | – | – | ||
| split_leakage | pass | 0 shared ids | 0 | – | – | no id may appear in more than one split | |
| G2 | Metric floors | fail | |||||
| deviance_skill | fail | -0.1991 | min 0.005 | [-0.2195, -0.165] | – | ||
| gini | fail | -0.1615 | min 0.1 | [-0.1901, -0.1206] | – | ||
| poisson_deviance | fail | 1.425 | max 0.9 | [1.303, 1.499] | – | ||
| G3 | No regression vs champion | fail | |||||
| poisson_deviance | fail | 1.425 | relative worsening <= 0.02 | [0.6452, 0.8515] | 0.804 | relative worsening +77.1867% (95% CI of worsening shown) | |
| gini | fail | -0.1615 | relative worsening <= 0.1 | [1.19, 1.39] | 0.5507 | relative worsening +129.3237% (95% CI of worsening shown) | |
| G4 | Calibration | fail | |||||
| decile_calibration_error | fail | 0.6431 | max 0.35 | – | – | ||
| G5 | Slice performance | fail | |||||
| driver_age=25-35 | insufficient | insufficient data | n >= 2000 | – | – | n=1011: slice too small to judge; reported, not passed silently | |
| driver_age=35-45 | insufficient | insufficient data | n >= 2000 | – | – | n=1481: slice too small to judge; reported, not passed silently | |
| driver_age=45-55 | insufficient | insufficient data | n >= 2000 | – | – | n=1477: slice too small to judge; reported, not passed silently | |
| driver_age=55-65 | insufficient | insufficient data | n >= 2000 | – | – | n=969: slice too small to judge; reported, not passed silently | |
| driver_age=65-75 | insufficient | insufficient data | n >= 2000 | – | – | n=420: slice too small to judge; reported, not passed silently | |
| driver_age=<25 | insufficient | insufficient data | n >= 2000 | – | – | n=513: slice too small to judge; reported, not passed silently | |
| driver_age=>=75 | insufficient | insufficient data | n >= 2000 | – | – | n=129: slice too small to judge; reported, not passed silently | |
| vehicle_age=1-3 | insufficient | insufficient data | n >= 2000 | – | – | n=1094: slice too small to judge; reported, not passed silently | |
| vehicle_age=10-15 | insufficient | insufficient data | n >= 2000 | – | – | n=933: slice too small to judge; reported, not passed silently | |
| vehicle_age=3-6 | insufficient | insufficient data | n >= 2000 | – | – | n=1781: slice too small to judge; reported, not passed silently | |
| vehicle_age=6-10 | insufficient | insufficient data | n >= 2000 | – | – | n=1590: slice too small to judge; reported, not passed silently | |
| vehicle_age=<1 | insufficient | insufficient data | n >= 2000 | – | – | n=202: slice too small to judge; reported, not passed silently | |
| vehicle_age=>=15 | insufficient | insufficient data | n >= 2000 | – | – | n=400: slice too small to judge; reported, not passed silently | |
| region=R21 | insufficient | insufficient data | n >= 2000 | – | – | n=299: slice too small to judge; reported, not passed silently | |
| region=R31 | insufficient | insufficient data | n >= 2000 | – | – | n=257: slice too small to judge; reported, not passed silently | |
| region=R42 | insufficient | insufficient data | n >= 2000 | – | – | n=281: slice too small to judge; reported, not passed silently | |
| region=R52 | insufficient | insufficient data | n >= 2000 | – | – | n=293: slice too small to judge; reported, not passed silently | |
| region=R72 | insufficient | insufficient data | n >= 2000 | – | – | n=293: slice too small to judge; reported, not passed silently | |
| region=R73 | insufficient | insufficient data | n >= 2000 | – | – | n=292: slice too small to judge; reported, not passed silently | |
| region=R94 | insufficient | insufficient data | n >= 2000 | – | – | n=259: slice too small to judge; reported, not passed silently | |
| region=other | fail | 2.285 | |gap| <= 0.45 vs 1.0 | [2.124, 2.48] | – | n=4026, gap 128.5% | |
| bonus_malus=55-70 | fail | 2.172 | |gap| <= 0.45 vs 1.0 | [1.957, 2.345] | – | n=2884, gap 117.2% | |
| bonus_malus=70-90 | insufficient | insufficient data | n >= 2000 | – | – | n=746: slice too small to judge; reported, not passed silently | |
| bonus_malus=90-110 | insufficient | insufficient data | n >= 2000 | – | – | n=104: slice too small to judge; reported, not passed silently | |
| bonus_malus=<55 | pass | 0.9957 | |gap| <= 0.45 vs 1.0 | [0.8687, 1.165] | – | n=2241, gap 0.4% | |
| bonus_malus=>=110 | insufficient | insufficient data | n >= 2000 | – | – | n=25: slice too small to judge; reported, not passed silently | |
| exposure=0.25-0.75 | fail | 2.467 | |gap| <= 0.45 vs 1.0 | [2.214, 2.706] | – | n=3200, gap 146.7% | |
| exposure=0.75-1 | fail | 2.237 | |gap| <= 0.45 vs 1.0 | [2.094, 2.415] | – | n=2083, gap 123.7% | |
| exposure=<0.25 | insufficient | insufficient data | n >= 2000 | – | – | n=715: slice too small to judge; reported, not passed silently | |
| exposure=>=1 | insufficient | insufficient data | n >= 2000 | – | – | n=2: slice too small to judge; reported, not passed silently | |
| G6 | Performance budget | pass | |||||
| p95_latency_ms | pass | 3.734 | max 60 | – | – | single-row predict(), n=100 | |
| artifact_mb | pass | 0.193 | max 50 | – | – | ||
| G7 | Reproducibility | pass | |||||
| same_hash | pass | 60528a9499f3939d | 60528a9499f3939d | – | – | identical hash on retrain with same seed | |
| G8 | Model card | pass | |||||
| card_exists | pass | – | – | – | – | ||
| card_matches_candidate | pass | – | – | – | – | card must name this candidate hash (regenerated per candidate) | |
| required_sections | pass | all present | 9 sections | – | – |
Candidate c2c40600b7290013 (glm) pass
- data version
synthetic-7c8426d485-n40000- git
bcbdf51a7d5d- seed
- 20260101
- gate config
7a4a29aea97dd405- lockfile
325e385df3aac6ed- evidence
- gate_report.json · gate_report.md
Candidate vs champion (holdout, exposure-weighted)
| metric | candidate | 95% CI | champion |
|---|---|---|---|
| deviance_skill | 0.32327 | [0.2906, 0.3472] | no champion at the time |
| gini | 0.5507 | [0.5228, 0.583] | no champion at the time |
| poisson_deviance | 0.804 | [0.7641, 0.8432] | no champion at the time |
All checks
| gate | check | status | measured | threshold | 95% CI | champion | detail |
|---|---|---|---|---|---|---|---|
| G1 | Data validation | pass | |||||
| schema[train] | pass | – | – | – | – | 24001 rows valid | |
| schema[val] | pass | – | – | – | – | 4000 rows valid | |
| schema[holdout] | pass | – | – | – | – | 6000 rows valid | |
| row_floor[train] | pass | 24001 | 10000 | – | – | ||
| row_floor[val] | pass | 4000 | 1000 | – | – | ||
| row_floor[holdout] | pass | 6000 | 1000 | – | – | ||
| null_rate | pass | 0 | 0.0 | – | – | ||
| split_leakage | pass | 0 shared ids | 0 | – | – | no id may appear in more than one split | |
| G2 | Metric floors | pass | |||||
| deviance_skill | pass | 0.3233 | min 0.005 | [0.2906, 0.3472] | – | ||
| gini | pass | 0.5507 | min 0.1 | [0.5228, 0.583] | – | ||
| poisson_deviance | pass | 0.804 | max 0.9 | [0.7641, 0.8432] | – | ||
| G3 | No regression vs champion | skip | No champion exists: first promotion. G3 skipped and recorded as skipped. | ||||
| regression | skip | – | – | – | – | No champion exists: first promotion. G3 skipped and recorded as skipped. | |
| G4 | Calibration | pass | |||||
| decile_calibration_error | pass | 0.05228 | max 0.35 | – | – | ||
| G5 | Slice performance | warn | |||||
| driver_age=25-35 | insufficient | insufficient data | n >= 2000 | – | – | n=1011: slice too small to judge; reported, not passed silently | |
| driver_age=35-45 | insufficient | insufficient data | n >= 2000 | – | – | n=1481: slice too small to judge; reported, not passed silently | |
| driver_age=45-55 | insufficient | insufficient data | n >= 2000 | – | – | n=1477: slice too small to judge; reported, not passed silently | |
| driver_age=55-65 | insufficient | insufficient data | n >= 2000 | – | – | n=969: slice too small to judge; reported, not passed silently | |
| driver_age=65-75 | insufficient | insufficient data | n >= 2000 | – | – | n=420: slice too small to judge; reported, not passed silently | |
| driver_age=<25 | insufficient | insufficient data | n >= 2000 | – | – | n=513: slice too small to judge; reported, not passed silently | |
| driver_age=>=75 | insufficient | insufficient data | n >= 2000 | – | – | n=129: slice too small to judge; reported, not passed silently | |
| vehicle_age=1-3 | insufficient | insufficient data | n >= 2000 | – | – | n=1094: slice too small to judge; reported, not passed silently | |
| vehicle_age=10-15 | insufficient | insufficient data | n >= 2000 | – | – | n=933: slice too small to judge; reported, not passed silently | |
| vehicle_age=3-6 | insufficient | insufficient data | n >= 2000 | – | – | n=1781: slice too small to judge; reported, not passed silently | |
| vehicle_age=6-10 | insufficient | insufficient data | n >= 2000 | – | – | n=1590: slice too small to judge; reported, not passed silently | |
| vehicle_age=<1 | insufficient | insufficient data | n >= 2000 | – | – | n=202: slice too small to judge; reported, not passed silently | |
| vehicle_age=>=15 | insufficient | insufficient data | n >= 2000 | – | – | n=400: slice too small to judge; reported, not passed silently | |
| region=R21 | insufficient | insufficient data | n >= 2000 | – | – | n=299: slice too small to judge; reported, not passed silently | |
| region=R31 | insufficient | insufficient data | n >= 2000 | – | – | n=257: slice too small to judge; reported, not passed silently | |
| region=R42 | insufficient | insufficient data | n >= 2000 | – | – | n=281: slice too small to judge; reported, not passed silently | |
| region=R52 | insufficient | insufficient data | n >= 2000 | – | – | n=293: slice too small to judge; reported, not passed silently | |
| region=R72 | insufficient | insufficient data | n >= 2000 | – | – | n=293: slice too small to judge; reported, not passed silently | |
| region=R73 | insufficient | insufficient data | n >= 2000 | – | – | n=292: slice too small to judge; reported, not passed silently | |
| region=R94 | insufficient | insufficient data | n >= 2000 | – | – | n=259: slice too small to judge; reported, not passed silently | |
| region=other | pass | 0.9656 | |gap| <= 0.45 vs 1.0 | [0.9093, 1.009] | – | n=4026, gap 3.4% | |
| bonus_malus=55-70 | pass | 0.9613 | |gap| <= 0.45 vs 1.0 | [0.8824, 1.035] | – | n=2884, gap 3.9% | |
| bonus_malus=70-90 | insufficient | insufficient data | n >= 2000 | – | – | n=746: slice too small to judge; reported, not passed silently | |
| bonus_malus=90-110 | insufficient | insufficient data | n >= 2000 | – | – | n=104: slice too small to judge; reported, not passed silently | |
| bonus_malus=<55 | pass | 0.9587 | |gap| <= 0.45 vs 1.0 | [0.8409, 1.102] | – | n=2241, gap 4.1% | |
| bonus_malus=>=110 | insufficient | insufficient data | n >= 2000 | – | – | n=25: slice too small to judge; reported, not passed silently | |
| exposure=0.25-0.75 | pass | 1.01 | |gap| <= 0.45 vs 1.0 | [0.9436, 1.09] | – | n=3200, gap 1.0% | |
| exposure=0.75-1 | pass | 0.9559 | |gap| <= 0.45 vs 1.0 | [0.8736, 1.027] | – | n=2083, gap 4.4% | |
| exposure=<0.25 | insufficient | insufficient data | n >= 2000 | – | – | n=715: slice too small to judge; reported, not passed silently | |
| exposure=>=1 | insufficient | insufficient data | n >= 2000 | – | – | n=2: slice too small to judge; reported, not passed silently | |
| G6 | Performance budget | pass | |||||
| p95_latency_ms | pass | 5.269 | max 60 | – | – | single-row predict(), n=100 | |
| artifact_mb | pass | 0.005 | max 50 | – | – | ||
| G7 | Reproducibility | pass | |||||
| same_hash | pass | c2c40600b7290013 | c2c40600b7290013 | – | – | identical hash on retrain with same seed | |
| G8 | Model card | pass | |||||
| card_exists | pass | – | – | – | – | ||
| card_matches_candidate | pass | – | – | – | – | card must name this candidate hash (regenerated per candidate) | |
| required_sections | pass | all present | 9 sections | – | – |
Calibration and slices
Candidate 5fc4383bbf (gbm) pass
- ae_ratio
- 0.9946
- decile_calibration_error
- 0.08826
Data table for the chart
| decile | exposure | predicted | observed |
|---|---|---|---|
| 1 | 370.72 | 0.07936 | 0.06744 |
| 2 | 371.99 | 0.09943 | 0.09947 |
| 3 | 370.49 | 0.115 | 0.07288 |
| 4 | 371.73 | 0.1303 | 0.1399 |
| 5 | 371.71 | 0.1482 | 0.1614 |
| 6 | 371.07 | 0.1702 | 0.09971 |
| 7 | 371.29 | 0.201 | 0.2155 |
| 8 | 371.89 | 0.256 | 0.2635 |
| 9 | 371.24 | 0.3993 | 0.3825 |
| 10 | 371.59 | 1.43 | 1.51 |
Slices (actual / expected claims; 1.0 is calibrated)
| slice | level | n | A/E | 95% CI | flag |
|---|---|---|---|---|---|
| bonus_malus | 55-70 | 2884 | 0.94 | [0.859, 1.01] | |
| bonus_malus | 70-90 | 746 | 1.09 | – | insufficient data |
| bonus_malus | 90-110 | 104 | 1.12 | – | insufficient data |
| bonus_malus | <55 | 2241 | 0.958 | [0.843, 1.1] | |
| bonus_malus | >=110 | 25 | 0.83 | – | insufficient data |
| driver_age | 25-35 | 1011 | 1.02 | – | insufficient data |
| driver_age | 35-45 | 1481 | 0.828 | – | insufficient data |
| driver_age | 45-55 | 1477 | 0.983 | – | insufficient data |
| driver_age | 55-65 | 969 | 1.07 | – | insufficient data |
| driver_age | 65-75 | 420 | 0.713 | – | insufficient data |
| driver_age | <25 | 513 | 1.06 | – | insufficient data |
| driver_age | >=75 | 129 | 0.522 | – | insufficient data |
| exposure | 0.25-0.75 | 3200 | 1.03 | [0.959, 1.1] | |
| exposure | 0.75-1 | 2083 | 0.96 | [0.875, 1.03] | |
| exposure | <0.25 | 715 | 1 | – | insufficient data |
| exposure | >=1 | 2 | 2.57 | – | insufficient data |
| region | R21 | 299 | 0.912 | – | insufficient data |
| region | R31 | 257 | 1.17 | – | insufficient data |
| region | R42 | 281 | 0.892 | – | insufficient data |
| region | R52 | 293 | 0.836 | – | insufficient data |
| region | R72 | 293 | 0.945 | – | insufficient data |
| region | R73 | 292 | 1.3 | – | insufficient data |
| region | R94 | 259 | 1.25 | – | insufficient data |
| region | other | 4026 | 0.973 | [0.916, 1.02] | |
| vehicle_age | 1-3 | 1094 | 0.97 | – | insufficient data |
| vehicle_age | 10-15 | 933 | 1.05 | – | insufficient data |
| vehicle_age | 3-6 | 1781 | 1.04 | – | insufficient data |
| vehicle_age | 6-10 | 1590 | 0.938 | – | insufficient data |
| vehicle_age | <1 | 202 | 1.13 | – | insufficient data |
| vehicle_age | >=15 | 400 | 0.905 | – | insufficient data |
Candidate 60528a9499 (gbm_degraded) fail
- ae_ratio
- 2.359
- decile_calibration_error
- 0.6431
Data table for the chart
| decile | exposure | predicted | observed |
|---|---|---|---|
| 1 | 371.09 | 0.08834 | 0.3288 |
| 2 | 371.26 | 0.09733 | 0.3394 |
| 3 | 371.43 | 0.1024 | 0.4012 |
| 4 | 371.56 | 0.1077 | 0.4629 |
| 5 | 370.67 | 0.114 | 0.4101 |
| 6 | 372 | 0.122 | 0.3387 |
| 7 | 371.53 | 0.1307 | 0.2422 |
| 8 | 370.8 | 0.1422 | 0.2184 |
| 9 | 371.44 | 0.1603 | 0.1588 |
| 10 | 371.94 | 0.2122 | 0.1129 |
Slices (actual / expected claims; 1.0 is calibrated)
| slice | level | n | A/E | 95% CI | flag |
|---|---|---|---|---|---|
| bonus_malus | 55-70 | 2884 | 2.17 | [1.96, 2.34] | |
| bonus_malus | 70-90 | 746 | 6.36 | – | insufficient data |
| bonus_malus | 90-110 | 104 | 12.9 | – | insufficient data |
| bonus_malus | <55 | 2241 | 0.996 | [0.869, 1.17] | |
| bonus_malus | >=110 | 25 | 19.7 | – | insufficient data |
| driver_age | 25-35 | 1011 | 3.21 | – | insufficient data |
| driver_age | 35-45 | 1481 | 1.19 | – | insufficient data |
| driver_age | 45-55 | 1477 | 1.31 | – | insufficient data |
| driver_age | 55-65 | 969 | 0.919 | – | insufficient data |
| driver_age | 65-75 | 420 | 0.464 | – | insufficient data |
| driver_age | <25 | 513 | 14.9 | – | insufficient data |
| driver_age | >=75 | 129 | 0.249 | – | insufficient data |
| exposure | 0.25-0.75 | 3200 | 2.47 | [2.21, 2.71] | |
| exposure | 0.75-1 | 2083 | 2.24 | [2.09, 2.41] | |
| exposure | <0.25 | 715 | 2.59 | – | insufficient data |
| exposure | >=1 | 2 | 3.93 | – | insufficient data |
| region | R21 | 299 | 2.18 | – | insufficient data |
| region | R31 | 257 | 2.8 | – | insufficient data |
| region | R42 | 281 | 2.34 | – | insufficient data |
| region | R52 | 293 | 2.21 | – | insufficient data |
| region | R72 | 293 | 1.85 | – | insufficient data |
| region | R73 | 292 | 3.06 | – | insufficient data |
| region | R94 | 259 | 3.2 | – | insufficient data |
| region | other | 4026 | 2.28 | [2.12, 2.48] | |
| vehicle_age | 1-3 | 1094 | 2.13 | – | insufficient data |
| vehicle_age | 10-15 | 933 | 2.49 | – | insufficient data |
| vehicle_age | 3-6 | 1781 | 2.41 | – | insufficient data |
| vehicle_age | 6-10 | 1590 | 2.49 | – | insufficient data |
| vehicle_age | <1 | 202 | 2.65 | – | insufficient data |
| vehicle_age | >=15 | 400 | 1.84 | – | insufficient data |
Candidate c2c40600b7 (glm) pass
- ae_ratio
- 0.985
- decile_calibration_error
- 0.05228
Data table for the chart
| decile | exposure | predicted | observed |
|---|---|---|---|
| 1 | 370.65 | 0.06529 | 0.08094 |
| 2 | 371.68 | 0.08852 | 0.06995 |
| 3 | 371.29 | 0.106 | 0.08619 |
| 4 | 371.79 | 0.1228 | 0.1399 |
| 5 | 370.88 | 0.1406 | 0.1429 |
| 6 | 371.7 | 0.1637 | 0.1345 |
| 7 | 371.45 | 0.1968 | 0.1938 |
| 8 | 371.44 | 0.2529 | 0.2665 |
| 9 | 371.46 | 0.4027 | 0.3715 |
| 10 | 371.38 | 1.52 | 1.527 |
Slices (actual / expected claims; 1.0 is calibrated)
| slice | level | n | A/E | 95% CI | flag |
|---|---|---|---|---|---|
| bonus_malus | 55-70 | 2884 | 0.961 | [0.882, 1.03] | |
| bonus_malus | 70-90 | 746 | 1.07 | – | insufficient data |
| bonus_malus | 90-110 | 104 | 1.14 | – | insufficient data |
| bonus_malus | <55 | 2241 | 0.959 | [0.841, 1.1] | |
| bonus_malus | >=110 | 25 | 0.595 | – | insufficient data |
| driver_age | 25-35 | 1011 | 1 | – | insufficient data |
| driver_age | 35-45 | 1481 | 0.881 | – | insufficient data |
| driver_age | 45-55 | 1477 | 0.982 | – | insufficient data |
| driver_age | 55-65 | 969 | 1.21 | – | insufficient data |
| driver_age | 65-75 | 420 | 0.74 | – | insufficient data |
| driver_age | <25 | 513 | 1 | – | insufficient data |
| driver_age | >=75 | 129 | 0.54 | – | insufficient data |
| exposure | 0.25-0.75 | 3200 | 1.01 | [0.944, 1.09] | |
| exposure | 0.75-1 | 2083 | 0.956 | [0.874, 1.03] | |
| exposure | <0.25 | 715 | 1.02 | – | insufficient data |
| exposure | >=1 | 2 | 3.27 | – | insufficient data |
| region | R21 | 299 | 0.887 | – | insufficient data |
| region | R31 | 257 | 1.19 | – | insufficient data |
| region | R42 | 281 | 0.891 | – | insufficient data |
| region | R52 | 293 | 0.811 | – | insufficient data |
| region | R72 | 293 | 0.91 | – | insufficient data |
| region | R73 | 292 | 1.3 | – | insufficient data |
| region | R94 | 259 | 1.22 | – | insufficient data |
| region | other | 4026 | 0.966 | [0.909, 1.01] | |
| vehicle_age | 1-3 | 1094 | 0.871 | – | insufficient data |
| vehicle_age | 10-15 | 933 | 1.08 | – | insufficient data |
| vehicle_age | 3-6 | 1781 | 1.05 | – | insufficient data |
| vehicle_age | 6-10 | 1590 | 0.947 | – | insufficient data |
| vehicle_age | <1 | 202 | 1.13 | – | insufficient data |
| vehicle_age | >=15 | 400 | 0.92 | – | insufficient data |
Shadow and canary
All traffic below is simulated (held-out policies replayed in random order).
Candidate 5fc4383bbfa34776
Shadow pass
- requests
- 3600 (minimum 3000)
- agreement within tolerance
- 0.694
- rank correlation (Spearman)
- 0.916
- mean latency delta (candidate − champion)
- -0.0189 ms
- exception rate
- 0
- poisson_deviance vs champion (labelled rows)
- candidate 0.8339 vs champion 0.8434 (n=3600)
- evidence
- shadow_summary.json
| check | status | value | threshold |
|---|---|---|---|
| spearman | pass | 0.9158 | >= 0.85 |
| exception_rate | pass | 0 | <= 0.01 |
| latency_delta_ms | pass | -0.01893 | <= 30 |
| poisson_deviance_vs_champion | pass | -0.01126 | relative worsening <= 0.05 |
Canary pass
- traffic
- 10% canary: 372 canary / 3228 control requests
- error rate
- 0
- p95 latency
- 0.0365 ms
- prediction PSI vs shadow baseline
- 0.0259
- poisson_deviance vs control
- candidate 0.7661 vs champion 0.7797
- evidence
- canary_summary.json
| guardrail | status | value | threshold |
|---|---|---|---|
| error_rate | pass | 0 | <= 0.02 |
| p95_latency_ms | pass | 0.03651 | <= 50 |
| prediction_psi_vs_shadow | pass | 0.02587 | <= 0.2 |
| poisson_deviance_vs_control | pass | -0.01753 | relative worsening <= 0.1 |
Drift
Drift below is injected on purpose into simulated traffic (synthetic: true).
Run incident
| window | level | rows | start (UTC) | scenario | prediction PSI | validation failure rate | label drift (rel. change) | evidence |
|---|---|---|---|---|---|---|---|---|
| 6 | ok | 1200 | 2026-10-08T01:07:06Z | none (synthetic) | 0.0134 | 0 | 0.044 (n.s.) | json · Evidently report |
| 7 | ok | 1200 | 2026-10-08T01:07:06Z | none (synthetic) | 0.00859 | 0 | 0.004 (n.s.) | json · Evidently report |
| 8 | ok | 1200 | 2026-10-08T01:07:06Z | none (synthetic) | 0.0053 | 0 | 0.102 (n.s.) | json · Evidently report |
| 9 | watch | 1200 | 2026-10-08T01:07:06Z | young_driver_surge (synthetic) | 0.189 | 0 | 0.065 (n.s.) | json · Evidently report |
| 10 | alert | 1200 | 2026-10-08T01:07:06Z | young_driver_surge (synthetic) | 0.766 | 0 | 0.071 (n.s.) | json · Evidently report |
| 11 | alert | 1200 | 2026-10-08T01:07:07Z | young_driver_surge (synthetic) | 1.88 | 0 | 0.039 (n.s.) | json · Evidently report |
| 12 | alert | 1200 | 2026-10-08T01:07:12Z | young_driver_surge (synthetic) | 1.93 | 0 | 0.001 (n.s.) | json · Evidently report |
| 13 | alert | 1200 | 2026-10-08T01:07:12Z | young_driver_surge (synthetic) | 1.83 | 0 | 0.087 | json · Evidently report |
Per-feature detail at the worst window (13)
| feature | level | PSI | KS statistic | new-category rate | out-of-range rate |
|---|---|---|---|---|---|
| Area | ok | 0.00715 | – | 0 | – |
| BonusMalus | alert | 0.498 | 0.266 | – | 0.00833 |
| Density | ok | 0.0314 | 0.0433 | – | 0 |
| DrivAge | alert | 2.09 | 0.652 | – | 0 |
| Region | ok | 0.0463 | – | 0 | – |
| VehAge | ok | 0.021 | 0.0228 | – | 0 |
| VehBrand | ok | 0.0202 | – | 0 | – |
| VehGas | ok | 9e-06 | – | 0 | – |
| VehPower | ok | 0.0228 | 0.0577 | – | 0 |
Scenario matrix: declared vs observed (synthetic)
| scenario | what changes | declared expectation | window levels | result |
|---|---|---|---|---|
| covariate_shift | DrivAge shifts upward by ~2 standard deviations | alert on feature:DrivAge | ok → ok → alert → alert → alert → alert | detected |
| new_category | unseen levels appear in Area | alert on feature:Area | ok → ok → alert → alert → alert → alert | detected |
| schema_break | DrivAge is sent as a string | alert on validation | ok → ok → alert → alert → alert → alert | detected |
| label_shift | delayed labels show a higher base rate | alert on label | ok → ok → alert → alert → alert → alert | detected |
| silent_decay | Vehicle age drifts up slowly (older vehicles gradually over-represented) | alert on feature:VehAge, watch before alert | ok → ok → ok → ok → ok → ok → watch → watch → alert → alert | detected |
| young_driver_surge | A younger cohort enters the book: young drivers are over-represented in traffic | alert on feature:DrivAge | ok → ok → ok → alert → alert → alert → alert | detected |
| new_vehicle_brand | Unseen VehBrand values appear | alert on feature:VehBrand | ok → ok → alert → alert → alert → alert | detected |
| bonus_malus_schema_break | BonusMalus arrives as a string | alert on validation | ok → ok → alert → alert → alert → alert | detected |
| claims_frequency_shift | Delayed labels show a higher base claim rate | alert on label | ok → ok → alert → watch → alert → alert | detected |
Every declared expectation was met in this run. See docs/MONITORING.md for sensitivity limits.
Incident story
Synthetic drift on simulated traffic, driven by the audit log and drift reports.
- Healthy champion
Window 6, champion5fc4383bbf, level ok. A/E 1.04, decile calibration error 0.115, deviance 0.8893 (n=1200) - Drift begins (synthetic)
Scenarioyoung_driver_surgestarts at window 9 with intensity 0.09; level watch. - Monitor reaches watch
Window 9 (intensity 0.09): watch. A/E 0.935, decile calibration error 0.0841, deviance 0.9643 (n=1200) - Monitor reaches alert
Window 10 (intensity 0.3, 2026-10-08T01:07:06Z): alert. A/E 0.929, decile calibration error 0.162, deviance 1.175 (n=1200) - Rollback recommended
Audit #4 at 2026-10-08T01:07:11Z: Drift alert in window 10 (drivers: DrivAge, prediction). Rollback to previous champion c2c40600b7290013 recommended. - Rollback executed
Audit #5 at 2026-10-08T01:07:12Z:5fc4383bbf→c2c40600b7. Drift alert from synthetic scenario 'young_driver_surge'; restoring previous champion per runbook. - Restored state
Window 12: servingc2c40600b7, level alert. A/E 0.999, decile calibration error 0.141, deviance 1.416 (n=1200). The input drift is still present after rollback: rolling back restores the previous known-good model, it does not fix the data.
Data table for the chart
| window | phase | level | serving champion | A/E | decile error | rows |
|---|---|---|---|---|---|---|
| 6 | before | ok | 5fc4383bbf | 1.04 | 0.115 | 1200 |
| 7 | before | ok | 5fc4383bbf | 1 | 0.0623 | 1200 |
| 8 | during | ok | 5fc4383bbf | 0.898 | 0.159 | 1200 |
| 9 | during | watch | 5fc4383bbf | 0.935 | 0.0841 | 1200 |
| 10 | during | alert | 5fc4383bbf | 0.929 | 0.162 | 1200 |
| 11 | during | alert | 5fc4383bbf | 0.961 | 0.141 | 1200 |
| 12 | after | alert | c2c40600b7 | 0.999 | 0.141 | 1200 |
| 13 | after | alert | c2c40600b7 | 0.913 | 0.137 | 1200 |
Calibration: Before (healthy), window 6
Calibration: At first alert, window 10
Calibration: After rollback, window 12
Drift is injected into simulated traffic; nothing here is real-world behavior.
Model card
Card for candidate c2c40600b7290013 (the model most recently in service among those with a card). Generated, not hand-written.
Model card: claim frequency glm / candidate c2c40600b7290013
Synthetic data. This candidate was trained on a generated stand-in dataset, not real insurance data. Numbers below describe the stand-in only.
Intended use
Educational and portfolio demonstration of a gated model release process for a motor third-party-liability claim-frequency model (expected claims per exposure-year). It exists to exercise gates, shadow/canary, drift monitoring and rollback.
Out-of-scope use
Not for pricing, underwriting, reserving, or any decision about real people. No claim of regulatory compliance, actuarial sign-off or fairness certification. Severity (claim cost) is not modeled, so nothing here estimates expected loss.
Data
- Source: Generated by
examples/claims_frequency/data.py(labeled synthetic). - Data version:
synthetic-7c8426d485-n40000 - Rows per split: {'train': 24001, 'val': 4000, 'holdout': 6000, 'stream': 5999}
- Split: policy- and risk-profile-grouped, claim-stratified; NO temporal split (dataset has no dates)
- Known gaps: the dataset has no timestamps. A time-based split is impossible, so temporal generalization was not evaluated. Simulated production traffic replays held-out policies in random order and is not real production behavior.
- One row per policy; severity data not used; exposure is capped at 1 year, claim counts at 4 (row counts per cleaning rule are in the data-quality report).
Training procedure
- Family:
glm; seed20260101; params:{"alpha": 0.0001, "recalibrate": true} - Target: claim rate (claims per exposure-year), Poisson loss, exposure as sample weight (equivalent to an offset). Tuning used grouped cross-validation inside the training split only; the holdout is evaluated once per candidate.
- Optional scale-factor recalibration is fitted on the validation split and is part of the candidate (and its hash).
Metrics
Evaluated on the holdout, exposure-weighted, 95% bootstrap CIs (policies resampled).
| metric | value | 95% CI |
|---|---|---|
| poisson_deviance | 0.80400 | [0.76413, 0.84317] |
| deviance_skill | 0.32327 | [0.29062, 0.34725] |
| gini | 0.55070 | [0.52279, 0.58302] |
Slices and calibration
Calibration (exposure-weighted decile error, relative to portfolio frequency): 0.05228; overall actual/expected: 0.98497.
| slice | level | n | A/E | 95% CI | flag |
|---|---|---|---|---|---|
| driver_age | 25-35 | 1011 | 1.004 | insufficient data | |
| driver_age | 35-45 | 1481 | 0.881 | insufficient data | |
| driver_age | 45-55 | 1477 | 0.982 | insufficient data | |
| driver_age | 55-65 | 969 | 1.208 | insufficient data | |
| driver_age | 65-75 | 420 | 0.740 | insufficient data | |
| driver_age | <25 | 513 | 1.001 | insufficient data | |
| driver_age | >=75 | 129 | 0.540 | insufficient data | |
| vehicle_age | 1-3 | 1094 | 0.871 | insufficient data | |
| vehicle_age | 10-15 | 933 | 1.080 | insufficient data | |
| vehicle_age | 3-6 | 1781 | 1.051 | insufficient data | |
| vehicle_age | 6-10 | 1590 | 0.947 | insufficient data | |
| vehicle_age | <1 | 202 | 1.130 | insufficient data | |
| vehicle_age | >=15 | 400 | 0.920 | insufficient data | |
| region | R21 | 299 | 0.887 | insufficient data | |
| region | R31 | 257 | 1.190 | insufficient data | |
| region | R42 | 281 | 0.891 | insufficient data | |
| region | R52 | 293 | 0.811 | insufficient data | |
| region | R72 | 293 | 0.910 | insufficient data | |
| region | R73 | 292 | 1.301 | insufficient data | |
| region | R94 | 259 | 1.223 | insufficient data | |
| region | other | 4026 | 0.966 | [0.909, 1.009] | |
| bonus_malus | 55-70 | 2884 | 0.961 | [0.882, 1.035] | |
| bonus_malus | 70-90 | 746 | 1.066 | insufficient data | |
| bonus_malus | 90-110 | 104 | 1.144 | insufficient data | |
| bonus_malus | <55 | 2241 | 0.959 | [0.841, 1.102] | |
| bonus_malus | >=110 | 25 | 0.595 | insufficient data | |
| exposure | 0.25-0.75 | 3200 | 1.010 | [0.944, 1.090] | |
| exposure | 0.75-1 | 2083 | 0.956 | [0.874, 1.027] | |
| exposure | <0.25 | 715 | 1.024 | insufficient data | |
| exposure | >=1 | 2 | 3.266 | insufficient data |
Explainability
Describes the fitted model, not causal effects. Computed on a validation sample, not the holdout.
Permutation importance (increase in exposure-weighted deviance when a feature is shuffled):
| feature | deviance increase |
|---|---|
| DrivAge | 0.424056 |
| BonusMalus | 0.085794 |
| Density | 0.008884 |
| VehAge | 0.008623 |
| VehGas | 0.007170 |
| VehPower | 0.004453 |
| Exposure | 0.004240 |
| Area | 0.002342 |
| VehBrand | 0.002023 |
| Region | -0.000998 |
Largest GLM coefficients (rate ratio = exp(coef)):
| term | coef | rate ratio |
|---|---|---|
| bins__DrivAge_2 | 1.7690 | 5.865 |
| bins__DrivAge_1 | 1.7130 | 5.546 |
| bins__BonusMalus_8 | 1.4861 | 4.420 |
| bins__BonusMalus_1 | -1.0805 | 0.339 |
| bins__BonusMalus_7 | 0.9058 | 2.474 |
| bins__BonusMalus_2 | -0.8206 | 0.440 |
| bins__DrivAge_10 | -0.6762 | 0.509 |
| bins__DrivAge_11 | -0.6175 | 0.539 |
| bins__DrivAge_9 | -0.6105 | 0.543 |
| bins__BonusMalus_3 | -0.5626 | 0.570 |
| bins__DrivAge_13 | -0.5196 | 0.595 |
| bins__DrivAge_12 | -0.4742 | 0.622 |
| bins__DrivAge_14 | -0.4521 | 0.636 |
| bins__DrivAge_3 | 0.4474 | 1.564 |
| bins__DrivAge_4 | 0.4170 | 1.517 |
Partial dependence, DrivAge: 18.0: 1.1604, 29.0: 0.3273, 36.0: 0.1676, 42.0: 0.1562, 47.0: 0.1709, 53.0: 0.1662, 60.0: 0.1064, 75.0: 0.1244
Partial dependence, BonusMalus: 50.0: 0.1718, 51.0: 0.2228, 53.0: 0.2228, 56.0: 0.2228, 58.0: 0.2228, 62.0: 0.2883, 68.0: 0.2883, 89.0: 0.5429
Partial dependence, VehAge: 0.0: 0.3219, 2.0: 0.3605, 3.0: 0.3080, 5.0: 0.2779, 6.0: 0.2779, 8.0: 0.2717, 11.0: 0.2273, 20.0: 0.2331
Partial dependence, Density: 5.0: 0.1943, 39.27: 0.2374, 91.0: 0.2589, 170.0: 0.2764, 331.0: 0.2963, 672.44: 0.3192, 1539.73: 0.3482, 9043.6: 0.4195
Limitations
- No timestamps: temporal generalization, seasonality and trend were not evaluated.
- Simulated traffic and drift: shadow/canary/drift results come from replaying held-out policies and from synthetic perturbations; they are demonstrations of the machinery, not evidence about real production.
- Frequency only; no severity. One country/product/period; extrapolation elsewhere is untested.
- Small slices are flagged
insufficient dataand not judged. Bootstrap CIs resample policies and ignore policy-to-policy dependence. - The audit log is a local file: tamper-evident, not tamper-proof.
Sensitive features and proxy risk
DrivAge and Region (and, indirectly, Density/Area) can act as proxies for protected characteristics (age, and in some jurisdictions ethnicity, religion or socio-economic status). They are kept for realism and reported by slice; this card does not claim the model is fair, compliant, or fit for pricing or underwriting.
Provenance
| field | value |
|---|---|
| candidate hash | c2c40600b7290013 |
| git SHA | bcbdf51a7d5d |
| data version | synthetic-7c8426d485-n40000 |
| seed | 20260101 |
| lockfile hash | 325e385df3aac6ed |
| gate config hash | 7a4a29aea97dd405 |
Card for candidate 5fc4383bbf
Model card: claim frequency gbm / candidate 5fc4383bbfa34776
Synthetic data. This candidate was trained on a generated stand-in dataset, not real insurance data. Numbers below describe the stand-in only.
Intended use
Educational and portfolio demonstration of a gated model release process for a motor third-party-liability claim-frequency model (expected claims per exposure-year). It exists to exercise gates, shadow/canary, drift monitoring and rollback.
Out-of-scope use
Not for pricing, underwriting, reserving, or any decision about real people. No claim of regulatory compliance, actuarial sign-off or fairness certification. Severity (claim cost) is not modeled, so nothing here estimates expected loss.
Data
- Source: Generated by
examples/claims_frequency/data.py(labeled synthetic). - Data version:
synthetic-7c8426d485-n40000 - Rows per split: {'train': 24001, 'val': 4000, 'holdout': 6000, 'stream': 5999}
- Split: policy- and risk-profile-grouped, claim-stratified; NO temporal split (dataset has no dates)
- Known gaps: the dataset has no timestamps. A time-based split is impossible, so temporal generalization was not evaluated. Simulated production traffic replays held-out policies in random order and is not real production behavior.
- One row per policy; severity data not used; exposure is capped at 1 year, claim counts at 4 (row counts per cleaning rule are in the data-quality report).
Training procedure
- Family:
gbm; seed20260101; params:{"n_estimators": 600, "learning_rate": 0.05, "num_leaves": 16, "min_child_samples": 400, "subsample": 0.8, "subsample_freq": 1, "colsample_bytree": 0.8, "reg_lambda": 5.0, "num_threads": 4, "recalibrate": true} - Target: claim rate (claims per exposure-year), Poisson loss, exposure as sample weight (equivalent to an offset). Tuning used grouped cross-validation inside the training split only; the holdout is evaluated once per candidate.
- Optional scale-factor recalibration is fitted on the validation split and is part of the candidate (and its hash).
Metrics
Evaluated on the holdout, exposure-weighted, 95% bootstrap CIs (policies resampled).
| metric | value | 95% CI |
|---|---|---|
| deviance_skill | 0.32472 | [0.29245, 0.34820] |
| gini | 0.55080 | [0.52287, 0.58032] |
| poisson_deviance | 0.80227 | [0.76450, 0.83489] |
Three-model comparison (same holdout)
| model | candidate | deviance | CI | skill vs constant | Gini | CI |
|---|---|---|---|---|---|---|
| Constant baseline | c7e902deae28e708 | 1.18807 | [1.11657, 1.23138] | 0.00000 | 0.00000 | [0.00000, 0.00000] |
| Poisson GLM | c2c40600b7290013 | 0.80400 | [0.76413, 0.84317] | 0.32327 | 0.55070 | [0.52279, 0.58302] |
| LightGBM | 5fc4383bbfa34776 | 0.80227 | [0.76450, 0.83489] | 0.32472 | 0.55080 | [0.52287, 0.58032] |
Lowest deviance: LightGBM. Differences smaller than the CI width are not evidence of a real difference; see the paired-bootstrap G3 check for the champion/challenger decision. Holdout is one random grouped split, not a time split.
Slices and calibration
Calibration (exposure-weighted decile error, relative to portfolio frequency): 0.08826; overall actual/expected: 0.99463.
| slice | level | n | A/E | 95% CI | flag |
|---|---|---|---|---|---|
| bonus_malus | 55-70 | 2884 | 0.940 | [0.859, 1.010] | |
| bonus_malus | 70-90 | 746 | 1.094 | insufficient data | |
| bonus_malus | 90-110 | 104 | 1.119 | insufficient data | |
| bonus_malus | <55 | 2241 | 0.958 | [0.843, 1.098] | |
| bonus_malus | >=110 | 25 | 0.830 | insufficient data | |
| driver_age | 25-35 | 1011 | 1.017 | insufficient data | |
| driver_age | 35-45 | 1481 | 0.828 | insufficient data | |
| driver_age | 45-55 | 1477 | 0.983 | insufficient data | |
| driver_age | 55-65 | 969 | 1.067 | insufficient data | |
| driver_age | 65-75 | 420 | 0.713 | insufficient data | |
| driver_age | <25 | 513 | 1.057 | insufficient data | |
| driver_age | >=75 | 129 | 0.522 | insufficient data | |
| exposure | 0.25-0.75 | 3200 | 1.027 | [0.959, 1.100] | |
| exposure | 0.75-1 | 2083 | 0.960 | [0.875, 1.032] | |
| exposure | <0.25 | 715 | 1.003 | insufficient data | |
| exposure | >=1 | 2 | 2.575 | insufficient data | |
| region | R21 | 299 | 0.912 | insufficient data | |
| region | R31 | 257 | 1.168 | insufficient data | |
| region | R42 | 281 | 0.892 | insufficient data | |
| region | R52 | 293 | 0.836 | insufficient data | |
| region | R72 | 293 | 0.945 | insufficient data | |
| region | R73 | 292 | 1.304 | insufficient data | |
| region | R94 | 259 | 1.247 | insufficient data | |
| region | other | 4026 | 0.973 | [0.916, 1.020] | |
| vehicle_age | 1-3 | 1094 | 0.970 | insufficient data | |
| vehicle_age | 10-15 | 933 | 1.050 | insufficient data | |
| vehicle_age | 3-6 | 1781 | 1.036 | insufficient data | |
| vehicle_age | 6-10 | 1590 | 0.938 | insufficient data | |
| vehicle_age | <1 | 202 | 1.128 | insufficient data | |
| vehicle_age | >=15 | 400 | 0.905 | insufficient data |
Explainability
Describes the fitted model, not causal effects. Computed on a validation sample, not the holdout.
Permutation importance (increase in exposure-weighted deviance when a feature is shuffled):
| feature | deviance increase |
|---|---|
| DrivAge | 0.409138 |
| BonusMalus | 0.081809 |
| VehGas | 0.013334 |
| Density | 0.006920 |
| VehPower | 0.004433 |
| VehAge | 0.004273 |
| VehBrand | 0.001130 |
| Area | 0.000184 |
| Exposure | 0.000113 |
| Region | -0.001787 |
TreeSHAP via LightGBM pred_contrib (log-rate scale). Mean |SHAP| by feature:
| feature | mean abs SHAP |
|---|---|
| DrivAge | 0.4602 |
| BonusMalus | 0.1996 |
| Density | 0.1336 |
| VehAge | 0.1002 |
| VehPower | 0.0645 |
| VehGas | 0.0496 |
| Region | 0.0484 |
| VehBrand | 0.0295 |
| Area | 0.0131 |
| Exposure | 0.0109 |
Partial dependence, DrivAge: 18.0: 1.1539, 29.0: 0.3415, 36.0: 0.1755, 42.0: 0.1715, 47.0: 0.1719, 53.0: 0.1696, 60.0: 0.1245, 75.0: 0.1275
Partial dependence, BonusMalus: 50.0: 0.2017, 51.0: 0.2017, 53.0: 0.2198, 56.0: 0.2389, 58.0: 0.2435, 62.0: 0.2706, 68.0: 0.3267, 89.0: 0.5546
Partial dependence, VehAge: 0.0: 0.3342, 2.0: 0.3332, 3.0: 0.3258, 5.0: 0.2925, 6.0: 0.2885, 8.0: 0.2785, 11.0: 0.2561, 20.0: 0.2509
Partial dependence, Density: 5.0: 0.2205, 39.27: 0.2462, 91.0: 0.2488, 170.0: 0.2833, 331.0: 0.2999, 672.44: 0.3231, 1539.73: 0.3477, 9043.6: 0.3729
Limitations
- No timestamps: temporal generalization, seasonality and trend were not evaluated.
- Simulated traffic and drift: shadow/canary/drift results come from replaying held-out policies and from synthetic perturbations; they are demonstrations of the machinery, not evidence about real production.
- Frequency only; no severity. One country/product/period; extrapolation elsewhere is untested.
- Small slices are flagged
insufficient dataand not judged. Bootstrap CIs resample policies and ignore policy-to-policy dependence. - The audit log is a local file: tamper-evident, not tamper-proof.
Sensitive features and proxy risk
DrivAge and Region (and, indirectly, Density/Area) can act as proxies for protected characteristics (age, and in some jurisdictions ethnicity, religion or socio-economic status). They are kept for realism and reported by slice; this card does not claim the model is fair, compliant, or fit for pricing or underwriting.
Provenance
| field | value |
|---|---|
| candidate hash | 5fc4383bbfa34776 |
| git SHA | bcbdf51a7d5d |
| data version | synthetic-7c8426d485-n40000 |
| seed | 20260101 |
| lockfile hash | 325e385df3aac6ed |
| gate config hash | 7a4a29aea97dd405 |
Card for candidate 60528a9499
Model card: claim frequency gbm_degraded / candidate 60528a9499f3939d
Synthetic data. This candidate was trained on a generated stand-in dataset, not real insurance data. Numbers below describe the stand-in only.
Intended use
Educational and portfolio demonstration of a gated model release process for a motor third-party-liability claim-frequency model (expected claims per exposure-year). It exists to exercise gates, shadow/canary, drift monitoring and rollback.
Out-of-scope use
Not for pricing, underwriting, reserving, or any decision about real people. No claim of regulatory compliance, actuarial sign-off or fairness certification. Severity (claim cost) is not modeled, so nothing here estimates expected loss.
Data
- Source: Generated by
examples/claims_frequency/data.py(labeled synthetic). - Data version:
synthetic-7c8426d485-n40000 - Rows per split: {'train': 24001, 'val': 4000, 'holdout': 6000, 'stream': 5999}
- Split: policy- and risk-profile-grouped, claim-stratified; NO temporal split (dataset has no dates)
- Known gaps: the dataset has no timestamps. A time-based split is impossible, so temporal generalization was not evaluated. Simulated production traffic replays held-out policies in random order and is not real production behavior.
- One row per policy; severity data not used; exposure is capped at 1 year, claim counts at 4 (row counts per cleaning rule are in the data-quality report).
Training procedure
- Family:
gbm_degraded; seed20260101; params:{"n_estimators": 100, "learning_rate": 0.05, "num_leaves": 16, "min_child_samples": 400, "num_threads": 4, "recalibrate": false, "corrupt_preprocessing": true} - Target: claim rate (claims per exposure-year), Poisson loss, exposure as sample weight (equivalent to an offset). Tuning used grouped cross-validation inside the training split only; the holdout is evaluated once per candidate.
- Optional scale-factor recalibration is fitted on the validation split and is part of the candidate (and its hash).
Metrics
Evaluated on the holdout, exposure-weighted, 95% bootstrap CIs (policies resampled).
| metric | value | 95% CI |
|---|---|---|
| poisson_deviance | 1.42458 | [1.30299, 1.49901] |
| deviance_skill | -0.19907 | [-0.21953, -0.16500] |
| gini | -0.16149 | [-0.19006, -0.12056] |
Slices and calibration
Calibration (exposure-weighted decile error, relative to portfolio frequency): 0.64306; overall actual/expected: 2.35865.
| slice | level | n | A/E | 95% CI | flag |
|---|---|---|---|---|---|
| driver_age | 25-35 | 1011 | 3.212 | insufficient data | |
| driver_age | 35-45 | 1481 | 1.194 | insufficient data | |
| driver_age | 45-55 | 1477 | 1.307 | insufficient data | |
| driver_age | 55-65 | 969 | 0.919 | insufficient data | |
| driver_age | 65-75 | 420 | 0.464 | insufficient data | |
| driver_age | <25 | 513 | 14.948 | insufficient data | |
| driver_age | >=75 | 129 | 0.249 | insufficient data | |
| vehicle_age | 1-3 | 1094 | 2.134 | insufficient data | |
| vehicle_age | 10-15 | 933 | 2.489 | insufficient data | |
| vehicle_age | 3-6 | 1781 | 2.409 | insufficient data | |
| vehicle_age | 6-10 | 1590 | 2.492 | insufficient data | |
| vehicle_age | <1 | 202 | 2.653 | insufficient data | |
| vehicle_age | >=15 | 400 | 1.838 | insufficient data | |
| region | R21 | 299 | 2.179 | insufficient data | |
| region | R31 | 257 | 2.797 | insufficient data | |
| region | R42 | 281 | 2.340 | insufficient data | |
| region | R52 | 293 | 2.209 | insufficient data | |
| region | R72 | 293 | 1.855 | insufficient data | |
| region | R73 | 292 | 3.062 | insufficient data | |
| region | R94 | 259 | 3.204 | insufficient data | |
| region | other | 4026 | 2.285 | [2.124, 2.480] | |
| bonus_malus | 55-70 | 2884 | 2.172 | [1.957, 2.345] | |
| bonus_malus | 70-90 | 746 | 6.359 | insufficient data | |
| bonus_malus | 90-110 | 104 | 12.880 | insufficient data | |
| bonus_malus | <55 | 2241 | 0.996 | [0.869, 1.165] | |
| bonus_malus | >=110 | 25 | 19.684 | insufficient data | |
| exposure | 0.25-0.75 | 3200 | 2.467 | [2.214, 2.706] | |
| exposure | 0.75-1 | 2083 | 2.237 | [2.094, 2.415] | |
| exposure | <0.25 | 715 | 2.587 | insufficient data | |
| exposure | >=1 | 2 | 3.928 | insufficient data |
Explainability
Describes the fitted model, not causal effects. Computed on a validation sample, not the holdout.
Permutation importance (increase in exposure-weighted deviance when a feature is shuffled):
| feature | deviance increase |
|---|---|
| Density | 0.003261 |
| VehBrand | 0.001447 |
| VehAge | 0.001337 |
| Region | 0.000633 |
| Area | 0.000491 |
| Exposure | 0.000043 |
| VehGas | 0.000000 |
| VehPower | -0.000070 |
| BonusMalus | -0.024224 |
| DrivAge | -0.031138 |
TreeSHAP via LightGBM pred_contrib (log-rate scale). Mean |SHAP| by feature:
| feature | mean abs SHAP |
|---|---|
| DrivAge | 0.3813 |
| BonusMalus | 0.1830 |
| Density | 0.0948 |
| VehAge | 0.0763 |
| VehPower | 0.0292 |
| Region | 0.0291 |
| VehGas | 0.0209 |
| VehBrand | 0.0105 |
| Area | 0.0038 |
| Exposure | 0.0029 |
Partial dependence, DrivAge: 18.0: 0.1156, 29.0: 0.1156, 36.0: 0.1156, 42.0: 0.1156, 47.0: 0.1156, 53.0: 0.1228, 60.0: 0.1313, 75.0: 0.1852
Partial dependence, BonusMalus: 50.0: 0.1457, 51.0: 0.1457, 53.0: 0.1456, 56.0: 0.1190, 58.0: 0.1175, 62.0: 0.1172, 68.0: 0.1172, 89.0: 0.1172
Partial dependence, VehAge: 0.0: 0.1461, 2.0: 0.1461, 3.0: 0.1456, 5.0: 0.1215, 6.0: 0.1209, 8.0: 0.1190, 11.0: 0.1159, 20.0: 0.1159
Partial dependence, Density: 5.0: 0.1057, 39.27: 0.1131, 91.0: 0.1131, 170.0: 0.1312, 331.0: 0.1333, 672.44: 0.1386, 1539.73: 0.1407, 9043.6: 0.1407
Limitations
- No timestamps: temporal generalization, seasonality and trend were not evaluated.
- Simulated traffic and drift: shadow/canary/drift results come from replaying held-out policies and from synthetic perturbations; they are demonstrations of the machinery, not evidence about real production.
- Frequency only; no severity. One country/product/period; extrapolation elsewhere is untested.
- Small slices are flagged
insufficient dataand not judged. Bootstrap CIs resample policies and ignore policy-to-policy dependence. - The audit log is a local file: tamper-evident, not tamper-proof.
Sensitive features and proxy risk
DrivAge and Region (and, indirectly, Density/Area) can act as proxies for protected characteristics (age, and in some jurisdictions ethnicity, religion or socio-economic status). They are kept for realism and reported by slice; this card does not claim the model is fair, compliant, or fit for pricing or underwriting.
Provenance
| field | value |
|---|---|
| candidate hash | 60528a9499f3939d |
| git SHA | bcbdf51a7d5d |
| data version | synthetic-7c8426d485-n40000 |
| seed | 20260101 |
| lockfile hash | 325e385df3aac6ed |
| gate config hash | 7a4a29aea97dd405 |
About and limitations
- No timestamps in the dataset. The public frequency dataset has no date column, so a time-based split is impossible. Temporal generalization was not evaluated.
- Traffic is simulated. "Production" traffic replays held-out policies in random order. Shadow and canary results demonstrate the machinery, not real production behavior.
- Drift is synthetic. Every scenario is injected on purpose and tagged
synthetic: true. - The audit log is a local file. It is tamper-evident, not tamper-proof: anyone who can rewrite the whole file can recompute the chain. Anchor the head hash somewhere you trust if that matters.
- This page is read-only. It is generated from JSON reports and the audit log, never writes to them, makes no network requests and loads no external resources.
- No compliance or fairness claim. Driver age and region can proxy for protected characteristics; no actuarial sign-off, pricing or underwriting advice is implied.
- Verification in the browser. The audit chain is verified when this page is built. The button below recomputes it in your browser with Web Crypto; if your browser does not provide Web Crypto for
file://pages, the button is hidden and the build-time result is the only one shown.