Statistical broadsheet Indian Premier League
1,243 matches
37 grounds
577 bowlers
Seasons 2008–2026

The IPL in 284,465 deliveries

I went through every legal delivery bowled in this competition since 2008, first to work out which players were genuinely better than the situations they were handed, and then to test the things everybody repeats about how a T20 innings actually works.

The players

Two batters can average the same and be playing different sports

Strike rates mostly describe the job a player was given rather than how well he did it, so every delivery here is scored against what the whole league managed in the same season, the same phase and against the same type of bowling. The column that comes out is runs above par per hundred balls, and the largest effects it finds are not about quality at all.

100 100 120 120 140 140 160 160 180 180 EQUAL AGAINST BOTH MUCH BETTER AGAINST PACE MUCH BETTER AGAINST SPIN de VILLIERS BUTTLER MS DHONI du PLESSIS S MARSH MAXWELL POORAN VIRAT KOHLI KLAASEN SURYAKUMAR GAYLE McCULLUM Y PATHAN ROHIT SHARMA STRIKE RATE AGAINST PACE STRIKE RATE AGAINST SPIN
Batters with at least 700 balls faced against pace and 500 against spin. Distance from the dashed line is the size of the imbalance.
Best against pace relative to spin
Batterv pacev spinSwing
Brendon McCullum1,341 / 831 balls142115+25
MS Dhoni2,421 / 1,514 balls154110+25
Faf du Plessis2,093 / 1,413 balls146120+21
AB de Villiers2,003 / 1,389 balls165131+21
Rohit Sharma3,363 / 2,130 balls143116+18
Parthiv Patel1,596 / 754 balls127108+17
Quinton de Kock1,649 / 894 balls142122+16
Suryakumar Yadav1,671 / 1,400 balls161134+14
Best against spin relative to pace
Batterv pacev spinSwing
Heinrich Klaasen707 / 546 balls161175−31
Shaun Marsh1,246 / 616 balls125148−26
Dwayne Smith1,114 / 641 balls127150−25
Nicholas Pooran936 / 599 balls163166−24
Glenn Maxwell942 / 867 balls149162−24
Yusuf Pathan1,505 / 724 balls139152−23
Sai Sudharsan1,124 / 550 balls146156−22
Nitish Rana1,336 / 892 balls135142−18

Heinrich Klaasen scores at 175 against spin and 161 against pace, while Brendon McCullum ran the other way at 142 and 115, and both gaps are larger than the distance between the best and worst batters in the competition overall.

The big names

The five best-known careers in the competition, adjusted for the overs they were given

100 125 150 175 200 ROHIT SHARMA 1-6 126 7-16 122 17-20 181 VIRAT KOHLI 1-6 130 7-16 126 17-20 189 MS DHONI 1-6 72 BALLS 7-16 101 17-20 175 THE PLAYER LEAGUE PAR
Strike rate by phase against the league par for that same phase and era. All three sit at or below par until the last four overs.

Virat Kohli

6,907 balls · 9,316 runs

+3.6
runs above par / 100
26%
harder to dismiss
.326
dot rate, par .382

The scoring number is good rather than remarkable, and the wicket number is where his case sits, because he is dismissed a quarter less often per ball than his situations predict while facing more deliveries than anyone above him on that list.

He gets there by wasting fewer balls rather than hitting more of them, since his boundary rate is actually below par, and his four most recent seasons read 139.8, 154.7, 144.7 and 166.0, which are the best of a nineteen-year career arriving at the end of it.

Rohit Sharma

5,493 balls · 7,284 runs

+2.4
runs above par / 100
143 / 116
strike rate, pace / spin
180.5
strike rate at the death

Slightly above par overall, and all of it earned at the death, where he strikes at 180.5 against a par of 152.9 while his powerplay rate of 126.2 sits below par for an opener.

The split that would change a team plan is the other one: he is nearly ten runs per hundred above par against pace and nearly nine below it against spin, an eighteen-run swing that puts him among the most pace-dependent batters in the league.

MS Dhoni

3,935 balls · 5,409 runs

41%
harder to dismiss
154 / 110
strike rate, pace / spin
−1.5
runs above par / 100

The most double-edged record in the dataset, because he is the hardest man here to get out with any real volume behind him, and once you adjust for the fact that he almost never batted before the seventh over he is a fractionally below-par scorer.

His middle-overs strike rate of 100.5 against a par of 121.5 gets lost inside the finisher reputation, and the spin problem is the second largest in the competition, since he survives it comfortably at a dismissal every 37.9 balls and simply cannot score off it.

6 7 8 9 10 JASPRIT BUMRAH 6.5 −1.4 1-6 6.3 −1.7 7-16 7.9 −1.9 17-20 R ASHWIN 6.9 −0.4 1-6 6.8 −0.5 7-16 7.8 −0.9 17-20 RUNS PER OVER
Runs conceded per over by phase, against the league par for that phase and era. The red figure is runs saved.

Jasprit Bumrah

3,631 balls · 187 wickets · pace

28.4
runs saved / 100 balls
7.00
economy, par 8.70
.146
boundary rate, par .205

No batter in this newspaper beats their par by as much as Bumrah beats his, and the record is built on preventing runs rather than taking wickets, since his strike rate sits a shade below par for the overs he bowls.

He is under par in all three phases, bowls a dot with 52.1 per cent of his powerplay deliveries, and concedes 7.86 an over at the death where the league concedes 9.78, which is nearly two full runs an over better than the phase he specialises in.

Ravichandran Ashwin

4,710 balls · 183 wickets · off spin

9.3
runs saved / 100 balls
6.95
economy, par 7.50
.118
boundary rate, par .145

A containment record almost entirely, achieved by not being hit to the fence: his boundary rate is the second lowest of any bowler here with three thousand deliveries or more, behind only Sunil Narine.

He takes wickets at about eight per cent below par, which is what happens when a bowler works a phase where wickets are cheap and declines to chase them, and he owns the most lopsided rivalry in the dataset, having dismissed Kohli once in 149 deliveries.

Above and below par

Who actually beat the situation they were handed

Batting · runs above par per 100 balls
BatterBallsSRPar+/−
Virender Sehwag1,740155.9117.0+38.9
Andre Russell1,505174.6144.6+30.0
Chris Gayle3,318148.4122.9+25.5
Glenn Maxwell1,809155.4131.8+23.6
AB de Villiers3,392151.2129.2+22.0
Nicholas Pooran1,535164.3146.3+18.0
Dwayne Smith1,755135.6120.7+14.8
David Warner4,675139.8125.8+14.0
Shane Watson2,797137.9124.6+13.4
Suresh Raina4,022136.9124.7+12.1
Bowling · runs saved per 100 balls
BowlerBallsEconParSaved
Jasprit Bumrahpace3,6316.998.70+28.4
Lasith Malingapace2,8276.748.02+21.4
Sunil Narinespin4,6516.667.87+20.2
Jofra Archerpace1,5787.809.01+20.1
Dale Steynpace2,1766.567.72+19.3
Bhuvneshwar Kumarpace4,6007.428.49+17.8
Mustafizur Rahmanpace1,3647.838.64+13.5
Muttiah Muralitharanspin1,5246.427.22+13.3
Rashid Khanspin3,5307.177.94+12.9
Krunal Pandyaspin2,7017.337.90+9.4

Virender Sehwag beat his situations by 39 runs per hundred balls in an era when the league scored 1.2 a ball, and paid for it by getting out about a third more often than par. At the other end Jasprit Bumrah is clear of every bowler by seven full runs per hundred. Both tables ask for 1,500 deliveries; drop that to 400 and Vaibhav Suryavanshi arrives at the top of the batting one on 447 balls, 75 runs per hundred above par, which is nearly double Sehwag.

All or nothing

The most destructive batter in the competition is also one of the most wasteful

A strike rate hides the shape of an innings completely, so the chart below splits it in two: how often a batter finds the boundary, and how often he fails to score at all, both measured against what the situation was worth. Almost everybody sits on a diagonal where more boundaries means fewer dots, and Andre Russell does not.

He hits 8.8 percentage points more boundaries than par and also leaves 7.6 points more balls unscored, which is the most extreme position anyone occupies on either axis. Yuvraj Singh and Kieron Pollard sit in the same corner with less of the upside. The opposite corner belongs to batters nobody describes as explosive, where Shubman Gill and Virat Kohli cut six or seven points off the par dot rate while hitting an entirely ordinary number of boundaries.

-8 -4 +0 +4 +8 -4 -2 +0 +2 +4 +6 +8 +10 MORE DOTS THAN PAR ALL OR NOTHING FEW DOTS, FEW BOUNDARIES RUSSELL POORAN WATSON GAYLE YUVRAJ MAXWELL GILL POLLARD WARNER KOHLI DOT BALLS ABOVE OR BELOW PAR, PERCENTAGE POINTS BOUNDARIES ABOVE OR BELOW PAR
Batters with at least 1,500 deliveries. Right of the vertical line is more dot balls than the situation deserved, above the horizontal line is more boundaries.

Time in

Almost every batter needs ten balls, and one of them never did

Splitting each career at the tenth delivery of an innings shows how large the warm-up cost really is. Shane Watson struck at 104 for his first ten balls and 164 after them, a swing of sixty runs per hundred, and David Miller, Nicholas Pooran and Kieron Pollard are all within a run or two of the same gap.

The interesting name is at the bottom. Yashasvi Jaiswal is the only batter in this sample who is not faster once set, scoring at 154 early and 152 later, which is a different way of batting rather than a better one.

100 125 150 175 200 SHANE WATSON +60 DAVID MILLER +60 NICHOLAS POORAN +59 KIERON POLLARD +58 AB DE VILLIERS +54 CHRIS GAYLE +53 S BADRINATH +49 MIKE HUSSEY +48 GLENN MAXWELL +48 RAHUL TRIPATHI +10 HARDIK PANDYA +5 YASHASVI JAISWAL -2 FIRST 10 BALLS AFTER THAT SWING STRIKE RATE
Strike rate over the first ten balls of an innings against every ball after that, for batters with at least 400 early deliveries and 700 later ones.

The one the thresholds keep out

Vaibhav Suryavanshi has faced 447 deliveries across 2025 and 2026, which is short of the 400-early and 700-later cut this chart uses, so he does not appear on it. His split is worth printing anyway. He strikes at 207.8 over his first ten balls and 245.3 after them, so he warms up like everybody else, by about 37 runs per hundred.

What makes him an outlier is where he starts. His cold rate of 207.8 is faster than any batter on the chart manages once set, Nicholas Pooran included at 195.3. Across all 447 balls he is 75 runs per hundred above par, roughly double Virender Sehwag’s figure at the top of the leaderboard earlier, and he is dismissed 16 per cent more often than par to do it. Two seasons is not a career, and this is the number most likely to move.

Clearing the rope

Some bowlers are close to impossible to hit over the fence

Economy rates fold together singles, boundaries and sixes, so it is worth pulling out the shot that decides most T20 matches on its own. Jasprit Bumrah concedes a six on 3.7 per cent of his deliveries where the situations he bowls in are worth 6.4, a reduction of 42 per cent, and Dale Steyn and Lasith Malinga are barely behind him.

At the other end Karn Sharma is hit for a six roughly half as often again as par, and the list beside him is mostly part-time seam. It is worth noting that Andre Russell appears on the wrong end of this table and the right end of the batting one, which is a reasonable summary of his career.

+43% JASPRIT BUMRAH +40% DALE STEYN +36% LASITH MALINGA +32% SUNIL NARINE +32% BHUVNESHWAR KUMAR +30% ARSHDEEP SINGH +29% ZAHEER KHAN +25% KRUNAL PANDYA -12% KULDEEP YADAV -13% RAJAT BHATIA -20% JAYDEV UNADKAT -21% HARDIK PANDYA -22% ANDRE RUSSELL -48% KARN SHARMA PAR CONCEDES MORE SIXES CONCEDES FEWER SIXES
Sixes conceded against the par rate for the same phase, season and type of bowling. Bowlers with at least 1,500 deliveries.
Wickets that needed no fielder
BowlerWktsBowled or lbw
Rashid Khan17947%
Lasith Malinga17044%
Sunil Narine20740%
Mitchell Starc7639%
Piyush Chawla18938%
Jacques Kallis6537%
Rajat Bhatia7135%
Kieron Pollard6812%
Andre Russell12312%
Josh Hazlewood7214%
Best by phase · runs saved per 100 balls
BowlerPhaseEconSaved
Jofra ArcherOvers 1-66.88+26.9
Bhuvneshwar KumarOvers 1-66.29+25.2
Jasprit BumrahOvers 1-66.51+24.1
Lasith MalingaOvers 1-65.88+20.7
Jasprit BumrahOvers 7-166.34+28.9
Sunil NarineOvers 7-166.46+19.0
Rashid KhanOvers 7-166.77+14.8
Pat CumminsOvers 7-167.47+14.7
Jasprit BumrahOvers 17-207.86+32.0
Lasith MalingaOvers 17-207.47+29.9
Sunil NarineOvers 17-207.08+27.8
Dale SteynOvers 17-207.95+22.0

Rashid Khan takes 47 per cent of his wickets bowled or lbw, against 12 per cent for Andre Russell and Shardul Thakur, whose wickets almost always need someone to catch the ball. Bumrah and Sunil Narine are the only two bowlers who appear in the top six of all three phases.

Under a target

The best chaser in the competition is not the man you would name

Comparing each batter with himself, first innings against second, produces a spread of more than forty runs per hundred balls between the ends of it. Yusuf Pathan and Dwayne Smith were both around 23 runs per hundred better chasing than setting; AB de Villiers, whose reputation is built on chases, was 21 runs per hundred worse.

Better chasing than setting
BatterSettingChasingGap
Yusuf Pathan−1+22+23
Dwayne Smith+4+27+23
Adam Gilchrist+12+33+21
Axar Patel−20−4+16
Virender Sehwag+31+46+15
Nicholas Pooran+11+25+15
David Miller−8+6+13
Kieron Pollard+3+14+12
Better setting than chasing
BatterSettingChasingGap
AB de Villiers+31+10−21
Murali Vijay+11−6−17
Yuvraj Singh+9−5−15
Kane Williamson+1−11−12
Steve Smith+4−7−11
Gautam Gambhir+9−1−11
Chris Gayle+30+20−10
Ruturaj Gaikwad+2−6−9

Treat this one more lightly than the rest of the page. Half a career is a much smaller sample than a whole one, and a chase that finishes early takes its easiest deliveries with it.

The dry spell

A batter who has not hit a boundary for a while is not due one, he is in trouble

The idea that pressure builds until it bursts is so embedded in how the game is described that nobody checks it, and when you do check it the line runs the other way in every phase of the innings. In the middle overs a boundary arrives on 17.2 per cent of the balls that immediately follow another boundary, and on 11.8 per cent of the balls that follow fifteen or more without one, which is a difference of about a third.

Nothing resets. A long gap between boundaries is not tension accumulating, it is evidence about the pitch, the bowler and the batter that keeps being confirmed the longer it goes on.

0.10 0.14 0.18 0.22 0.26 0 1-2 3-5 6-9 10-14 15+ BALLS SINCE THE LAST BOUNDARY BY EITHER BATTER CHANCE THE NEXT BALL GOES TO THE BOUNDARY OVERS 1-6 OVERS 7-16 OVERS 17-20
Every legal ball since 2008, grouped by how many deliveries have passed since either batter last found the fence. The three lines are the three phases of an innings, measured separately so the shape is not just the powerplay showing through.

The bunny

A bowler who has dismissed you twice before is no more likely to do it again

Broadcast graphics love this match-up, and there is nothing behind it. Splitting every delivery by what had already happened between that particular batter and that particular bowler produces three groups that are indistinguishable, and not merely close: the run rates agree to three decimal places, at 1.332, 1.334 and 1.334 runs per ball.

Sixty-six thousand deliveries went into this and it still refuses to separate, which is a useful reminder that the longest individual rivalry in nineteen seasons of this competition has produced 149 balls, or about four overs a year.

0.00 0.02 0.04 0.06 4.40% NEVER OUT TO HIM 26,142 BALLS 4.48% OUT ONCE BEFORE 23,368 BALLS 4.32% OUT TWICE OR MORE 17,235 BALLS CHANCE THIS BALL TAKES HIS WICKET
Deliveries where the batter had faced this bowler at least a dozen times before, split by how often the bowler had already dismissed him. The dashed line is the first group's rate, carried across for comparison.

Getting set

The set batter scores faster and is also far likelier to get out

Being set is treated as a safety condition, as though the danger were all at the start of an innings and passed once a batter had seen a few deliveries, but the risk per ball climbs steadily the longer he stays. A batter on nothing goes at 1.05 runs a ball and is dismissed on 4.5 per cent of them; a batter past fifty goes at 1.86 and is dismissed on 7.4 per cent.

Both lines rise together because they are the same decision seen from two sides. The batter who has worked out the conditions starts taking the risks that the conditions allow, and the extra wickets are the price of the extra runs rather than a failure of concentration.

1.0 1.2 1.4 1.6 1.8 2.0 0-4 5-9 10-19 20-29 30-39 40-59 60+ RUNS THE BATTER ALREADY HAS 4% 5% 6% 7% 8% RUNS PER BALL CHANCE OF BEING OUT
Runs per ball on the left scale, chance of dismissal on the right, plotted against the score the batter had already reached before the delivery.

Two new batters

The three balls after a wicket are the safest in the innings, not the most dangerous

Commentary treats a fresh batter as a wicket waiting to happen, and captains attack accordingly, but in the powerplay the dismissal rate in the first three balls after a wicket falls is 3.3 per cent against 4.6 per cent a dozen balls later, and the middle overs show the same reversal at 3.0 against 4.9.

What actually collapses is the scoring, not the survival. Those same powerplay deliveries produce 0.88 runs a ball against 1.46 later in the phase, so the wicket buys the fielding side a period of quiet rather than a period of danger, and the danger returns only once the new pair has settled and started playing shots again.

3% 5% 7% 9% 0-2 3-5 6-11 12-23 24+ BALLS SINCE THE LAST WICKET FELL CHANCE THIS BALL TAKES A WICKET OVERS 1-6 OVERS 7-16 OVERS 17-20 A NEW PAIR AT THE CREASE
Chance of a wicket falling, grouped by how many balls have passed since the previous one, measured separately within each phase.

Momentum

The ball after a six really is different, and it is dangerous at both ends

Of all the folklore I tested this is the one that survived. Restricting the comparison to consecutive balls in the same over, so that the bowler and the field setting cannot change underneath the measurement, the ball after a six goes for 1.72 runs against a league average of 1.34, and finds the boundary on 24.5 per cent of occasions against 17.7.

The wicket line moves with it rather than against it, from 5.2 per cent on an average ball to 6.3 after a six, so a batter who has just cleared the rope is genuinely more likely to do something spectacular in either direction. Whatever is happening here, it is not calm.

0% 7% 14% 21% 28% 16.8 5.0 DOT 16.7 5.2 SINGLE 20.2 4.9 FOUR 24.5 6.3 SIX 17.7 5.2 ANY BALL WHAT HAPPENED ON THE PREVIOUS BALL OF THE SAME OVER BOUNDARY NEXT BALL WICKET NEXT BALL
Deliveries compared with the one immediately before them in the same over, so the bowler and the field are held constant. Dashed lines mark the all-ball averages.

Who bowls the seventeenth

Spin is the cheaper option almost to the end of the innings, and captains stop using it long before that

Spinners concede fewer runs per ball than seamers in every single over from the seventh to the eighteenth, and the gap is at its widest around the sixteenth, where spin goes at 1.43 and pace at 1.51. Over the same stretch the share of overs handed to spin falls from around sixty per cent to under a fifth.

The pattern only inverts in the last two overs, where pace finally becomes the better option, by which point spin is bowling about nine per cent of the deliveries anyway. There is a plausible defence of all this, since a captain is protecting his best seamers for the finish and a spinner who goes for four sixes has no way back, but the raw ledger says the seventeenth over is being given to the more expensive bowler.

1.0 1.2 1.4 1.6 1.8 PACE SPIN SHARE OF OVERS GIVEN TO SPIN 1 3 5 7 9 11 13 15 17 19 OVER OF THE INNINGS RUNS CONCEDED PER BALL
Runs conceded per ball by each type, over by over, with the share of overs actually given to spin shown beneath.

The last over

The most explosive over of the innings is also the one with the most dot balls

Sixes and dots are usually described as opposites, and across the closing overs they climb together. The twentieth over produces a six on 11.3 per cent of its deliveries, more than any other, and it also produces more balls that score nothing than the eighteenth or nineteenth do.

This is what an innings looks like when every ball is being attacked without exception. The yorker that beats the swing and the slower ball that is missed entirely both come out as dots, and a wicket falls on 13.6 per cent of the deliveries, roughly one in seven.

0% 9% 18% 27% 36% 16 17 18 19 20 OVER OF THE INNINGS BALLS HIT FOR SIX BALLS THAT SCORE NOTHING BALLS THAT TAKE A WICKET
The closing five overs of an innings, showing the share of deliveries producing a six, a dot and a wicket.

Chasing

Knowing the target makes the closing overs harder, not easier

A chasing side is supposed to have the advantage of information, and through the powerplay it does score a little faster, but at the death the position reverses and the team batting first scores 1.62 runs a ball against the chasing team's 1.52.

Some of that is arithmetic rather than nerves, because a chase that is comfortably won stops needing runs, but the wicket rate does not fall to match, which suggests the gap is not only about teams coasting home.

1.2 1.3 1.4 1.5 1.6 1.7 1.24 1.29 OVERS 1-6 1.26 1.27 OVERS 7-16 1.62 1.52 OVERS 17-20 BATTING FIRST CHASING RUNS PER BALL
Runs per ball by phase, split by whether the batting side was setting a target or chasing one.

The ceiling

Past about twelve an over, asking for more runs does not produce more runs

Teams respond to a rising required rate exactly as you would expect until the target passes roughly twelve an over, and then the response stops. Chases needing twelve to fifteen score 1.48 runs a ball; chases needing more than fifteen score 1.46, which is slightly less.

The dismissal rate keeps climbing regardless, from 6.0 per cent to 8.2, so the extra ambition converts into wickets rather than runs. There appears to be a hard limit on how fast a T20 side can be made to score, and beyond it the only thing that changes is how quickly the innings ends.

1.2 1.3 1.4 1.5 UNDER 6 6-8 8-10 10-12 12-15 15+ RUNS PER OVER STILL REQUIRED 4% 6% 8% RUNS PER BALL SCORED CHANCE OF BEING OUT SCORING STOPS RESPONDING
Chases with at least two overs remaining, grouped by the run rate still required. Runs scored on the left, dismissals on the right.

The boundary boom

Scoring has risen by a fifth in nineteen years and almost none of it is fours

The competition's run rate has gone up by about twenty per cent since 2008, and the natural assumption is that batting has improved across the board. It has not. The share of deliveries hit for six has risen by seventy-nine per cent, from one ball in twenty-one to one in twelve, while the share hit for four has moved by six per cent, which over nineteen seasons is close to nothing.

3% 6% 9% 12% 15% FOURS SIXES 08 09 12 15 18 21 24 26 SHARE OF ALL DELIVERIES +79% SINCE 2008 +6%
Share of all legal deliveries producing a four and a six, by season.

The pitch report

Grounds differ enormously in runs and hardly at all in wickets

Venues are separated by a third in scoring, from 1.16 runs a ball at Kingsmead to 1.51 at Mullanpur, which is the sort of gap that dominates a pitch report. Plot those same grounds against how often a wicket falls there and the relationship disappears.

A high-scoring ground is not a ground where batting is safe, and a low-scoring one is not a ground where bowlers get rewarded. Sides appear to adjust their risk to the conditions until the wicket column comes out roughly level everywhere, which also explains something I found when trying to predict deliveries: knowing which stadium a ball was bowled at added nothing at all.

4.5% 5.0% 5.5% 6.0% 1.15 1.25 1.35 1.45 RUNS PER BALL AT THIS GROUND NO SLOPE AT ALL CHANCE OF A WICKET PER BALL AT THIS GROUND KINGSMEAD MULLANPUR
The twenty grounds with at least 3,000 deliveries. Horizontal position is scoring, vertical is the wicket rate, and the dashed line is the average across all of them.

What "par" means here

Every delivery is compared with what the whole league managed in the same season, the same phase of the innings and against the same type of bowling, which removes the twenty per cent scoring inflation of the last nineteen years and most of the advantage of batting in a particular slot. It does not remove the quality of the opposition, the ground or the state of the match, so a batter who spent his career against weak attacks will still come out looking good.

Colophon

Ball-by-ball records for 1,243 matches across seasons 2008 to 2026, covering 295,557 deliveries of which 284,465 were legal. Bowling actions and batting stances were reconstructed separately for all 577 bowlers, since the source data records neither. Every group compared above holds at least a few thousand deliveries unless stated otherwise, and the individual head-to-head figures are printed because they are interesting rather than because they are evidence, since even the longest rivalry here is only 149 balls.

The par comparisons and the counter-intuitive findings are descriptions of what happened, not causal claims. Where this page says that something adds nothing, that comes from prediction models fitted on the 2008 to 2023 seasons and scored on the two most recent ones, and those seasons were looked at more than once while the work was going on, so the smaller differences should be treated as worth checking again rather than settled.

The working, unedited

The unedited working behind the run figures: how the outcome of each delivery was modelled, what was tried, and what failed. Code, output and figures exactly as they ran.

Ball-level run prediction

Predicts the conditional distribution of runs off the bat given that a delivery is legal: how likely each of 0, 1, 2, 3, 4 and 6 is.

Same parquet as the wicket notebook, built by build_ball_data.py, and the same rule: every feature describes the state of the match before the delivery, and career/venue/matchup features are snapshots from strictly earlier matches. The checks below verify the intended pre-ball construction; evaluation leakage is discussed separately.

Three framing decisions, all of which matter more than the model choice:

  • The outcome is categorical, not continuous. Runs off the bat take six values and the gaps are real: a ball is 40× more likely to go for 4 than for 3. A regression on runs spends its capacity on a number line that is mostly empty, and its residuals are wildly non-Gaussian. We fit a 6-way multinomial and read expected runs off the predicted distribution when a point estimate is wanted.
  • Legal deliveries only. This is a conditional model: legality is not known before the bowler releases the ball. An unconditional live predictor or simulator also needs a legality/extras model. Wides and no-balls are 3.8% of rows.
  • The bar is the training prior, not accuracy. Predicting "1 run" on every ball scores 38% accuracy and is worth nothing. We score multiclass log loss against the prior, plus RMSE of expected runs and calibration.

Unlike wickets, there is real signal here: expect roughly a 3–4% log-loss gain rather than 2%. Whether a batter scores is more predictable than whether they get out.

Evaluation status

This is a strong exploratory backtest, not a pristine one-shot holdout. Earlier versions of the analysis repeatedly inspected 2025–26 during ablation, pruning, calibration choice and style-feature experiments. Temporal splitting prevents future rows entering model fitting, but repeated analyst use of the final period can still bias model-selection claims. Small deltas should be confirmed by a future season or a rolling-origin evaluation.

Final comparisons therefore include paired 95% bootstrap intervals that resample whole matches. Random-seed variation only measures optimiser stability and is not a confidence interval.

[1]
import lightgbm as lgb
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl
import xgboost as xgb
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import log_loss, roc_auc_score
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, SplineTransformer, StandardScaler

# feature lists and the labelling rule live in the script so notebook and script never drift
from model_baseline import CATEGORICAL, LINEAR, SPLINE
from model_runs import CATS, FEATS, OUTCOMES, expected_runs, frames, label
from evaluation import clustered_log_loss_gain

BLUE, ORANGE, GREY, GREEN = "#2a78d6", "#eb6834", "#8a8a85", "#3f9e6a"
plt.rcParams.update({"figure.dpi": 120, "axes.spines.top": False,
                     "axes.spines.right": False, "axes.grid": True,
                     "grid.alpha": 0.25, "grid.linewidth": 0.6, "font.size": 9})

raw = pl.read_parquet("ball_data.parquet")
df = label(raw)          # legal balls only, runs off the bat -> class index 0..5
print(f"{raw.height:,} deliveries -> {df.height:,} legal  "
      f"({100 * (1 - df.height / raw.height):.1f}% wides/no-balls dropped)")
print(f"{df['match_id'].n_unique():,} matches · {df['year'].min()}-{df['year'].max()}")
print(f"mean runs off the bat: {df['_runs'].mean():.4f} per legal ball")


# Fail early if the generated table no longer matches the notebook's assumptions.
required = {"runs_batter", "is_legal", "match_id", "year", "over", "legal_balls"}
assert not (required - set(raw.columns)), f"missing columns: {sorted(required - set(raw.columns))}"
assert raw["runs_batter"].null_count() == 0
assert set(df["_runs"].unique().to_list()) == set(OUTCOMES)
assert df.select(pl.struct(["match_id", "innings", "over", "ball_in_over"]).n_unique()).item() == df.height
295,557 deliveries -> 284,465 legal  (3.8% wides/no-balls dropped)
1,243 matches · 2008-2026
mean runs off the bat: 1.3340 per legal ball

The shape of the outcome

Three quarters of legal balls are a dot or a single. Boundaries are 18% of balls but 61% of the runs, which is why a model that only gets the mean right is not much use: the interesting question at any point in an innings is how likely is a boundary, not what is the average.

[2]
mix = (df.group_by("_runs").agg(pl.len().alias("n"))
         .with_columns((pl.col("n") / df.height).alias("share"),
                       (pl.col("_runs") * pl.col("n")).alias("runs"))
         .with_columns((pl.col("runs") / pl.col("runs").sum()).alias("run_share"))
         .sort("_runs"))
display(mix)

fig, axes = plt.subplots(1, 2, figsize=(9.5, 3.2))
x = np.arange(len(OUTCOMES))
axes[0].bar(x - 0.2, mix["share"], width=0.4, color=BLUE, label="share of balls")
axes[0].bar(x + 0.2, mix["run_share"], width=0.4, color=ORANGE, label="share of runs")
axes[0].set(xticks=x, xlabel="runs off the bat", ylabel="share",
            title="18% of balls, 61% of the runs")
axes[0].set_xticklabels(OUTCOMES)
axes[0].legend(frameon=False, fontsize=8)

by_over = (df.filter(pl.col("over") < 20).group_by("over")
             .agg(pl.col("_runs").mean().alias("runs per ball"),
                  (pl.col("_runs") >= 4).mean().alias("boundary rate"),
                  (pl.col("_runs") == 0).mean().alias("dot rate")).sort("over"))
for col, colour in [("runs per ball", BLUE), ("boundary rate", ORANGE), ("dot rate", GREEN)]:
    axes[1].plot(by_over["over"] + 1, by_over[col], "o-", color=colour, lw=2, ms=4, label=col)
axes[1].set(xlabel="over", title="Powerplay, squeeze, death",
            xticks=range(1, 21, 2))
axes[1].legend(frameon=False, fontsize=8)
plt.tight_layout()
shape: (6, 5)
_runsnsharerunsrun_share
i64u32f64i64f64
01063030.37369400.0
11090530.3833621090530.287386
2182590.064187365180.096235
38420.0029625260.006657
4343400.1207181373600.361983
6156680.055079940080.247738
notebook figure

Note the shape in the right panel, because it is the single most important thing the model has to learn and it is not the monotone climb that the wicket rate showed. Scoring rises through the powerplay, falls off a cliff the moment the field goes out (boundary rate halves from 0.220 in over 6 to 0.114 in over 7), then climbs back through the middle and steeply at the death.

The dot rate does not mirror it: it keeps falling straight through that cliff (0.431 to 0.362) and only turns up in the last over. Batters do not stop scoring when the field spreads, they switch from boundaries to rotating strike. Boundary rate in over 5 (0.222) is close to over 20 (0.248) and the balls that produce them have almost nothing else in common, which is why is_powerplay turns out to be worth more to the model than over itself.

Split: temporal, never random

Balls from the same over are near-duplicates. A random split would let the model memorise a match and report a score it cannot reproduce on a future game.

[3]
levels, (train, val, test) = frames()
ytr, yva, yte = (d["y"].to_numpy() for d in (train, val, test))

PRIOR = np.bincount(ytr, minlength=len(OUTCOMES)) / len(ytr)   # everything is scored against this
LABELS = list(range(len(OUTCOMES)))
runs_te = np.array(OUTCOMES, dtype=float)[yte]                 # actual runs on each test ball

for name, d, y in [("train", train, ytr), ("val", val, yva), ("test", test, yte)]:
    print(f"{name:6s} {d.height:>7,} balls  {d['year'].min()}-{d['year'].max()}  "
          f"mean runs {np.array(OUTCOMES, float)[y].mean():.4f}")

print("\ntrain prior:  " + "  ".join(f"{v}:{p:.4f}" for v, p in zip(OUTCOMES, PRIOR)))

NULL = np.tile(PRIOR / PRIOR.sum(), (len(yte), 1))
NULL_LL = log_loss(yte, NULL, labels=LABELS)
NULL_RMSE = float(np.sqrt(np.mean((expected_runs(NULL) - runs_te) ** 2)))
print(f"null log loss {NULL_LL:.5f} · null E[runs] RMSE {NULL_RMSE:.4f}")


assert train["year"].max() < val["year"].min() <= val["year"].max() < test["year"].min()
assert set(train["match_id"].unique()).isdisjoint(set(val["match_id"].unique()))
assert set(train["match_id"].unique()).isdisjoint(set(test["match_id"].unique()))
train  234,995 balls  2008-2023  mean runs 1.2923
val     16,299 balls  2024-2024  mean runs 1.5048
test    33,171 balls  2025-2026  mean runs 1.5448

train prior:  0:0.3805  1:0.3839  2:0.0653  3:0.0032  4:0.1174  6:0.0498
null log loss 1.40050 · null E[runs] RMSE 1.8607

Baseline: multinomial logistic regression with splines

Same preprocessing as the wicket baseline (spline bases on the features whose effect is curved, plain standardised terms on the rest), but the head is a 6-way softmax instead of a binary logit. Splines earn their place more here than they did for wickets: over has a genuine non-monotone effect on scoring (up, down, flat, up), and no single linear coefficient can express that.

[4]
FE = SPLINE + LINEAR + CATEGORICAL
X = lambda d: d.select(FE).to_pandas()

# SimpleImputer runs first: the chase features (rrr, balls_left, ...) are null for every
# first-innings ball, and SplineTransformer cannot take NaN. add_indicator keeps a flag so
# the model still knows the value was missing rather than median.
pre = ColumnTransformer([
    ("spline", Pipeline([
        ("impute", SimpleImputer(strategy="median", add_indicator=True)),
        ("spline", SplineTransformer(n_knots=6, degree=3, include_bias=False)),
        ("scale", StandardScaler()),
    ]), SPLINE),
    ("linear", Pipeline([
        ("impute", SimpleImputer(strategy="median", add_indicator=True)),
        ("scale", StandardScaler()),
    ]), LINEAR),
    ("cat", OneHotEncoder(handle_unknown="ignore", drop="first"), CATEGORICAL),
])

lr = Pipeline([("pre", pre), ("lr", LogisticRegression(max_iter=1000, C=1.0))])
lr.fit(X(train), ytr)

print(f"{len(FE)} features -> {lr.named_steps['pre'].transform(X(val)).shape[1]} columns "
      f"-> {len(OUTCOMES)} class probabilities")
36 features -> 103 columns -> 6 class probabilities

Scoring

Four numbers, because no single one covers a distributional forecast:

  • Multiclass log loss against the training prior. The headline.
  • RMSE of expected runs, the price of using the model as a point predictor. It will barely move, because the irreducible variance of a single ball is enormous, and that is the point.
  • Boundary AUC: can the model rank the balls that go for 4 or 6? This is where the usable signal lives.
  • Dot AUC, the same question at the other end.
[5]
def norm(p):
    """Guard against float drift so log_loss does not complain."""
    return p / p.sum(axis=1, keepdims=True)


def report(name, y, p, runs=None):
    p = norm(p)
    ll = log_loss(y, p, labels=LABELS)
    null_ll = log_loss(y, np.tile(PRIOR, (len(y), 1)), labels=LABELS)
    runs = np.array(OUTCOMES, float)[y] if runs is None else runs
    rmse = np.sqrt(np.mean((expected_runs(p) - runs) ** 2))
    bnd = roc_auc_score((y >= 4).astype(int), p[:, 4] + p[:, 5])
    dot = roc_auc_score((y == 0).astype(int), p[:, 0])
    print(f"{name:6s} logloss {ll:.5f}  vs null {null_ll:.5f} "
          f"({100 * (null_ll - ll) / null_ll:+.2f}%)   E[runs] RMSE {rmse:.4f}   "
          f"boundary AUC {bnd:.4f}   dot AUC {dot:.4f}")


for name, d, y in [("train", train, ytr), ("val", val, yva), ("test", test, yte)]:
    report(name, y, lr.predict_proba(X(d)))
train  logloss 1.28270  vs null 1.33259 (+3.74%)   E[runs] RMSE 1.6035   boundary AUC 0.6296   dot AUC 0.6292
val    logloss 1.34244  vs null 1.39681 (+3.89%)   E[runs] RMSE 1.7766   boundary AUC 0.6344   dot AUC 0.6104
test   logloss 1.35260  vs null 1.40050 (+3.42%)   E[runs] RMSE 1.8236   boundary AUC 0.6154   dot AUC 0.6115

Calibration

For a distributional model this is the thing that matters. Two views: the point estimate (when the model says 1.6 expected runs, do those balls average 1.6?) and the boundary probability (when it says 25%, is it 25%?).

[6]
p_lr = norm(lr.predict_proba(X(test)))


def cal_curve(score, actual, bins=10):
    return (pl.DataFrame({"s": score, "a": actual.astype(float)})
              .with_columns(pl.col("s").qcut(bins, labels=[str(i) for i in range(bins)])
                              .alias("b"))
              .group_by("b").agg(pl.len().alias("n"),
                                 pl.col("s").mean().alias("predicted"),
                                 pl.col("a").mean().alias("actual")).sort("b"))


e_lr = cal_curve(expected_runs(p_lr), runs_te)
b_lr = cal_curve(p_lr[:, 4] + p_lr[:, 5], (yte >= 4).astype(int))
display(e_lr)

fig, axes = plt.subplots(1, 2, figsize=(8.6, 4.2))
for ax, c, lo, hi, xl, t in [
        (axes[0], e_lr, 0.8, 2.0, "predicted E[runs]", "Expected runs (deciles)"),
        (axes[1], b_lr, 0.05, 0.32, "predicted P(boundary)", "Boundary probability (deciles)")]:
    ax.plot([lo, hi], [lo, hi], color=GREY, lw=1, ls="--")
    ax.plot(c["predicted"], c["actual"], "o-", color=BLUE, lw=2, ms=6)
    ax.set(xlabel=xl, ylabel="observed", title=t, xlim=(lo, hi), ylim=(lo, hi))
    ax.set_aspect("equal")
plt.tight_layout()
shape: (10, 4)
bnpredictedactual
catu32f64f64
"0"33180.8481140.950271
"1"33171.0831071.203497
"2"33171.2281531.343081
"3"33171.3421571.424178
"4"33171.4420821.49593
"5"33171.5431651.565873
"6"33171.6550891.715104
"7"33171.7965291.79801
"8"33172.003891.907748
"9"33172.4975942.044619
notebook figure

Gradient boosting

Trees should do better here for the same reasons as in the wicket model: thresholds found without a spline basis, and interactions an additive model cannot express. The interaction that matters most for runs is over × wickets_in_hand: at over 18 with 8 wickets in hand the batting side swings at everything, and at over 18 with 1 wicket left it does not. The logistic regression has to average those two situations together.

ground joins as a native categorical. The wicket notebook found venue worth nothing, but noted that grounds differ substantially in scoring (1.30 to 1.51 runs per ball), so this is the model where it should finally pay. The ablation below tests that.

[7]
def frame(d, feats=FEATS):
    """Trees take features raw: no splines, no scaling, NaN handled natively."""
    Xd = d.select(feats).to_pandas()
    for c in CATS:
        if c in feats:
            Xd[c] = pd.Categorical(Xd[c], categories=levels[c])
    return Xd


Xtr, Xva, Xte = (frame(d) for d in (train, val, test))

PARAMS = dict(objective="multiclass", num_class=len(OUTCOMES), n_estimators=3000,
              learning_rate=0.03, num_leaves=15, min_child_samples=400, subsample=0.8,
              subsample_freq=1, colsample_bytree=0.7, reg_lambda=10.0, verbose=-1)

lgbm = lgb.LGBMClassifier(**PARAMS, random_state=0)
lgbm.fit(Xtr, ytr, eval_X=Xva, eval_y=yva, eval_metric="multi_logloss",
         callbacks=[lgb.early_stopping(100, verbose=False)])

xgbm = xgb.XGBClassifier(
    objective="multi:softprob", num_class=len(OUTCOMES), n_estimators=3000,
    learning_rate=0.03, max_depth=4, min_child_weight=50, subsample=0.8,
    colsample_bytree=0.7, reg_lambda=10.0, enable_categorical=True, tree_method="hist",
    eval_metric="mlogloss", early_stopping_rounds=100, random_state=0)
xgbm.fit(Xtr, ytr, eval_set=[(Xva, yva)], verbose=False)

print(f"LightGBM stopped at {lgbm.best_iteration_} rounds "
      f"({lgbm.best_iteration_ * len(OUTCOMES)} trees, one per class per round)")
print(f"XGBoost  stopped at {xgbm.best_iteration} rounds")
LightGBM stopped at 258 rounds (1548 trees, one per class per round)
XGBoost  stopped at 444 rounds
[8]
preds = {
    "null (train prior)": NULL,
    "multinomial + splines": p_lr,
    "lightgbm": norm(lgbm.predict_proba(Xte)),
    "xgboost": norm(xgbm.predict_proba(Xte)),
}
preds["lgb + xgb"] = (preds["lightgbm"] + preds["xgboost"]) / 2

rows = []
for name, p in preds.items():
    p = norm(p)                       # float drift alone is enough to upset log_loss
    ll = log_loss(yte, p, labels=LABELS)
    null_row = name.startswith("null")
    rows.append({"model": name, "logloss": round(ll, 5),
                 "gain_%": round(100 * (NULL_LL - ll) / NULL_LL, 2),
                 "E_runs_RMSE": round(float(np.sqrt(np.mean((expected_runs(p) - runs_te) ** 2))), 4),
                 "boundary_AUC": None if null_row else round(roc_auc_score((yte >= 4).astype(int), p[:, 4] + p[:, 5]), 4),
                 "dot_AUC": None if null_row else round(roc_auc_score((yte == 0).astype(int), p[:, 0]), 4)})
display(pl.DataFrame(rows))
shape: (5, 6)
modelloglossgain_%E_runs_RMSEboundary_AUCdot_AUC
strf64f64f64f64f64
"null (train prior)"1.40050.01.8607nullnull
"multinomial + splines"1.35263.421.82360.61540.6115
"lightgbm"1.351423.51.82050.61580.6156
"xgboost"1.350843.551.82010.61560.6158
"lgb + xgb"1.350823.551.82020.61590.6158

Where the gain is, class by class

The single log-loss number hides which parts of the distribution improved. Splitting it into six one-vs-rest problems is worth doing because the answer is not uniform.

[9]
p_lgb = preds["lightgbm"]
rows = []
for i, v in enumerate(OUTCOMES):
    yb = (yte == i).astype(int)
    base = np.full(len(yb), PRIOR[i])
    rows.append({"runs": v, "share": round(float(yb.mean()), 4),
                 "AUC": round(roc_auc_score(yb, p_lgb[:, i]), 4),
                 "logloss": round(log_loss(yb, p_lgb[:, i], labels=[0, 1]), 5),
                 "vs_prior_%": round(100 * (log_loss(yb, base, labels=[0, 1])
                                            - log_loss(yb, p_lgb[:, i], labels=[0, 1]))
                                     / log_loss(yb, base, labels=[0, 1]), 2)})
display(pl.DataFrame(rows))
shape: (6, 5)
runsshareAUCloglossvs_prior_%
i64f64f64f64f64
00.33950.61560.623813.18
10.3820.61470.646032.86
20.05660.57870.216180.96
30.00180.63090.013143.34
40.13810.59990.394312.26
60.0820.62820.278354.93

The six is the most predictable outcome in this backtest, at 4.9% over its prior and AUC 0.628, comfortably ahead of the dot (3.2%) and the four (2.3%). That ordering makes sense: a six needs a batter who clears the rope and a situation that asks them to, and both are things the feature set knows. A four is much more often a good ball that beat the field.

The twos are the floor at 1.0%, which is also as expected: running two is a function of where the ball goes and how fast the fielder is, not of the match state. (The 3s row reads 2.8% but that class is 0.18% of balls, so it is noise.)

Did it capture the shape?

Compare the model's mean prediction against the observed rate, bucketed by a real feature. As in the wicket notebook this is deliberately not a partial-dependence sweep: over, legal_balls, phase and is_powerplay all encode "how late is it", the fit splits that effect across them, and varying one alone describes deliveries that never occur.

[10]
scored = test.with_columns(pl.Series("e_runs", expected_runs(p_lgb)),
                           pl.Series("p_bnd", p_lgb[:, 4] + p_lgb[:, 5]))


def shape(expr, xlabel, ax, pred, obs, min_n=300):
    g = (scored.with_columns(expr.alias("bucket"))
                .group_by("bucket").agg(pl.len().alias("n"),
                                        pl.col(pred).mean().alias("predicted"),
                                        obs.mean().alias("actual"))
                .filter(pl.col("n") >= min_n).sort("bucket"))
    ax.plot(g["bucket"], g["actual"], "o-", color=BLUE, lw=2, ms=5, label="observed")
    ax.plot(g["bucket"], g["predicted"], "s--", color=ORANGE, lw=2, ms=5, label="model")
    ax.set(xlabel=xlabel)
    ax.legend(frameon=False, fontsize=8)
    return g


fig, axes = plt.subplots(1, 3, figsize=(12.5, 3.4))
shape(pl.col("over") + 1, "over", axes[0], "e_runs", pl.col("_runs"))
axes[0].set(ylabel="runs per ball", title="Runs by over", xticks=range(1, 21, 3))
shape(pl.col("over") + 1, "over", axes[1], "p_bnd", (pl.col("_runs") >= 4).cast(pl.Float64))
axes[1].set(ylabel="P(boundary)", title="Boundary rate by over", xticks=range(1, 21, 3))
shape(pl.col("bat_balls") // 5 * 5, "balls faced by the striker", axes[2], "e_runs", pl.col("_runs"))
axes[2].set(ylabel="runs per ball", title="Getting your eye in")
plt.tight_layout()
notebook figure

Importance, and why you should not trust it

Gain importance rewards a feature for every split it wins, which flatters high-cardinality categoricals: ground has 37 levels, so it gets many chances to carve off a slice of noise. The ablation is the check that matters: retrain without a feature group and see whether the loss actually moves. Here that diagnostic was run on the final period, so it is exploratory and must not be mistaken for an untouched test result.

[11]
gains = sorted(zip(FEATS, lgbm.booster_.feature_importance("gain")), key=lambda t: -t[1])
total = sum(v for _, v in gains)
top = [(k, 100 * v / total) for k, v in gains[:12]][::-1]

fig, ax = plt.subplots(figsize=(6, 3.6))
colours = [ORANGE if k in ("ground", "venue_runs_per_ball", "venue_wicket_rate") else BLUE
           for k, _ in top]
ax.barh([k for k, _ in top], [v for _, v in top], color=colours, height=0.7)
ax.set(xlabel="% of total gain", title="LightGBM importance (orange = venue, see ablation)")
ax.grid(axis="y", visible=False)
plt.tight_layout()
notebook figure
[12]
def ablate(drop):
    feats = [f for f in FEATS if f not in drop]
    m = lgb.LGBMClassifier(**PARAMS, random_state=0)
    m.fit(frame(train, feats), ytr, eval_X=frame(val, feats), eval_y=yva,
          eval_metric="multi_logloss", callbacks=[lgb.early_stopping(100, verbose=False)])
    p = norm(m.predict_proba(frame(test, feats)))
    ll = log_loss(yte, p, labels=LABELS)
    return {"dropped": ", ".join(drop) if drop else "(nothing)",
            "logloss": round(ll, 5),
            "gain_%": round(100 * (NULL_LL - ll) / NULL_LL, 2),
            "boundary_AUC": round(roc_auc_score((yte >= 4).astype(int), p[:, 4] + p[:, 5]), 4)}


display(pl.DataFrame([
    ablate([]),
    ablate(["ground"]),
    ablate(["ground", "venue_runs_per_ball", "venue_wicket_rate"]),
    ablate(["h2h_dismissals", "h2h_balls"]),
    ablate(["bat_career_sr", "bat_career_dismissal_rate", "bat_career_balls"]),
]))
shape: (5, 4)
droppedloglossgain_%boundary_AUC
strf64f64f64
"(nothing)"1.351423.50.6158
"ground"1.348823.690.6171
"ground, venue_runs_per_ball, v…1.349853.620.6169
"h2h_dismissals, h2h_balls"1.351553.50.6154
"bat_career_sr, bat_career_dism…1.356113.170.6127

Permutation importance, and the collinearity trap

Gain importance is measured while the trees are built. Permutation importance is measured on a finished model: shuffle one column of the final-period data (an exploratory diagnostic, not a locked test) so it keeps its distribution but no longer lines up with its rows, re-score, and charge the feature whatever that cost, expressed as a share of the model's skill (the log-loss ground it gains over predicting the prior).

The catch is that a shuffled feature can be covered for by a surviving duplicate. Four columns here encode the clock (over, legal_balls, is_powerplay, phase), so breaking any one leaves the model three others. Shuffling them as a block is the honest measurement.

[13]
rng = np.random.default_rng(0)
ll_base = log_loss(yte, p_lgb, labels=LABELS)
skill = NULL_LL - ll_base


def perm_pct(cols, repeats=3):
    """% of skill lost when `cols` are shuffled together. Returns (mean, sd) over repeats."""
    Xp, losses = Xte.copy(), []
    for _ in range(repeats):
        idx = rng.permutation(len(Xp))      # one index for the whole group, so a collinear set
        for c in cols:                      # stays internally consistent (legal_balls is still
            Xp[c] = Xte[c].values[idx]      # 6 x over) and only its link to the outcome breaks
        losses.append(log_loss(yte, norm(lgbm.predict_proba(Xp)), labels=LABELS))
    for c in cols:
        Xp[c] = Xte[c]
    return 100 * (np.mean(losses) - ll_base) / skill, 100 * np.std(losses) / skill


groups = {f: [f] for f in FEATS}
groups["clock (over+legal_balls+is_powerplay+phase)"] = ["over", "legal_balls", "is_powerplay", "phase"]

rows = []
for name, cols in groups.items():
    pct, sd = perm_pct(cols)
    rows.append({"feature": name, "pct_of_skill": round(pct, 1), "sd": round(sd, 1)})

perm = pl.DataFrame(rows).sort("pct_of_skill", descending=True)
display(perm.head(18))
shape: (18, 3)
featurepct_of_skillsd
strf64f64
"clock (over+legal_balls+is_pow…54.21.6
"is_powerplay"12.40.2
"team_runs"10.60.7
"team_wickets"8.50.3
"legal_balls"7.90.5
"rrr"2.40.2
"phase"2.10.2
"wickets_in_hand"1.90.1
"bat_sr"1.10.1
"bat_career_balls"1.10.2

Reading it

The clock block is 53% of skill. Individually is_powerplay is 11.9%, legal_balls 7.7%, phase 2.0% and over 0.6%. They sum to 22%, less than half the joint figure, because each was measured while the other three were still standing. When the model is told how late it is, it is telling you most of what it knows.

Note also which of the four survives being measured alone: is_powerplay, a single bit, beats the over number by 20×. The field restriction is the thing; the over is mostly a proxy for it.

The next tier is genuinely different information rather than another view of the clock: team_runs (11%), team_wickets (8%), bat_career_sr (6%). That third one is the interesting entry and it has no counterpart in the wicket model: who is on strike matters for scoring in a way it barely does for dismissals. Dropping the three batter-career columns in the ablation above costs more than dropping anything else.

At the bottom, the same negatives as the wicket notebook: venue_runs_per_ball, bat_is_new, bowl_runs and balls_left all score at or below zero. Permuting them helps slightly, the signature of a column the model has no use for.


A second model, cut to what survived

Two findings feed into this:

  1. Most features do nothing. Two thirds of the table scores inside the shuffle noise, and every one of them still costs the model capacity.
  2. ground is the third-largest feature by gain and the ablation says dropping it improves the test loss. Venue does not earn its place here either.

So: keep the 15 that carry weight, and regularise harder now that there is less to overfit to (reg_lambda 10 → 30, min_child_samples 400 → 1000, learning rate 0.03 → 0.02). Averaging 5 seeds removes the run-to-run wobble.

[14]
KEEP = ["is_powerplay", "team_runs", "team_wickets", "legal_balls", "bat_career_sr",
        "balls_since_wicket", "bat_runs", "bat_dot_pct", "bat_balls",
        "bat_career_dismissal_rate", "bowl_career_runs_per_ball", "rrr", "phase",
        "bat_sr", "bat_career_balls"]
print(f"{len(FEATS)} features -> {len(KEEP)}:  {', '.join(KEEP)}\n")

NEW = dict(PARAMS, n_estimators=4000, learning_rate=0.02, min_child_samples=1000,
           reg_lambda=30.0)


def ensemble(feats, params, n_seeds=5):
    """Fit n_seeds models and average their predicted distributions.

    Returns predictions on val as well as test: the calibration step below needs
    a set of predictions the model was not fitted on and that is not the test set.
    """
    a, b, c = (frame(d, feats) for d in (train, val, test))
    pv, pt = [], []
    for s in range(n_seeds):
        m = lgb.LGBMClassifier(**{**params, "random_state": s})
        m.fit(a, ytr, eval_X=b, eval_y=yva, eval_metric="multi_logloss",
              callbacks=[lgb.early_stopping(150, verbose=False)])
        pv.append(norm(m.predict_proba(b)))
        pt.append(norm(m.predict_proba(c)))
    return np.mean(pv, axis=0), np.mean(pt, axis=0)


p_new_va, p_new = ensemble(KEEP, NEW)

rows = []
for name, p in [("null (train prior)", NULL),
                ("multinomial + splines", p_lr),
                ("old: 37 features", p_lgb),
                ("new: 15 features, 5 seeds", p_new)]:
    p = norm(p)
    ll = log_loss(yte, p, labels=LABELS)
    null_row = name.startswith("null")
    rows.append({"model": name, "logloss": round(ll, 5),
                 "gain_%": round(100 * (NULL_LL - ll) / NULL_LL, 2),
                 "E_runs_RMSE": round(float(np.sqrt(np.mean((expected_runs(p) - runs_te) ** 2))), 4),
                 "boundary_AUC": None if null_row else round(roc_auc_score((yte >= 4).astype(int), p[:, 4] + p[:, 5]), 4)})
display(pl.DataFrame(rows))
37 features -> 15:  is_powerplay, team_runs, team_wickets, legal_balls, bat_career_sr, balls_since_wicket, bat_runs, bat_dot_pct, bat_balls, bat_career_dismissal_rate, bowl_career_runs_per_ball, rrr, phase, bat_sr, bat_career_balls
shape: (4, 5)
modelloglossgain_%E_runs_RMSEboundary_AUC
strf64f64f64f64
"null (train prior)"1.40050.01.8607null
"multinomial + splines"1.35263.421.82360.6154
"old: 37 features"1.351423.51.82050.6158
"new: 15 features, 5 seeds"1.349663.631.81760.6164
[15]
# Calibration of the final model, both views, against the baseline.
fig, axes = plt.subplots(1, 2, figsize=(8.6, 4.2))
for ax, score, actual, lo, hi, xl, t in [
        (axes[0], lambda p: expected_runs(p), runs_te, 0.8, 2.1,
         "predicted E[runs]", "Expected runs (deciles)"),
        (axes[1], lambda p: p[:, 4] + p[:, 5], (yte >= 4).astype(int), 0.05, 0.33,
         "predicted P(boundary)", "Boundary probability (deciles)")]:
    ax.plot([lo, hi], [lo, hi], color=GREY, lw=1, ls="--")
    for nm, p, colour, mk in [("multinomial + splines", p_lr, BLUE, "o"),
                              ("lgbm, 15 features", p_new, ORANGE, "s")]:
        c = cal_curve(score(p), actual)
        ax.plot(c["predicted"], c["actual"], marker=mk, ls="-", color=colour, lw=2, ms=5, label=nm)
    ax.set(xlabel=xl, ylabel="observed", title=t, xlim=(lo, hi), ylim=(lo, hi))
    ax.set_aspect("equal")
axes[0].legend(frameon=False, fontsize=8, loc="upper left")
plt.tight_layout()
notebook figure

The era problem

Both calibration curves sit above the diagonal almost everywhere, and the shape plots show the same thing: the model has the shape right and the level too low. This is not a modelling error, it is drift. The training seasons averaged 1.29 runs per ball and the test seasons average 1.54. T20 scoring rose about 20% over the window, and a model fitted mostly on 2008–2019 cricket carries the old level with it.

[16]
by_year = (pl.concat([train, val, test]).group_by("year")
             .agg(pl.col("_runs").mean().alias("actual")).sort("year"))
e_new = expected_runs(p_new)
print(f"train mean {np.array(OUTCOMES, float)[ytr].mean():.4f}  ->  "
      f"test mean {runs_te.mean():.4f}  (+{100 * (runs_te.mean() / np.array(OUTCOMES, float)[ytr].mean() - 1):.1f}%)")
print(f"model mean prediction on test: {e_new.mean():.4f}  "
      f"({e_new.mean() - runs_te.mean():+.4f} per ball)")

fig, ax = plt.subplots(figsize=(6.5, 3))
ax.plot(by_year["year"], by_year["actual"], "o-", color=BLUE, lw=2, ms=4, label="observed")
ax.axvspan(2024.5, 2026.5, color=ORANGE, alpha=0.10)
ax.scatter(test.select(pl.col("year")).to_series().unique().sort(),
           [pl.DataFrame({"y": test["year"], "e": e_new}).filter(pl.col("y") == yr)["e"].mean()
            for yr in sorted(test["year"].unique().to_list())],
           color=ORANGE, marker="s", zorder=3, label="model, test seasons")
ax.set(xlabel="season", ylabel="runs per ball", title="Scoring drifted up; the model did not")
ax.legend(frameon=False, fontsize=8)
plt.tight_layout()
train mean 1.2923  ->  test mean 1.5448  (+19.5%)
model mean prediction on test: 1.4755  (-0.0693 per ball)
notebook figure

The obvious fix, training only on recent seasons, does not work. Refitting the 15-feature model on 2018–2023 scores 3.53% and on 2021–2023 scores 3.49%, both worse than the 3.64% from all seasons, and neither closes the level gap (mean prediction 1.462 and 1.456 against the full model's 1.473). Throwing away 60–80% of the rows costs more in variance than the stale level costs in bias, and the level is not really what the trees learned from the old seasons anyway: the shape transfers fine, which is why log loss holds up.

The fix is a recalibration step rather than a data change: learn the correction on the validation season and apply it to the test predictions, leaving the model alone. That is the next section.

Recalibrating on the most recent season

The model has the ranking right and the level wrong, so the correction should move the level without touching the ranking. Two candidates, both fitted on 2024 and applied unchanged to 2025–26. 2024 is also the validation season used for early stopping. It is separate from 2025–26, but it is not a fully independent calibration set; the apparent benefit is therefore exploratory. In a production pipeline, use a separate calibration window or out-of-fold predictions:

  • Prior shift. For each outcome, compare how often the model said it would happen on the validation season against how often it did, and scale that class by the ratio. Six numbers.
  • Temperature. Divide the logits by a scalar fitted to minimise validation log loss. The standard fix for over- or under-confidence.

They address different failures, and only one of them is the failure we have.

[17]
def prior_shift(p_fit, y_fit):
    """Per-class multiplicative correction: observed frequency / mean predicted probability."""
    obs = np.bincount(y_fit, minlength=len(OUTCOMES)) / len(y_fit)
    return obs / p_fit.mean(axis=0)


def temper(p, T):
    z = np.log(np.clip(p, 1e-12, None)) / T
    z -= z.max(axis=1, keepdims=True)
    return norm(np.exp(z))


W = prior_shift(p_new_va, yva)
grid = np.linspace(0.6, 1.6, 101)
T = float(grid[int(np.argmin([log_loss(yva, temper(p_new_va, t), labels=LABELS) for t in grid]))])

print("per-class weights fitted on 2024:")
for v, w in zip(OUTCOMES, W):
    print(f"  {v}: {w:.3f}")
print(f"\nfitted temperature: {T:.2f}")

p_cal = norm(p_new * W)

rows = []
for name, p in [("null (train prior)", NULL),
                ("15 features, uncalibrated", p_new),
                (f"+ temperature (T={T:.2f})", temper(p_new, T)),
                ("+ prior shift", p_cal)]:
    p = norm(p)
    ll = log_loss(yte, p, labels=LABELS)
    null_row = name.startswith("null")
    rows.append({"model": name, "logloss": round(ll, 5),
                 "gain_%": round(100 * (NULL_LL - ll) / NULL_LL, 2),
                 "mean_E_runs": round(float(expected_runs(p).mean()), 4),
                 "E_runs_RMSE": round(float(np.sqrt(np.mean((expected_runs(p) - runs_te) ** 2))), 4),
                 "boundary_AUC": None if null_row else round(roc_auc_score((yte >= 4).astype(int), p[:, 4] + p[:, 5]), 4)})
display(pl.DataFrame(rows))
print(f"actual mean runs per ball on test: {runs_te.mean():.4f}")

fig, axes = plt.subplots(1, 2, figsize=(8.6, 4.2))
for ax, score, actual, lo, hi, xl, t in [
        (axes[0], lambda p: expected_runs(p), runs_te, 0.8, 2.3,
         "predicted E[runs]", "Expected runs (deciles)"),
        (axes[1], lambda p: p[:, 4] + p[:, 5], (yte >= 4).astype(int), 0.05, 0.36,
         "predicted P(boundary)", "Boundary probability (deciles)")]:
    ax.plot([lo, hi], [lo, hi], color=GREY, lw=1, ls="--")
    for nm, p, colour, mk in [("uncalibrated", p_new, ORANGE, "s"),
                              ("prior shift", p_cal, GREEN, "D")]:
        cc = cal_curve(score(p), actual)
        ax.plot(cc["predicted"], cc["actual"], marker=mk, ls="-", color=colour, lw=2, ms=5, label=nm)
    ax.set(xlabel=xl, ylabel="observed", title=t, xlim=(lo, hi), ylim=(lo, hi))
    ax.set_aspect("equal")
axes[0].legend(frameon=False, fontsize=8, loc="upper left")
plt.tight_layout()
per-class weights fitted on 2024:
  0: 0.939
  1: 1.027
  2: 1.013
  3: 0.688
  4: 1.029
  6: 1.133

fitted temperature: 1.04
shape: (4, 6)
modelloglossgain_%mean_E_runsE_runs_RMSEboundary_AUC
strf64f64f64f64f64
"null (train prior)"1.40050.01.29231.8607null
"15 features, uncalibrated"1.349663.631.47551.81760.6164
"+ temperature (T=1.04)"1.349423.651.50851.81640.6163
"+ prior shift"1.347573.781.55311.81710.6163
actual mean runs per ball on test: 1.5448
notebook figure

Prior shift appears to address the level error here; temperature does not. Six numbers fitted on one season take the model from 3.63% to 3.78% over the prior, and the mean prediction from 1.476 to 1.553 against an actual 1.545: the level error goes from 0.071 runs per ball to 0.008. Boundary AUC is unchanged at 0.616, which is the point: the ranking was never the problem and the correction did not disturb it.

Temperature buys 0.01% and should be read as a null result. T comes out at 1.04, a slight flattening, and the mean prediction it produces (1.507) drifts up only as a side effect of pushing mass toward the rare high-scoring classes. It is fixing confidence, and the model was not overconfident, it was aimed at the wrong era.

Look at the weights themselves, because they are the era effect in six numbers: sixes ×1.14, fours ×1.03, dots ×0.94. The change in T20 batting between the training seasons and now is almost entirely six-hitting, which is what anyone watching would have told you, and the fitted correction reflects that same era shift. (The 3s weight of 0.69 is noise: that class is 0.3% of balls.)

Two caveats, both visible in the plot. The correction is a single global adjustment, so it lands the average and slightly over-corrects the top decile: the highest-scoring situations now predict 2.27 expected runs against an observed 2.03, where before they were roughly right. Trading a small error everywhere for a larger one in the tail is the right trade for log loss, but if the death overs are what you care about, fit the shift on those balls only.

And 2024 itself averaged 1.505 runs per ball against the test seasons' 1.545, so the correction is fitted on an era already slightly behind the one it is applied to. It still overshoots the mean rather than undershooting it, because the multiplicative form amplifies the tail classes. In production this should be refitted every season.

An innings, ball by ball

What the model is actually for. Expected runs and boundary probability tracked through a single chase, with the delivered runs underneath.

[18]
mid = test.filter(pl.col("innings") == 2)["match_id"][0]
inn = (test.with_columns(pl.Series("e_runs", expected_runs(p_cal)),
                         pl.Series("p_bnd", p_cal[:, 4] + p_cal[:, 5]))
           .filter((pl.col("match_id") == mid) & (pl.col("innings") == 2))
           .with_row_index("ball"))

fig, axes = plt.subplots(2, 1, figsize=(9.5, 4.6), sharex=True,
                         gridspec_kw={"height_ratios": [2, 1]})
axes[0].plot(inn["ball"], inn["e_runs"], color=BLUE, lw=1.6, label="E[runs]")
ax2 = axes[0].twinx()
ax2.plot(inn["ball"], inn["p_bnd"], color=ORANGE, lw=1.2, alpha=0.8, label="P(boundary)")
ax2.set_ylabel("P(boundary)", color=ORANGE)
ax2.grid(False)
axes[0].set(ylabel="E[runs]", title=f"{inn['batting_team'][0]} chasing "
            f"{inn['target_runs'][0]} v {inn['bowling_team'][0]} ({inn['date'][0]})")
axes[1].bar(inn["ball"], inn["_runs"], color=GREY, width=0.9)
axes[1].set(xlabel="ball of the innings", ylabel="runs", yticks=[0, 2, 4, 6])
plt.tight_layout()
notebook figure

Bowler type: the feature that was not in the data

The next-steps list above called bowler type the largest missing signal, on the grounds that the powerplay/middle-overs shape is substantially a pace-versus-spin split the model has to infer from the over number. Cricsheet does not record it, so it had to be fetched.

fetch_bowler_style.py builds player_style.csv from three sources chained together: Cricsheet's own register maps every player to a Cricinfo ID, Wikidata maps that ID to an en.wikipedia article, and the article's infobox carries bowling and batting fields in prose that normalises cleanly. Cricinfo itself is blocked to scripted requests, so its ID is used only as a join key. Wikidata's own bowling style property (P2545) exists and is unusable - 43 of our 577 bowlers, and a value distribution that is visibly bot-damaged. 25 uncapped players with no article were filled in by hand; see docs/BOWLER_STYLE_TODO.md. Coverage is 577/577 bowlers, every delivery.

None of it is derived from the ball data, so there is nothing to leak: a bowler's action and a batter's stance are fixed properties of the player, known before the season starts.

Four columns come out of it, and the last two are the interesting ones:

column
bowler_type pace / spin
bat_hand the striker's stance, for 99.7% of balls faced
same_handed right-arm to a right-hander, left to a left
turn_into_batter which way a spinner's ball moves relative to the bat

turn_into_batter is the real cricket variable. Finger spin turns into a batter of the same handedness as the bowling arm - an off-break into a right-hander, orthodox into a left-hander - and wrist spin turns into the opposite one. It is null for pace, where turn is not the variable.

[19]
style = df                      # already legal-only, with _runs, from the first cell

fig, axes = plt.subplots(1, 3, figsize=(12.5, 3.4))

# --- when each type bowls, which is what the model has been inferring ---
mix = (style.filter(pl.col("over") < 20).group_by("over")
            .agg((pl.col("bowler_type") == "spin").mean().alias("spin")).sort("over"))
axes[0].bar(mix["over"] + 1, mix["spin"], color=BLUE, width=0.72)
axes[0].set(xlabel="over", ylabel="share of balls bowled by spin",
            title="Spin bowls the middle", xticks=range(1, 21, 3))

# --- runs by type and phase ---
g = (style.group_by(["bowler_type", "is_powerplay"])
          .agg(pl.col("_runs").mean().alias("rpb"),
               (pl.col("_runs") >= 4).mean().alias("bnd")).sort("bowler_type", "is_powerplay"))
x = np.arange(2)
for i, (t, colour) in enumerate([("pace", ORANGE), ("spin", GREEN)]):
    v = g.filter(pl.col("bowler_type") == t).sort("is_powerplay")
    axes[1].bar(x + (i - 0.5) * 0.38, v["rpb"], width=0.36, color=colour, label=t)
axes[1].set(xticks=x, ylabel="runs per ball", title="The gap is a middle-overs gap")
axes[1].set_xticklabels(["overs 7-20", "powerplay"])
axes[1].legend(frameon=False, fontsize=8)

# --- the matchup ---
m = (style.filter(pl.col("bowler_type") == "spin")
          .drop_nulls("turn_into_batter").group_by("turn_into_batter")
          .agg(pl.len().alias("n"), pl.col("_runs").mean().alias("rpb"),
               (pl.col("_runs") >= 4).mean().alias("bnd")).sort("turn_into_batter"))
axes[2].bar(["turns away", "turns in"], m["rpb"], color=[GREY, BLUE], width=0.55)
for i, (r, n) in enumerate(zip(m["rpb"], m["n"])):
    axes[2].annotate(f"{r:.3f}\n({n:,} balls)", (i, r), ha="center", va="bottom",
                     xytext=(0, 2), textcoords="offset points", fontsize=8)
axes[2].set(ylabel="runs per ball", title="Spin: which way it turns", ylim=(0, 1.55))
plt.tight_layout()

display(style.group_by("bowler_type").agg(
    pl.len().alias("balls"), pl.col("_runs").mean().round(4).alias("runs_per_ball"),
    (pl.col("_runs") >= 4).mean().round(4).alias("boundary_rate"),
    (pl.col("_runs") == 0).mean().round(4).alias("dot_rate"),
    pl.col("wicket").mean().round(4).alias("wicket_rate")).sort("bowler_type"))
shape: (2, 6)
bowler_typeballsruns_per_ballboundary_ratedot_ratewicket_rate
stru32f64f64f64f64
"pace"1821161.37110.19420.38910.0541
"spin"1023491.26790.14310.34620.0461
notebook figure

The raw gaps are big. Pace concedes 1.442 runs per ball outside the powerplay against spin's 1.271, with boundary rates of 18.5% and 13.6%. Spin turning into the bat goes for 1.322 against 1.230 turning away - the ball that comes on with the turn is the one you can hit through the line.

The middle panel is the warning, though: inside the powerplay pace and spin are level (1.266 vs 1.253). The type gap is a middle-overs gap, and the model already knows what over it is.

[20]
STYLE = ["bowler_type"]
MATCHUP = ["bowler_type", "bowler_arm", "bat_hand", "same_handed", "turn_into_batter"]
CATS_X = CATS + ["bowler_type", "bowler_arm", "bat_hand"]
levels_x = {c: sorted(raw[c].unique().drop_nulls().to_list()) for c in CATS_X}


def frame_x(d, feats):
    Xd = d.select(feats).to_pandas()
    for c in CATS_X:
        if c in feats:
            Xd[c] = pd.Categorical(Xd[c], categories=levels_x[c])
    return Xd


def ensemble_x(feats, params=NEW, n_seeds=5):
    a, b, c = (frame_x(d, feats) for d in (train, val, test))
    ps = []
    for s in range(n_seeds):
        m = lgb.LGBMClassifier(**{**params, "random_state": s})
        m.fit(a, ytr, eval_X=b, eval_y=yva, eval_metric="multi_logloss",
              callbacks=[lgb.early_stopping(150, verbose=False)])
        ps.append(norm(m.predict_proba(c)))
    return np.mean(ps, axis=0)


rows = []
style_predictions = {}
for name, feats in [("15 features (baseline)", KEEP),
                    ("+ bowler_type", KEEP + STYLE),
                    ("+ bowler_type, bat_hand", KEEP + ["bowler_type", "bat_hand"]),
                    ("+ full matchup", KEEP + MATCHUP)]:
    p = norm(ensemble_x(feats))
    style_predictions[name] = p
    ll = log_loss(yte, p, labels=LABELS)
    rows.append({"model": name, "n_feats": len(feats), "logloss": round(ll, 5),
                 "gain_%": round(100 * (NULL_LL - ll) / NULL_LL, 2),
                 "boundary_AUC": round(roc_auc_score((yte >= 4).astype(int), p[:, 4] + p[:, 5]), 4)})
    if name == "+ full matchup":
        p_style = p
display(pl.DataFrame(rows))
shape: (4, 5)
modeln_featsloglossgain_%boundary_AUC
stri64f64f64f64
"15 features (baseline)"151.349663.630.6164
"+ bowler_type"161.348353.720.6174
"+ bowler_type, bat_hand"171.348333.720.6174
"+ full matchup"201.348113.740.6168
[21]
# Paired uncertainty: resample matches, not individual balls.
comparisons = [
    ("pruned vs old 37-feature model", p_lgb, p_new),
    ("prior shift vs uncalibrated", p_new, p_cal),
    ("+ bowler type vs pruned", p_new, style_predictions["+ bowler_type"]),
    ("+ full matchup vs pruned", p_new, style_predictions["+ full matchup"]),
]
uncertainty = []
for name, reference, candidate in comparisons:
    ci = clustered_log_loss_gain(
        yte, reference, candidate, test["match_id"].to_numpy(), n_boot=5_000
    )
    uncertainty.append({
        "comparison": name,
        "gain_x1e4": round(10_000 * ci.estimate, 2),
        "95%_low_x1e4": round(10_000 * ci.low, 2),
        "95%_high_x1e4": round(10_000 * ci.high, 2),
    })
display(pl.DataFrame(uncertainty))
shape: (4, 4)
comparisongain_x1e495%_low_x1e495%_high_x1e4
strf64f64f64
"pruned vs old 37-feature model"17.591.7633.03
"prior shift vs uncalibrated"20.8514.8326.98
"+ bowler type vs pruned"13.124.6221.6
"+ full matchup vs pruned"15.547.623.48

Positive values favour the named candidate; the units are $10^{-4}$ log-loss points. A 95% interval that crosses zero means the apparent ordering is not stable across matches. This is a more demanding and useful check than comparing random seeds on the same fixed rows.

What it was worth

Smaller than the raw gaps suggest, but stable across matches within this explored final period.

model log loss gain over prior
15 features (baseline) 1.34966 +3.63%
+ bowler_type 1.34835 +3.72%
+ bowler_type, bat_hand 1.34833 +3.72%
+ full matchup 1.34811 +3.74%

The match-clustered interval for bowler type versus the pruned model is wholly positive on 2025–26. That is better evidence than seed variation, although it is still exploratory because this period was inspected while choosing features. A future-season or rolling-origin repeat is needed for an unbiased confirmation.

The size of the effect is modest because the model already has correlated proxies: bowl_career_runs_per_ball partly identifies bowler type, while is_powerplay and phase identify when that type tends to bowl. The new information is the residual within-bowler and within-phase matchup.

Adding batting hand to bowler type barely changes the point estimate. The full matchup improves over the no-style baseline, but this comparison does not isolate whether its extra gain over bowler_type alone is reliable. A direct paired interval between those two variants would be the appropriate test before making that narrower claim.

Style still earns consideration because it is actionable: type and matchup are levers a captain can change when choosing who bowls the next over. That interpretability matters even when the aggregate log-loss increment is small.


Combining matchup features with independent calibration

The earlier style model and calibrator both looked useful, but the calibrator shared 2024 with early stopping. This version gives every period exactly one job:

period job
2008–2022 estimate model parameters while selecting tree count against 2023
2023 early stopping only
2008–2023 refit the selected model at its fixed tree count
2024 fit six regularized class intercepts only
2025–26 evaluation only within this experiment

The calibration layer adds one bias to each class log-probability and renormalizes. It cannot change delivery rankings or relearn feature effects; it can only correct class prevalence drift. The rare classes are weakly regularized.

[22]
from model_runs_calibrated import fit_independent_style_model

independent = fit_independent_style_model(df, n_seeds=5, calibration_l2=1.0)
assert np.array_equal(independent.y_test, yte)

print(f"tree counts selected on 2023: {list(independent.best_iterations)}")
print("class biases fitted on 2024: " + "  ".join(
    f"{runs}:{bias:+.3f}" for runs, bias in zip(OUTCOMES, independent.bias)))

rows = []
for name, p in [
        ("null (train prior)", NULL),
        ("previous best: pruned + prior shift", p_cal),
        ("full matchup, independent raw", independent.raw_test),
        ("full matchup + independent bias", independent.calibrated_test)]:
    ll = log_loss(yte, p, labels=LABELS)
    null_row = name.startswith("null")
    rows.append({
        "model": name,
        "logloss": round(ll, 6),
        "gain_%": round(100 * (NULL_LL - ll) / NULL_LL, 3),
        "mean_E_runs": round(float(expected_runs(p).mean()), 4),
        "E_runs_RMSE": round(float(np.sqrt(np.mean((expected_runs(p) - runs_te) ** 2))), 4),
        "boundary_AUC": None if null_row else round(
            roc_auc_score((yte >= 4).astype(int), p[:, 4:].sum(axis=1)), 4),
    })
display(pl.DataFrame(rows))

comparisons = [
    ("independent calibration vs its raw model",
     independent.raw_test, independent.calibrated_test),
    ("combined model vs previous calibrated best", p_cal, independent.calibrated_test),
    ("combined model vs original 37-feature model", p_lgb, independent.calibrated_test),
]
rows = []
for name, reference, candidate in comparisons:
    ci = clustered_log_loss_gain(
        yte, reference, candidate, independent.test_match_ids, n_boot=5_000
    )
    rows.append({
        "comparison": name,
        "gain_x1e4": round(10_000 * ci.estimate, 2),
        "95%_low_x1e4": round(10_000 * ci.low, 2),
        "95%_high_x1e4": round(10_000 * ci.high, 2),
    })
display(pl.DataFrame(rows))
tree counts selected on 2023: [346, 383, 386, 339, 308]
class biases fitted on 2024: 0:-0.188  1:-0.104  2:-0.116  3:-0.500  4:-0.072  6:+0.000
shape: (4, 6)
modelloglossgain_%mean_E_runsE_runs_RMSEboundary_AUC
strf64f64f64f64f64
"null (train prior)"1.4004970.01.29231.8607null
"previous best: pruned + prior …1.3475743.7791.55311.81710.6163
"full matchup, independent raw"1.3484033.721.47051.81830.6157
"full matchup + independent bia…1.3465433.8531.55721.81740.6158
shape: (3, 4)
comparisongain_x1e495%_low_x1e495%_high_x1e4
strf64f64f64
"independent calibration vs its…18.612.6224.74
"combined model vs previous cal…10.312.9717.65
"combined model vs original 37-…48.7532.2865.32

Positive interval bounds mean the combined model wins across match resamples, not merely across random training seeds. The comparison with the previous best is the key row: it tests whether style and independently fitted calibration compound rather than duplicate one another.

The historical caveat still applies. Previous notebook versions inspected 2025–26, so this is a methodologically cleaner experiment but not a magically restored pristine holdout. A future IPL season or rolling-origin rerun remains the honest confirmation.


Reading the result

There is more signal in runs than in wickets. The independently calibrated full-matchup ensemble is the current best; its exact loss, gain, and uncertainty are in the table above, and the boundary AUC of 0.616 is the number to look at: given one ball that went to the fence and one that did not, the model ranks them correctly 62% of the time knowing nothing about the delivery itself. The logistic regression is not far behind at 3.42%, which suggests much of the structure here is additive; the trees' edge is small and should be read with the clustered interval above.

The point estimate barely improves, and that is not a failure. Expected-runs RMSE goes from 1.861 (prior) to 1.818. A single ball is 0, 1 or 4 and no amount of context changes that; the variance is in the outcome, not in the estimate. Anyone reporting run prediction as an RMSE is reporting a number that cannot move. The distribution is the product: P(boundary) on this ball, aggregated over an over or an innings, and its calibration should be checked explicitly rather than assumed.

Venue does not pay, even here. This was the ablation worth running: the wicket notebook found ground worthless but flagged that grounds differ by 1.30 to 1.51 runs per ball, so scoring was where it should have mattered. It does not. Dropping ground and both venue-rate columns makes the model slightly better. The between-ground scoring spread is real in the raw data but it is already implied by what the other features say about how the innings is going: teams and conditions are not independent of the state they produce.

What does pay, and did not for wickets, is the batter. bat_career_sr is the top non-clock feature after the two innings-state columns, and dropping the batter-career block costs more than any other ablation. Strike rate is a stable player property in a way that dismissal rate is not.

The independently calibrated matchup model is the strongest run component. The wicket and run models together are the useful object: one gives P(wicket) and the other the run distribution on the same delivery, which are useful inputs to a win-probability or simulation layer. A valid simulator still needs legality/extras and the joint dependence between runs and wickets; multiplying two independent marginal models is not enough.

The level was stale even though the shape was not, and that was the cheapest fix in the notebook. Scoring rose ~20% across the window, and the uncalibrated model predicted 1.476 runs per ball against an actual 1.545. AUC is blind to that because ranking can remain unchanged; log loss is not, and the calibration error is one reason it worsens, which is exactly why a calibration plot belongs in every writeup of a model like this. Regularized class biases fitted on the independent 2024 calibration window improve the full matchup model from 3.72% to 3.85% gain. The ranking barely changes; the win is calibration.

Next steps, roughly in order of expected value

  1. Bowler type (pace/spin) and handedness matchup - both done, see the section above. Together with independent calibration, they reach 3.85% gain. The style-only increment is modest; calibration contributes the larger final step. The raw effects are large and almost entirely confounded with the clock.
  2. Ordinal structure in the head. The six classes are not exchangeable: a model that confuses 4 and 6 is doing better than one that confuses 0 and 6, and multiclass log loss does not know that. Worth trying a cumulative-link head, or scoring with a distance-weighted loss.
  3. Model legality and extras, so the forecast is unconditional before the ball is bowled.
  4. Feed P(wicket) in as a feature, or fit the seven-way joint outcome (six run values plus dismissal) directly. The two processes compete for the same ball.

The working, unedited

The companion notebook, predicting whether a delivery takes a wicket rather than how many runs it concedes. Same data, same splits, a much harder problem.

Ball-level wicket prediction

Predicts, for every delivery, the probability that a wicket falls on it.

Data comes from build_ball_data.py, which flattens the Cricsheet JSON into one row per delivery. Every feature describes the state of the match before the ball is bowled, and career/venue/matchup features are snapshots from strictly earlier matches. The checks below verify the intended pre-ball construction; evaluation leakage is discussed separately.

Two things to keep in mind throughout:

  • The base rate is ~5%, so accuracy is a useless metric here: predicting "no wicket" on every ball scores 95%. We use log loss, PR-AUC and calibration instead.
  • This is a near-irreducible problem. A realistic ROC-AUC is 0.60–0.68. Anything much higher means something leaked.

Evaluation status

This is a strong exploratory backtest, not a pristine one-shot holdout. Earlier versions of the analysis repeatedly inspected 2025–26 while pruning features, comparing calibrators and testing additions. The temporal split prevents future rows entering model fitting, but repeated analyst use of the final period can still bias model-selection claims. Treat small improvements as hypotheses until they survive a future season or a rolling-origin evaluation.

Accordingly, the final comparisons now include paired 95% bootstrap intervals that resample whole matches. Seed-to-seed variation measures optimiser stability; it does not measure uncertainty about performance on future matches.

[1]
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (average_precision_score, brier_score_loss,
                             log_loss, roc_auc_score)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, SplineTransformer, StandardScaler

# feature lists live in the script so notebook and script never drift
from model_baseline import CATEGORICAL, LINEAR, SPLINE, TARGET
from evaluation import clustered_log_loss_gain

BLUE, ORANGE, GREY = "#2a78d6", "#eb6834", "#8a8a85"
plt.rcParams.update({"figure.dpi": 120, "axes.spines.top": False,
                     "axes.spines.right": False, "axes.grid": True,
                     "grid.alpha": 0.25, "grid.linewidth": 0.6, "font.size": 9})

df = pl.read_parquet("ball_data.parquet")
print(f"{df.height:,} deliveries · {df['match_id'].n_unique():,} matches · "
      f"{df['year'].min()}–{df['year'].max()} · {df.width} columns")
print(f"base rate: {df[TARGET].mean():.4f}")

# Fail early if the generated table no longer matches the notebook's assumptions.
required = {TARGET, "match_id", "year", "over", "legal_balls"}
assert not (required - set(df.columns)), f"missing columns: {sorted(required - set(df.columns))}"
assert df[TARGET].null_count() == 0
assert set(df[TARGET].unique().to_list()) <= {0, 1}
assert df.select(pl.struct(["match_id", "innings", "over", "ball_in_over"]).n_unique()).item() == df.height
295,557 deliveries · 1,243 matches · 2008–2026 · 69 columns
base rate: 0.0496

The signal we expect to find

Before modelling, look at the single strongest driver: how late in the innings it is.

[2]
by_over = (df.group_by("over").agg(pl.col(TARGET).mean().alias("rate"))
             .sort("over").filter(pl.col("over") < 20))

fig, ax = plt.subplots(figsize=(6.5, 3))
ax.bar(by_over["over"] + 1, by_over["rate"], color=BLUE, width=0.72)
ax.axhline(df[TARGET].mean(), color=GREY, lw=1, ls="--")
ax.annotate(f"league mean {df[TARGET].mean():.3f}", (0.7, df[TARGET].mean()),
            xytext=(0, 5), textcoords="offset points", color=GREY, fontsize=8)
ax.annotate(f"{by_over['rate'][-1]:.3f}", (20, by_over["rate"][-1]), ha="center",
            xytext=(0, 4), textcoords="offset points", color=BLUE, fontsize=8)
ax.set(xlabel="over", ylabel="wicket rate per ball",
       title="Wicket rate climbs steeply at the death", xticks=range(1, 21))
plt.tight_layout()
notebook figure

Split: temporal, never random

Balls from the same over are near-duplicates. A random split would let the model memorise a match and report a score it cannot reproduce on a future game.

[3]
train = df.filter(pl.col("year") <= 2023)
val   = df.filter(pl.col("year") == 2024)
test  = df.filter(pl.col("year") >= 2025)

for name, d in [("train", train), ("val", val), ("test", test)]:
    print(f"{name:6s} {d.height:>7,} balls  {d['year'].min()}–{d['year'].max()}  "
          f"rate {d[TARGET].mean():.4f}")

BASE = train[TARGET].mean()  # everything is scored against this constant

assert train["year"].max() < val["year"].min() <= val["year"].max() < test["year"].min()
assert set(train["match_id"].unique()).isdisjoint(set(val["match_id"].unique()))
assert set(train["match_id"].unique()).isdisjoint(set(test["match_id"].unique()))
train  243,656 balls  2008–2023  rate 0.0493
val     17,103 balls  2024–2024  rate 0.0516
test    34,798 balls  2025–2026  rate 0.0501

Why splines

Logistic regression is linear in each feature. Several of ours are not. The clearest case is the striker's score, where the wicket rate is roughly flat from 0–40 and then steps up once the batter is set and starts taking risks:

striker on 0–9 20–29 40–49 50–59 80–99 100+
wicket rate 0.0438 0.0488 0.0523 0.0629 0.0732 0.0771

A single linear coefficient fits one slope through that and misses both the flat stretch and the step. A cubic spline basis (6 knots) lets the model bend. Features with genuinely monotone-ish effects stay as plain standardised terms.

[4]
print(f"spline ({len(SPLINE)}):\n  " + ", ".join(SPLINE))
print(f"\nlinear ({len(LINEAR)}):\n  " + ", ".join(LINEAR))
print(f"\ncategorical: {CATEGORICAL}")

FEATS = SPLINE + LINEAR + CATEGORICAL
X = lambda d: d.select(FEATS).to_pandas()
y = lambda d: d[TARGET].to_numpy()
spline (9):
  bat_runs, bat_balls, over, crr, rrr, balls_since_boundary, bat_balls_since_boundary, team_wickets, bowl_balls

linear (26):
  team_runs, legal_balls, wickets_in_hand, balls_since_wicket, bat_dot_pct, bat_sr, bat_is_new, bowl_runs, bowl_wickets, runs_required, balls_left, rrr_minus_crr, is_powerplay, is_chase, innings, bat_career_balls, bat_career_dismissal_rate, bat_career_sr, bowl_career_balls, bowl_career_wicket_rate, bowl_career_runs_per_ball, venue_wicket_rate, venue_runs_per_ball, h2h_dismissals, h2h_balls, toss_won

categorical: ['phase']
[5]
# SimpleImputer runs first: the chase features (rrr, balls_left, ...) are null for every
# first-innings ball, and SplineTransformer cannot take NaN. add_indicator keeps a flag so
# the model still knows the value was missing rather than median.
pre = ColumnTransformer([
    ("spline", Pipeline([
        ("impute", SimpleImputer(strategy="median", add_indicator=True)),
        ("spline", SplineTransformer(n_knots=6, degree=3, include_bias=False)),
        ("scale", StandardScaler()),
    ]), SPLINE),
    ("linear", Pipeline([
        ("impute", SimpleImputer(strategy="median", add_indicator=True)),
        ("scale", StandardScaler()),
    ]), LINEAR),
    ("cat", OneHotEncoder(handle_unknown="ignore", drop="first"), CATEGORICAL),
])

model = Pipeline([("pre", pre), ("lr", LogisticRegression(max_iter=2000, C=1.0))])
model.fit(X(train), y(train))

print(f"{len(FEATS)} features → {model.named_steps['pre'].transform(X(val)).shape[1]} columns")
36 features → 103 columns

Scoring

The bar to clear is the null model: predicting the training base rate on every ball. A model that cannot beat that constant has learned nothing, however good its accuracy looks.

[6]
def report(name, d):
    yt, p = y(d), model.predict_proba(X(d))[:, 1]
    null_ll = log_loss(yt, np.full_like(p, BASE))
    ll = log_loss(yt, p)
    print(f"{name:6s} logloss {ll:.5f}  vs null {null_ll:.5f} ({100*(null_ll-ll)/null_ll:+.2f}%)"
          f"   ROC-AUC {roc_auc_score(yt, p):.4f}   PR-AUC {average_precision_score(yt, p):.4f}"
          f"   Brier {brier_score_loss(yt, p):.5f}")

for name, d in [("train", train), ("val", val), ("test", test)]:
    report(name, d)

print(f"\nPR-AUC baseline (random) = base rate = {test[TARGET].mean():.4f}")
train  logloss 0.19150  vs null 0.19657 (+2.58%)   ROC-AUC 0.6261   PR-AUC 0.0841   Brier 0.04635
val    logloss 0.20015  vs null 0.20334 (+1.57%)   ROC-AUC 0.6001   PR-AUC 0.0793   Brier 0.04862
test   logloss 0.19547  vs null 0.19887 (+1.71%)   ROC-AUC 0.6067   PR-AUC 0.0788   Brier 0.04726

PR-AUC baseline (random) = base rate = 0.0501

Calibration

For a probability model this matters more than AUC: when the model says 8%, do 8% of those balls actually produce a wicket? Points on the diagonal are perfectly calibrated.

[7]
p_test = model.predict_proba(X(test))[:, 1]
cal = (pl.DataFrame({"p": p_test, "y": y(test)})
         .with_columns(pl.col("p").qcut(10, labels=[str(i) for i in range(10)]).alias("bin"))
         .group_by("bin").agg(pl.len().alias("n"),
                              pl.col("p").mean().alias("predicted"),
                              pl.col("y").mean().alias("actual")).sort("bin"))
display(cal)

fig, ax = plt.subplots(figsize=(4.2, 4.2))
lim = float(max(cal["predicted"].max(), cal["actual"].max())) * 1.1
ax.plot([0, lim], [0, lim], color=GREY, lw=1, ls="--")
ax.annotate("perfect", (lim * 0.72, lim * 0.78), color=GREY, fontsize=8, rotation=38)
ax.plot(cal["predicted"], cal["actual"], "o-", color=BLUE, lw=2, ms=6)
ax.set(xlabel="predicted probability", ylabel="observed wicket rate",
       title="Calibration on 2025–26 (deciles)", xlim=(0, lim), ylim=(0, lim))
ax.set_aspect("equal")
plt.tight_layout()
shape: (10, 4)
binnpredictedactual
catu32f64f64
"0"34800.0252040.026437
"1"34800.0316490.031322
"2"34800.0357680.038218
"3"34790.0394860.041679
"4"34800.0433160.04023
"5"34800.0475910.05431
"6"34790.0530170.052601
"7"34800.0609630.052874
"8"34800.0758990.072989
"9"34800.1177750.090517
notebook figure

Did the spline capture the shape?

Compare the model's mean prediction against the observed rate, bucketed by the feature.

Note what this deliberately is not: a partial-dependence sweep holding every other feature at its median. That is misleading here, because over, legal_balls, phase and is_powerplay all encode "how late is it" and the fit splits that effect across them with offsetting signs. Varying one alone describes deliveries that never occur (a batter on 100 off the median 10 balls faced) and can even invert the curve. Grouping real deliveries keeps the correlation structure intact.

[8]
scored = test.with_columns(pl.Series("pred", p_test))

def shape(col, expr, xlabel, ax, min_n=300):
    g = (scored.with_columns(expr.alias("bucket"))
                .group_by("bucket").agg(pl.len().alias("n"),
                                        pl.col("pred").mean().alias("predicted"),
                                        pl.col(TARGET).mean().alias("actual"))
                .filter(pl.col("n") >= min_n).sort("bucket"))
    ax.plot(g["bucket"], g["actual"], "o-", color=BLUE, lw=2, ms=5, label="observed")
    ax.plot(g["bucket"], g["predicted"], "s--", color=ORANGE, lw=2, ms=5, label="model")
    ax.set(xlabel=xlabel, ylabel="wicket rate")
    ax.legend(frameon=False, fontsize=8)
    return g

fig, axes = plt.subplots(1, 2, figsize=(9.5, 3.4))
shape("bat_runs", (pl.col("bat_runs") // 10 * 10), "striker's score (10-run bands)", axes[0])
axes[0].set_title("Striker's score")
shape("over", (pl.col("over") + 1), "over", axes[1])
axes[1].set_title("Over")
axes[1].set_xticks(range(1, 21, 2))
plt.tight_layout()
notebook figure

Which features carry weight

Spline bases are summed back to their parent feature. Note these are standardised coefficients; read them as "how much the model leans on this", not as a causal effect.

[9]
names = model.named_steps["pre"].get_feature_names_out()
coefs = np.abs(model.named_steps["lr"].coef_[0])

agg = {}
for n, c in zip(names, coefs):
    key = n.split("__")[1].rsplit("_sp_", 1)[0]
    agg[key] = agg.get(key, 0.0) + float(c)
top = sorted(agg.items(), key=lambda t: -t[1])[:12][::-1]

fig, ax = plt.subplots(figsize=(6, 3.6))
ax.barh([k for k, _ in top], [v for _, v in top], color=BLUE, height=0.7)
ax.set(xlabel="summed |standardised coefficient|", title="What the model leans on")
ax.grid(axis="y", visible=False)
plt.tight_layout()
notebook figure

Gradient boosting using LightGBM and XGBoost

Trees should beat the logistic regression, for two reasons: they find thresholds like bat_runs >= 50 without being handed a spline basis, and they can express interactions (over × wickets_in_hand) that an additive model cannot.

Two deliberate choices, both about honesty of the probabilities:

  • No class rebalancing. scale_pos_weight or SMOTE would wreck calibration, and 14.6k positives is plenty of signal. We want a usable 8%, not a balanced confusion matrix.
  • Shallow and regularised, at 15 leaves / depth 4 and min_child_samples=400. The realistic ceiling here is a ~2% log-loss gain; an unconstrained booster spends its capacity memorising noise. Early stopping watches 2024, never the test seasons.
[10]
import lightgbm as lgb
import xgboost as xgb

from model_trees import CATS, FEATS, frames, to_pandas

# Trees take the features raw - no splines, no scaling, NaN handled natively - and `ground`
# joins as a native categorical. frames() shares one category vocabulary across the splits:
# Mullanpur debuted in 2024, so it is absent from the training frame and XGBoost errors on a
# level it has never seen. Only level *names* are shared, never target information.
levels, (tr, va, te) = frames()
Xtr, Xva, Xte = (to_pandas(d, levels) for d in (tr, va, te))
ytr, yva, yte = (d[TARGET].to_numpy() for d in (tr, va, te))

lgbm = lgb.LGBMClassifier(
    objective="binary", n_estimators=3000, learning_rate=0.02, num_leaves=15,
    min_child_samples=400, subsample=0.8, subsample_freq=1, colsample_bytree=0.7,
    reg_lambda=10.0, verbose=-1, random_state=0)
lgbm.fit(Xtr, ytr, eval_X=Xva, eval_y=yva, eval_metric="binary_logloss",
         callbacks=[lgb.early_stopping(100, verbose=False)])

xgbm = xgb.XGBClassifier(
    objective="binary:logistic", n_estimators=3000, learning_rate=0.02, max_depth=4,
    min_child_weight=50, subsample=0.8, colsample_bytree=0.7, reg_lambda=10.0,
    enable_categorical=True, tree_method="hist", eval_metric="logloss",
    early_stopping_rounds=100, random_state=0)
xgbm.fit(Xtr, ytr, eval_set=[(Xva, yva)], verbose=False)

print(f"LightGBM stopped at {lgbm.best_iteration_} trees")
print(f"XGBoost  stopped at {xgbm.best_iteration} trees")
LightGBM stopped at 230 trees
XGBoost  stopped at 279 trees
[11]
preds = {
    "null (base rate)": np.full(len(yte), BASE),
    "logistic + splines": p_test,
    "lightgbm": lgbm.predict_proba(Xte)[:, 1],
    "xgboost": xgbm.predict_proba(Xte)[:, 1],
}
preds["lgb + xgb"] = (preds["lightgbm"] + preds["xgboost"]) / 2

null_ll = log_loss(yte, preds["null (base rate)"])
rows = []
for name, p in preds.items():
    ll = log_loss(yte, p)
    rows.append({"model": name, "logloss": ll, "gain_%": 100 * (null_ll - ll) / null_ll,
                 "ROC_AUC": None if name.startswith("null") else roc_auc_score(yte, p),
                 "PR_AUC": None if name.startswith("null") else average_precision_score(yte, p)})
display(pl.DataFrame(rows).with_columns(pl.col("logloss").round(5), pl.col("gain_%").round(2),
                                        pl.col("ROC_AUC").round(4), pl.col("PR_AUC").round(4)))
shape: (5, 5)
modelloglossgain_%ROC_AUCPR_AUC
strf64f64f64f64
"null (base rate)"0.198870.0nullnull
"logistic + splines"0.195471.710.60670.0788
"lightgbm"0.194342.280.61770.084
"xgboost"0.194322.290.61830.084
"lgb + xgb"0.19432.290.61830.0843
[12]
def calibration(p, y, bins=10):
    return (pl.DataFrame({"p": p, "y": y})
              .with_columns(pl.col("p").qcut(bins, labels=[str(i) for i in range(bins)]).alias("b"))
              .group_by("b").agg(pl.col("p").mean().alias("pred"), pl.col("y").mean().alias("act"))
              .sort("b"))

fig, ax = plt.subplots(figsize=(4.4, 4.4))
lim = 0.14
ax.plot([0, lim], [0, lim], color=GREY, lw=1, ls="--")
ax.annotate("perfect", (lim * 0.74, lim * 0.79), color=GREY, fontsize=8, rotation=38)
for name, colour, marker in [("logistic + splines", BLUE, "o"), ("lightgbm", ORANGE, "s")]:
    c = calibration(preds[name], yte)
    ax.plot(c["pred"], c["act"], marker=marker, ls="-", color=colour, lw=2, ms=5, label=name)
ax.set(xlabel="predicted probability", ylabel="observed wicket rate",
       title="Calibration on 2025–26", xlim=(0, lim), ylim=(0, lim))
ax.set_aspect("equal")
ax.legend(frameon=False, fontsize=8, loc="upper left")
plt.tight_layout()
notebook figure

Importance

Gain importance rewards a feature for every split it wins, which flatters high-cardinality categoricals: ground has 37 levels, so it gets many chances to carve off a slice of noise. The ablation below is the check that matters: retrain without a feature group and see whether the test loss actually moves.

[13]
gains = sorted(zip(FEATS, lgbm.booster_.feature_importance("gain")), key=lambda t: -t[1])
total = sum(v for _, v in gains)
top = [(k, 100 * v / total) for k, v in gains[:12]][::-1]

fig, ax = plt.subplots(figsize=(6, 3.6))
colours = [ORANGE if k in ("ground", "venue_wicket_rate", "venue_runs_per_ball") else BLUE
           for k, _ in top]
ax.barh([k for k, _ in top], [v for _, v in top], color=colours, height=0.7)
ax.set(xlabel="% of total gain", title="LightGBM importance (orange = venue, see ablation)")
ax.grid(axis="y", visible=False)
plt.tight_layout()
notebook figure
[14]
import pandas as pd

def ablate(drop):
    feats = [f for f in FEATS if f not in drop]
    cats = [c for c in CATS if c not in drop]
    def prep(d):
        X = d.select(feats).to_pandas()
        for c in cats:
            X[c] = pd.Categorical(X[c], categories=levels[c])
        return X
    m = lgb.LGBMClassifier(**lgbm.get_params())
    m.fit(prep(tr), ytr, eval_X=prep(va), eval_y=yva, eval_metric="binary_logloss",
          callbacks=[lgb.early_stopping(100, verbose=False)])
    p = m.predict_proba(prep(te))[:, 1]
    return {"dropped": ", ".join(drop) if drop else "(nothing)",
            "logloss": round(log_loss(yte, p), 5),
            "gain_%": round(100 * (null_ll - log_loss(yte, p)) / null_ll, 2),
            "ROC_AUC": round(roc_auc_score(yte, p), 4)}

display(pl.DataFrame([
    ablate([]),
    ablate(["ground"]),
    ablate(["ground", "venue_wicket_rate", "venue_runs_per_ball"]),
    ablate(["h2h_dismissals", "h2h_balls"]),
]))
shape: (4, 4)
droppedloglossgain_%ROC_AUC
strf64f64f64
"(nothing)"0.194342.280.6177
"ground"0.194352.270.6182
"ground, venue_wicket_rate, ven…0.194352.270.6184
"h2h_dismissals, h2h_balls"0.194472.210.6155

Permutation importance, and the collinearity trap

Gain importance is measured while the trees are built. Permutation importance is measured on a model that is already finished: shuffle one column of the test set so it keeps its distribution but no longer lines up with its rows, re-score, and see what the damage is. It is expressed as a share of the model's skill, meaning the log-loss ground the model gains over predicting the base rate for every ball.

The catch is that a shuffled feature can be covered for by a surviving duplicate. legal_balls is over at ball resolution, so breaking either one alone leaves the model most of the clock, and both look far cheaper than they are. Shuffling them as a pair is the honest measurement, and the gap between the two readings is large enough to change the story.

[15]
# Shuffle a column so it keeps its distribution but no longer matches its row, re-score, and
# charge the feature whatever loss that cost, as a share of skill (null log loss - model log loss).
rng = np.random.default_rng(0)
ll_base = log_loss(yte, lgbm.predict_proba(Xte)[:, 1])
skill = null_ll - ll_base


def perm_pct(cols, repeats=5):
    """% of skill lost when `cols` are shuffled together. Returns (mean, sd) over repeats."""
    X, losses = Xte.copy(), []
    for _ in range(repeats):
        idx = rng.permutation(len(X))       # one index for the whole group, so a collinear pair
        for c in cols:                      # stays internally consistent (legal_balls is still
            X[c] = Xte[c].values[idx]       # 6 x over) and only its link to the outcome breaks
        losses.append(log_loss(yte, lgbm.predict_proba(X)[:, 1]))
    for c in cols:
        X[c] = Xte[c]
    return 100 * (np.mean(losses) - ll_base) / skill, 100 * np.std(losses) / skill


# Every feature alone, plus the two known-redundant pairs measured jointly.
groups = {f: [f] for f in FEATS}
groups["over + legal_balls"] = ["over", "legal_balls"]
groups["team_wickets + wickets_in_hand"] = ["team_wickets", "wickets_in_hand"]

rows = []
for name, cols in groups.items():
    pct, sd = perm_pct(cols)
    rows.append({"feature": name, "pct_of_skill": round(pct, 1), "sd": round(sd, 1)})

perm = pl.DataFrame(rows).sort("pct_of_skill", descending=True)
display(perm.head(15))
shape: (15, 3)
featurepct_of_skillsd
strf64f64
"over + legal_balls"66.65.5
"legal_balls"24.72.3
"bat_career_dismissal_rate"17.41.8
"over"13.42.2
"bat_career_balls"7.10.7
"bat_runs"2.10.5
"rrr_minus_crr"1.60.5
"team_runs"1.41.1
"bat_career_sr"1.30.7
"team_wickets + wickets_in_hand"1.10.8

Reading it

over + legal_balls is 68.2% of skill. Alone they are 13.9% and 23.5%, which sum to 37.4%, and that sum is a number worth distrusting: it is the two clocks each measured while the other was still standing, so each covered for the other and both came in cheap. Adding them does not recover the joint figure, it roughly halves it. The pair is the backbone of the model by a wide margin, and the 13.9 / 23.5 split between them is an artefact of the seed rather than a finding.

wickets_in_hand is −0.1%, so permuting it very slightly helps. That is the signature of a feature the model has no use for: it is 10 - team_wickets, and team_wickets (0.8% alone, 1.1% as a pair) is already carrying that information. Small negatives elsewhere, venue_wicket_rate at −0.8% and bowl_career_balls at −0.4%, say the same thing, and are consistent with the ablation above, where dropping the venue block did not cost anything either.

Note the sd column: over + legal_balls carries ±6.0 across five shuffles. Anything in this table under about a point is inside the noise, and there are a lot of features down there.


Reading the result

Gradient boosting beats the logistic regression, and by roughly the margin worth expecting: about a 2.3% log-loss gain over the base rate against the LR's 1.7%, ROC-AUC 0.618 vs 0.607. LightGBM and XGBoost land within 0.0001 log loss of each other and averaging them adds nothing. When two well-regularised boosters agree this closely, you are at the information ceiling of the feature set, not at a modelling-choice problem.

Put the gain in perspective: ball-level wicket timing is mostly irreducible randomness. A 2% improvement on log loss is real and it holds on seasons excluded from fitting (but repeatedly inspected during this exploratory analysis), but nobody is predicting individual wickets. The useful output is a calibrated probability rather than a yes/no call, which is why calibration and log loss lead here, and accuracy never appears.

What the ablation showed is the most useful negative result in the notebook: venue contributes nothing. ground looks like 9% of gain importance, but dropping it and both venue-rate features leaves test log loss unchanged. That matches the direct measurement: the true between-ground spread in wicket rate is only sd 0.0014 once binomial noise is subtracted. Stadium strongly affects scoring (1.30 to 1.51 runs/ball across grounds) and barely affects wickets per ball, because teams adjust their aggression to conditions.

Where the remaining signal actually is: over and legal_balls (how late it is), then the batter's career dismissal rate and how long the current pair has been in.

Next steps, roughly in order of expected value

  1. batting_position, still the biggest gap. A No. 10 is 2.4× more likely to get out per ball than an opener; team_wickets is only a proxy (r = 0.88) and bat_career_dismissal_rate under-reads tailenders badly, because the 120-ball prior shrinks exactly the players with the least history toward league average.
  2. Bowler type (pace/spin), not in Cricsheet but derivable from a bowler's own career pattern. Probably the largest missing signal after batting position.
  3. wicket_bowler as the target, dropping run-outs as a separate causal process.
  4. Monotone constraints on over and bat_runs in LightGBM. Cheap, and stops the model inventing non-monotone wiggles in sparse regions.

A second model, built from what the importances said

Three findings from above feed into this:

  1. Most features do nothing. 23 of the 37 score at or below 0.4% of skill, which is inside the shuffle noise. Every one of them still costs the model capacity to fit.
  2. ground is not worth its place. It looks like the third most important feature by gain, but the ablation showed dropping it changes test log loss by nothing.
  3. Redundant columns should go. wickets_in_hand is 10 - team_wickets, and crr is recoverable from team_runs and legal_balls.

As an exploratory candidate: cut to the 14 features that appeared to earn their place, and regularise harder now that there is less to overfit to (reg_lambda 10 → 30, min_child_samples 400 → 1000). Averaging 5 seeds costs nothing and removes the run-to-run wobble.

The cell below also tests batting_position, which FEATURES.md calls the biggest known gap.

[16]
# ---- the 14 features that survived the importance and ablation checks ----
DEAD = ["runs_required", "bat_sr", "bowl_career_balls", "bowl_career_wicket_rate",
        "bowl_career_runs_per_ball", "bat_is_new", "h2h_dismissals", "h2h_balls",
        "balls_since_boundary", "venue_wicket_rate", "venue_runs_per_ball",
        "bat_balls_since_boundary", "bowl_wickets", "bowl_runs", "balls_left",
        "wickets_in_hand", "is_powerplay", "is_chase", "innings", "bat_dot_pct",
        "toss_won", "crr", "ground"]
KEEP = [f for f in FEATS if f not in DEAD]
print(f"{len(FEATS)} features -> {len(KEEP)}:  {', '.join(KEEP)}\n")

NEW = dict(objective="binary", n_estimators=4000, learning_rate=0.02, num_leaves=15,
           min_child_samples=1000, subsample=0.8, subsample_freq=1, colsample_bytree=0.7,
           reg_lambda=30.0, verbose=-1)


def frame(d, feats):
    X = d.select(feats).to_pandas()
    for c in CATS:
        if c in feats:
            X[c] = pd.Categorical(X[c], categories=levels[c])
    return X


def ensemble(feats, params, n_seeds=5, data=(tr, va, te)):
    """Fit n_seeds models and average their test predictions."""
    a, b, c = (frame(d, feats) for d in data)
    ps = []
    for s in range(n_seeds):
        m = lgb.LGBMClassifier(**{**params, "random_state": s})
        m.fit(a, ytr, eval_X=b, eval_y=yva, eval_metric="binary_logloss",
              callbacks=[lgb.early_stopping(150, verbose=False)])
        ps.append(m.predict_proba(c)[:, 1])
    return np.mean(ps, axis=0)


p_new = ensemble(KEEP, NEW)

# ---- batting position: order in which each player first appears in the innings ----
# Not a leak: a batter's position is fixed the moment they walk out, before they face a ball.
def add_position(d):
    d = d.with_row_index("_i")
    order = (pl.concat([d.select(["match_id", "innings", "_i", pl.col("batter").alias("_p")]),
                        d.select(["match_id", "innings", "_i",
                                  pl.col("non_striker").alias("_p")])])
             .group_by(["match_id", "innings", "_p"]).agg(pl.col("_i").min().alias("_f"))
             .with_columns(pl.col("_f").rank("ordinal").over(["match_id", "innings"])
                           .cast(pl.Int32).alias("batting_position")))
    return d.join(order.select(["match_id", "innings", "_p", "batting_position"]),
                  left_on=["match_id", "innings", "batter"],
                  right_on=["match_id", "innings", "_p"], how="left")


pos = [add_position(d) for d in (tr, va, te)]
p_pos = ensemble(KEEP + ["batting_position"], NEW, n_seeds=3, data=pos)

# ---- compare against the model from earlier in this notebook ----
rows = []
for name, p in [("null (base rate)", np.full(len(yte), BASE)),
                ("old: 37 features", lgbm.predict_proba(Xte)[:, 1]),
                ("new: 14 features, 5 seeds", p_new),
                ("new + batting_position", p_pos)]:
    ll = log_loss(yte, p)
    rows.append({"model": name, "logloss": round(ll, 5),
                 "gain_%": round(100 * (null_ll - ll) / null_ll, 3),
                 "ROC_AUC": None if name.startswith("null") else round(roc_auc_score(yte, p), 4)})
display(pl.DataFrame(rows))
37 features -> 14:  bat_runs, bat_balls, over, rrr, team_wickets, bowl_balls, team_runs, legal_balls, balls_since_wicket, rrr_minus_crr, bat_career_balls, bat_career_dismissal_rate, bat_career_sr, phase
shape: (4, 4)
modelloglossgain_%ROC_AUC
strf64f64f64
"null (base rate)"0.198870.0null
"old: 37 features"0.194342.2750.6177
"new: 14 features, 5 seeds"0.194232.3340.6192
"new + batting_position"0.194252.3220.619

What it did and did not do

The pruned model and the batting-position variant land extremely close together. The exact table above is the source of truth; earlier prose embedded stale numbers from a previous run. Cutting features is operationally useful because it simplifies the model, but a tiny change in log loss should not be called a generalisation improvement merely because it is larger than seed-to-seed wobble. The match-clustered interval below is the relevant uncertainty check.

batting_position does not show a practically useful incremental signal here. Its raw effect is real, but bat_career_dismissal_rate, team_wickets, and bat_balls already encode much of the same information. This is evidence of redundancy in this feature set, not evidence that batting position never matters.

The broader conclusion is unchanged: the current features mainly identify when wickets become more likely. Material improvement probably needs delivery or shot information, not another small hyperparameter sweep.


Bowler type: the feature that was not in the data

The paragraph above ends by naming bowler type as the obvious candidate for new information, and noting it is not in Cricsheet. It is now in the data.

fetch_bowler_style.py builds player_style.csv by chaining three sources: Cricsheet's own register maps every player to a Cricinfo ID, Wikidata maps that ID to an en.wikipedia article, and the article's infobox carries bowling and batting fields in prose that normalises cleanly. Cricinfo itself refuses scripted requests, so its ID serves only as a join key. Wikidata's own bowling style property (P2545) exists and is unusable: 43 of our 577 bowlers, with a visibly bot-damaged value distribution (835 "left-arm orthodox spin" against 2 "off spin"). 25 uncapped players with no article were filled in by hand; see docs/BOWLER_STYLE_TODO.md. Coverage is 577/577 bowlers, every delivery.

Nothing here is derived from the ball data, so there is nothing to leak: a bowler's action and a batter's stance are fixed properties of the player, known before the season starts.

Four columns, of which the last two carry the actual cricket:

column
bowler_type pace / spin
bat_hand the striker's stance, on 99.7% of balls faced
same_handed right-arm to a right-hander, left to a left
turn_into_batter which way a spinner's ball moves relative to the bat

turn_into_batter is the variable a commentator would name: finger spin turns into a batter of the same handedness as the bowling arm (off-break into a right-hander, orthodox into a left-hander), and wrist spin turns into the opposite one. Null for pace, where turn is not the question.

One expectation to set before the numbers. The wicket rate gap between pace and spin is 0.0518 to 0.0455, which looks promising until you remember that the two bowl in different overs, and this model's single largest input is the clock. The same trap as batting_position.

[17]
# ---- the raw effect, and the confound that eats it ----
fig, axes = plt.subplots(1, 3, figsize=(12.5, 3.4))

spin_share = (df.filter(pl.col("over") < 20).group_by("over")
                .agg((pl.col("bowler_type") == "spin").mean().alias("spin")).sort("over"))
axes[0].bar(spin_share["over"] + 1, spin_share["spin"], color=BLUE, width=0.72)
axes[0].set(xlabel="over", ylabel="share of balls bowled by spin",
            title="Spin bowls the middle", xticks=range(1, 21, 3))

by_over = (df.filter(pl.col("over") < 20).group_by(["over", "bowler_type"])
             .agg(pl.len().alias("n"), pl.col(TARGET).mean().alias("rate"))
             .filter(pl.col("n") >= 300).sort("over"))
for t, colour in [("pace", ORANGE), ("spin", GREY)]:
    v = by_over.filter(pl.col("bowler_type") == t)
    axes[1].plot(v["over"] + 1, v["rate"], "o-", color=colour, lw=2, ms=4, label=t)
axes[1].set(xlabel="over", ylabel="wicket rate", title="Within an over, the gap mostly closes",
            xticks=range(1, 21, 3))
axes[1].legend(frameon=False, fontsize=8)

# stumpings are the one thing only spin does - a tell the target already contains
kinds = (df.filter(pl.col(TARGET) == 1).group_by("bowler_type")
           .agg((pl.col("dismissal_kind") == "stumped").mean().alias("stumped"),
                (pl.col("dismissal_kind") == "bowled").mean().alias("bowled"),
                (pl.col("dismissal_kind") == "caught").mean().alias("caught")))
x = np.arange(3)
for i, (t, colour) in enumerate([("pace", ORANGE), ("spin", GREY)]):
    v = kinds.filter(pl.col("bowler_type") == t)
    axes[2].bar(x + (i - 0.5) * 0.38, [v["caught"][0], v["bowled"][0], v["stumped"][0]],
                width=0.36, color=colour, label=t)
axes[2].set(xticks=x, ylabel="share of dismissals", title="How the wicket falls")
axes[2].set_xticklabels(["caught", "bowled", "stumped"])
axes[2].legend(frameon=False, fontsize=8)
plt.tight_layout()

display(df.group_by("bowler_type").agg(
    pl.len().alias("balls"), pl.col(TARGET).mean().round(4).alias("wicket_rate"),
    # dismissal_kind is null on non-wickets and mean() skips nulls, so this is
    # stumpings as a share of *dismissals*, not of balls
    (pl.col("dismissal_kind") == "stumped").mean().round(4).alias("stumped_share_of_wkts"),
    pl.col("over").mean().round(2).alias("mean_over")).sort("bowler_type"))
shape: (2, 5)
bowler_typeballswicket_ratestumped_share_of_wktsmean_over
stru32f64f64f64
"pace"1908320.05180.00099.02
"spin"1047250.04550.07969.52
notebook figure
[18]
# ---- does it help? same 14-feature model, same 5 seeds, one column at a time ----
# frame() above casts only the original CATS, so the new string columns need their own
# vocabulary; everything else is identical to the model in the cell above.
CATS_X = CATS + ["bowler_type", "bowler_arm", "bat_hand"]
levels_x = {**levels,
            **{c: sorted(df[c].unique().drop_nulls().to_list())
               for c in ("bowler_type", "bowler_arm", "bat_hand")}}
MATCHUP = ["bowler_type", "bowler_arm", "bat_hand", "same_handed", "turn_into_batter"]


def frame_x(d, feats):
    X = d.select(feats).to_pandas()
    for c in CATS_X:
        if c in feats:
            X[c] = pd.Categorical(X[c], categories=levels_x[c])
    return X


def ensemble_x(feats, params=NEW, n_seeds=5):
    a, b, c = (frame_x(d, feats) for d in (tr, va, te))
    ps = []
    for s in range(n_seeds):
        m = lgb.LGBMClassifier(**{**params, "random_state": s})
        m.fit(a, ytr, eval_X=b, eval_y=yva, eval_metric="binary_logloss",
              callbacks=[lgb.early_stopping(150, verbose=False)])
        ps.append(m.predict_proba(c)[:, 1])
    return np.mean(ps, axis=0)


rows = []
style_predictions = {}
for name, feats in [("14 features (baseline)", KEEP),
                    ("+ bowler_type", KEEP + ["bowler_type"]),
                    ("+ bowler_type, bat_hand", KEEP + ["bowler_type", "bat_hand"]),
                    ("+ full matchup", KEEP + MATCHUP)]:
    p = ensemble_x(feats)
    style_predictions[name] = p
    ll = log_loss(yte, p)
    rows.append({"model": name, "n_feats": len(feats), "logloss": round(ll, 5),
                 "gain_%": round(100 * (null_ll - ll) / null_ll, 3),
                 "ROC_AUC": round(roc_auc_score(yte, p), 4),
                 "PR_AUC": round(average_precision_score(yte, p), 4)})
display(pl.DataFrame(rows))
shape: (4, 6)
modeln_featsloglossgain_%ROC_AUCPR_AUC
stri64f64f64f64f64
"14 features (baseline)"140.194232.3340.61920.0839
"+ bowler_type"150.19422.3470.61940.0841
"+ bowler_type, bat_hand"160.194232.3320.61920.0838
"+ full matchup"190.194212.3440.61920.0838
[19]
# Paired uncertainty: resample matches, not individual balls.
old_test = lgbm.predict_proba(Xte)[:, 1]
comparisons = [
    ("pruned vs old 37-feature model", old_test, p_new),
    ("+ batting position vs pruned", p_new, p_pos),
    ("+ bowler type vs pruned", p_new, style_predictions["+ bowler_type"]),
    ("+ full matchup vs pruned", p_new, style_predictions["+ full matchup"]),
]
uncertainty = []
for name, reference, candidate in comparisons:
    ci = clustered_log_loss_gain(
        yte, reference, candidate, te["match_id"].to_numpy(), n_boot=5_000
    )
    uncertainty.append({
        "comparison": name,
        "gain_x1e4": round(10_000 * ci.estimate, 3),
        "95%_low_x1e4": round(10_000 * ci.low, 3),
        "95%_high_x1e4": round(10_000 * ci.high, 3),
    })
display(pl.DataFrame(uncertainty))
shape: (4, 4)
comparisongain_x1e495%_low_x1e495%_high_x1e4
strf64f64f64
"pruned vs old 37-feature model"1.161-0.7513.018
"+ batting position vs pruned"-0.244-0.7270.236
"+ bowler type vs pruned"0.264-0.2070.726
"+ full matchup vs pruned"0.2-0.5380.936

Positive values favour the named candidate; the units are $10^{-4}$ log-loss points. A 95% interval that crosses zero means the apparent ordering is not stable across matches. This is a more demanding and useful check than comparing random seeds on the same fixed test rows.

What it was worth

model log loss gain over null ROC-AUC
14 features (baseline) 0.19423 +2.33% 0.6192
+ bowler_type 0.19420 +2.35% 0.6194
+ bowler_type, bat_hand 0.19423 +2.33% 0.6192
+ full matchup 0.19421 +2.34% 0.6192

Nothing here survives the match-clustered uncertainty check. Every 95% interval crosses zero, including the pruned model versus the original. The simpler 14-feature model remains attractive for maintainability, but the score ordering should be read as indistinguishable on this sample.

The unconditional pace/spin gap is real: pace takes a wicket on 5.18% of deliveries and spin on 4.55%. It is also heavily confounded by when each type bowls. Once the model already knows the clock and batter state, bowling type adds no stable incremental wicket signal.

bat_hand is useful only as part of a matchup, not as an isolated property. Even the full matchup does not improve wicket log loss reliably here. The same information is more useful in the run model, where bowling type changes the outcome distribution more than it changes the marginal probability of a dismissal.

One important caveat on the dismissal-type chart: stumpings describe how a wicket fell, which is known only afterwards. They are shown for interpretation and are not used as model inputs.

Where that leaves the next steps

  1. wicket_bowler as the target, separating bowler-caused dismissals from run-outs.
  2. A rolling-origin evaluation across seasons, with feature choices made before each holdout.
  3. Delivery or shot information; the current state features cannot distinguish two otherwise similar balls whose physical execution differs by a few centimetres.