The Edition
Tournament State
Who wins Williamsport?
How these odds are computed
Next Up
Tonight's slate
Daily Feature
How hard is the climb?
One evidence-backed view of the road from district play to Williamsport
The next-level shock
Same teams, before and after advancingBlowouts thin out
10+ run margins · 2021–25What if? / Illustrative simulator
The champion against every rung
How this estimate works
Every rank on this site is a position on one ranked list, produced by one code path for every team. This is a retrospective rating of games already played, not a forecast — how it is built, how accurate it is out of sample, and its known limitations are all set out on .
Game day
Matchup
Player
Power Rankings
How good has each team looked so far — every team, at whatever stage its season reached
What this list is, and what it is not
This is a retrospective ranking. It asks how strongly has this team played in the games it actually played. It opens scoped to the teams on the regional / Williamsport pathway — the pool that produces the 20 Williamsport teams — including the ones whose 2026 season is already over; use the pathway filter below for the full state-and-regional population. It is not a list of who is best in the country right now, because most of these teams are no longer playing.
For the forward-looking question — who can still win something in 2026 — use .
| Rank | Team | Status | Type | Rating | RD | Off | Def | SOS | Percentile | Games | Record |
|---|
Still Alive
Only the teams that can still play a 2026 game — same ratings, forward-looking scope
| # | Team | Path to Williamsport | Type | Rating | RD | Record | Why it is still alive |
|---|
Regions & jurisdictions
State and provincial jurisdictions and regional brackets, all on the one unified list
| Region / jurisdiction | Type | Teams | With results | Best team | Best rating | Median | Player data | Both tiers |
|---|
Who wins each region
Every game already played is held at its real result; only genuinely undecided games are simulated
Brackets & results
Straight from the production input 2026_baseball_regionals.json — no game, score or pairing is inferred
What changed
Every rating this release moved, and the specific published game that moved it — kept as a record, so the ranks cited here are the ones in force at the time of the change
Newly rated
Teams with no rated identity at all in the currently shipped rankings
| Rank | Team | Type | Rating | Admitted game | Source |
|---|
Ratings that moved
Each caused by a specific admitted game, with the capture hash it was verified against
| Rank | Team | Type | Shipped | Candidate | Δ | Admitted game | Source |
|---|
Team profiles
Everything here is published data — absences read “not published”, never zero
| # | Team | Type | Record | GP | RS | RA | RS/G | RA/G | Diff | SB | CS |
|---|
| # | Team | Type | Rating | RD | Off | Def | SOS | Percentile | Region Win% |
|---|
Region Win% is shown only for teams whose regional bracket structure is sourced and currently live — most state-tournament-only teams show –, not 0.
Players
Four separate questions, each with its own qualification line — not one list that quietly answers a different question for every reader.
| # | Player | Team | G | AB | R | H | RBI | BB | SO | 2B | 3B | HR | TB | SB | CS | AVG | OBP | SLG | OPS |
|---|
| # | Player | Team | PA | BB% | K% | BB:K | ISO | XBH% | SB% | BABIP |
|---|
BABIP is descriptive at this sample size (2–3 games per player), not a skill signal — it says what happened on batted balls, not what to expect next.
| # | Player | Team | App | IP | H | R | ER | BB | SO | ERA6 | WHIP |
|---|
| # | Player | Team | BF | K:BB | K/6 | BB/6 | Strike% | K%−BB% | WP | HBP | BK |
|---|
No FIP here — a hard data gap, not a caution: pitcher-attributed home-runs-allowed isn’t in the schema yet, since attributing an opponent HR to the correct reliever needs play-by-play sequencing this project doesn’t capture from the box-score page alone.
| Player | Team | G | AVG | OPS | Batting | App | ERA6 | WHIP | Pitching |
|---|
Each row is two separate, already-qualified profiles for the same real player — this project does not join batting and pitching data into one merged record or one combined score (see ). Click either link to that profile's own full stat line.
Batting
| Player | Team | G | PA | AVG | OPS |
|---|
Pitching
| Player | Team | App | IP | ERA6 | WHIP |
|---|
Three audits
What this stage examined and what it decided, including the two that admitted nothing
Every evidence row, by decision
186 rows across 8 sources. 14 were admitted
Sources
Hash policy: a re-computed SHA-256 where a capture was retained, null where none exists
| Source | Publisher | Capture | Hash validated | Rows | Admitted |
|---|
The 14 admitted games
Every one carries a retained snapshot whose hash this stage re-computed and matched
| Row | Game | Date | Tournament | Tier verified | Capture SHA-256 |
|---|
Taiwan row-level tie-back
25 rows tied back to the captured bytes exactly — and still excluded
| Game | Row as captured | Date anchor | Score | Tie-back |
|---|
Three questions, answered straight
What a reader actually wants to know before trusting a number on this site
How accurate is this, really? 66.6% honest out-of-sample (Brier 0.213) — measured only on games the model never trained on. A do-nothing static region prior with no in-season updates at all scores 70.1%, better than the live model. That comparison is not buried; it is limitation 2, below.
Does it account for pitching, injuries, or who's actually available to throw? No. That adjustment was built, tested twice, and shipped as NULL both times — limitation 4.
Is any of this estimated or made up? No score, opponent, or record is fabricated anywhere in this package. Missing data — pitching for 92% of teams, any backtest at all for the state/provincial half — is reported as unknown, never filled in as zero or average. See Coverage and limitation 5 below.
Everything below is the same disclosure this page has always carried in full — the eight limitations, the decision record, what shipped and what didn't. Open it for the detail.
Open full methodology — all eight limitations, coverage, and the decision record
How the ratings are built
And the eight things that are wrong with them, stated by the package itself
Glicko-2 with stakes × margin-of-victory × recency weighting, seeded from a Bradley-Terry region-strength prior fit on 2013–2025 Williamsport games — extended from five seasons to thirteen, then human-reviewed and shipped (details below). The core update formula itself ran unmodified through every candidate tested and in the published build. Engine as-of date: 2026-08-16, and the true UTC date at the run is 2026-08-16. They first agreed on 2026-08-11: v1.0.47 corrected the production engine's own constant and v1.0.48 corrected the three further as-of constants this unified/live-site chain carries, so nothing is now discounted for a stale calendar. The constant is still hardcoded, so it goes stale again tomorrow.
Limitations, stated plainly
Quoted from executive_summary.md and coverage_and_limitations.md
- Honest out-of-sample accuracy is 66.6% (Brier 0.213). The often-quoted 77.7% is leaky and overstates real skill by ~11 points.
- The model is not clearly better than naive baselines, and is worse than the static region prior with no in-season updates at all (70.1%, Brier 0.224). Rating movement inside a single tournament should be read cautiously. Margin-of-victory weighting makes essentially no measurable difference (−0.25 accuracy points).
- The state/provincial half of this product has never been independently backtested. Same engine, no ground-truth check. The LOYO backtest validates the regional/Williamsport population only.
- The pitching adjustment is NULL. Two separate ablation tests both concluded the shipped model should stand. No rating here is adjusted for pitcher rest, workload or availability.
- Player data is thin: 33 teams (8.3% of the 398 teams this was measured across) have any harvested player-level data, all of it U.S. regional GameChanger box scores. Missing player stats mean unknown, never zero, everywhere in this package.
- Recency weighting was never backtested; its 10-day (regional) and 21-day (state) half-lives are documented choices, not fitted parameters.
- NEW, measured by this stage — the confidence label is inflated. In the shipped
glicko2.js, a down-weighted game shrinks rating deviation faster, so games the model deliberately discounted make it report more confidence. Disclosed, not fixed. It does not block publication: the behaviour is identical in the previously shipped model. - NEW — the engine’s
AS_OF_DATEis stale and its recency term is symmetric, so a game played after it is decayed as if equally far in the past. Using the true current date moves 334 of 398 ratings . That is a larger lever than every evidence-admission question in this release combined.
Offense, defense and strength of schedule
The three extra columns on the rankings table — what they are, and the one thing they are not for
Tested and not shipped
Three rigorous, predeclared-backtest model changes that did not beat production
This project runs every candidate model change through the shipped model's own strictly causal leave-one-year-out backtest before it can ship, with the pass/fail bar fixed before the numbers are computed. Three such candidates came back worse than production and were not shipped. They are shown here in full, not buried, because a project that only publishes its wins is not showing you its rigor. The candidates that beat their bar and shipped are in the next section.
recommendation.json and
ADDENDUM.md, named above.Tested and promoted
The two changes that cleared their predeclared bars, passed human review, and shipped on 2026-08-09
Investigated and settled
Not a rejected candidate and not a shipped change — a question asked, answered, and closed
Why 47 merged teams don’t contradict the rejected merges
The two merge candidates above were rejected — and the 47 merged rows on this site are real. Both are true. Here is how.
What was tested and rejected was a general model change: all ~76 bridgeable real-world teams, merged with naive full pass-through of state evidence into the regional tier. It failed for a measured reason — a state champion is by construction a team that won its state tournament, so full-weight merging inflated every bridged entrant’s rating upward and never downward, and partial bridge coverage turned that one-directional bias into an error the model cannot correct for. A follow-up then measured how that loss scales with the transfer weight λ, and found the bias channel shrinks superlinearly as λ falls.
The 47 merged rows on this site are a different object in two specific ways:
- Scope. This is not a population-wide model change. It is an explicitly human-authorized exception for 47 teams whose state↔regional identity was verified against littleleague.org primary sources — zero fuzzy matching, byte-equal name resolution only. It began as 22 pairs; on 2026-08-10 the user authorized extending it to the 25 further bridges that had since been verified to the same byte-equal standard against the same primary sources, which also collapsed six real-world teams that had been appearing twice. That extension was authorized against this project’s own backtest, which found the merge mechanism statistically indistinguishable from not merging at all — the stated reason was identity coherence, not predictive accuracy.
- Mechanism. The state→regional boundary does not use the full pass-through that failed. It uses a calibrated transfer weight taken from that follow-up’s own backtest — λ = 0.25, the setting that cut the measured inflation bias by ~75% — blending the state-side rating toward the region prior and widening rating deviation at the crossing, because evidence that transfers weakly should increase uncertainty, not confidence. The section→state boundary carries full weight for the opposite measured reason: section evidence was measured to predict state-tournament outcomes at 54–72 rating points per SD, a real coefficient, so discounting it there would throw away signal.
So the caution the two rejected findings established is not abandoned — it is applied narrowly: verified identities only, and a measured, cited transfer weight at the one boundary that needed one. The selection-bias mechanism those stages found is real and is why the merge scope stays limited to verified identities rather than all bridges. A control run with the merge switched off is the standing proof that a fully unmerged rebuild reproduces the published ratings exactly. It is not a claim that every other team’s rating number is immune to a change anywhere in the graph: Glicko-2 updates an opponent's rating using the other side's rating at the time of the game, so any change ripples forward through whoever played whoever next — the release that introduced the merge (22 teams at the time) moved 30 other teams' ratings this way (up to ~70 points), and the wider section carry-through moves a further, smaller set the same way. What stays fixed is each team's own inputs — its own games, its own seed, its own record — never a fabricated one.
The confidence-inflation measurement
Controlled calls to the unmodified updateRating(); only the per-game weight varies
| Per-game weight | Resulting rating | Resulting RD |
|---|
Coverage
What is complete, what is partial, what is not published
- Complete — the 2026 international and Canada regional brackets already carried in the production input: Latin America 29 games, Panama 50, Europe-Africa 28, Asia-Pacific 28, Caribbean 24, Australia 42, Japan 15, Mexico 37 scored of 45 cards, Curaçao 10, Canada 13.
- Partial / in progress — the ten U.S. regions. Six were mid-tournament at the snapshot; four (Great Lakes, Mountain, New England, Northwest) had no rated team identity at all until this candidate admitted their first games.
- Prior only at that snapshot — three U.S. regions had no published bracket when this record was frozen: West, Metro, Mid-Atlantic. Their official brackets have since been recovered and completed, and their teams are rated on the unified list; the prior-only treatment described here no longer applies.
- Not published — round labels for the 8 newly admitted GameChanger games. None was invented; the shipped classifier’s default weight of 1.0 applies. Two further scenarios measure the consequence.
- Player and pitcher data — 33 of 398 teams (8.3%), from 26 validated GameChanger box scores covering U.S. regional games only. State and provincial tournaments have no pitch-count coverage in this snapshot.
A team’s absence from the pitching data means no data was harvested. It never means zero pitches,
zero innings, or no workload concern. published_view_manifest.json records
zero_never_substituted_for_unknown: true.
How this release was built
The decision record behind this data release — what was recommended, and what had to be weighed against it
Why it was recommended
What had to be weighed against it
Not without a separate decision
Recovering nine empty tournament pages
How the list grew from 398 rated teams to 430, and what stayed empty on purpose
| Jurisdiction | Teams | Games | Result |
|---|
What this product is not
Four sentences the package ends on
- It is not a new model formula. The v1 Glicko-2 update ran unmodified in every stage on this site, including the unified build — what changed is the prior’s history depth, the source population, and 47 verified identities, each promoted through its own gate record.
- It is not a player-stat product. Only 33 teams have any harvested player-level data, all of it U.S. regional GameChanger box scores.
- It is not a general cross-tier fusion. 47 primary-source-verified real-world teams carry one continuous rating across tiers, each on an explicit advancement link plus a byte-equal jurisdiction label; every bridge that fails that standard — District of Columbia, Czechia National — stays unmerged.
- It is not a validated predictor of Williamsport outcomes beyond its measured 66.6% honest out-of-sample accuracy — see the limitations above before quoting any single rating.
Production notes
Typography, flags and data provenance — the fine print behind every page
Typography. Source Serif 4 sets display type and every numeral — it is the only face here with tabular figures, so a column of ratings actually lines up. Libre Franklin sets running text, team names and data column headers. Both are under the SIL Open Font License 1.1, subset to Latin and embedded directly in the page rather than fetched from a font host.
Flags. Every ranked team carries its state, provincial or national flag, resolved by jurisdiction or region — never by matching a team's own display name — against 97 verified flag images sourced from Wikipedia and (for the Chinese Taipei entrant) the IOC's own Olympic-banner convention. A team falls back to a neutral initials chip only if a future build ever ships without a matching flag; today, none do.
Data. Every rating, rank, record and game log on this page reads from one gate-verified computation covering all 428 teams. Profile evidence tables, section results and advancement notes are display enrichment layered on top of it, and the brackets reproduce the published tournament source unchanged. No game, score, pairing, bio or source is invented; missing values read “not published”, never zero.
Status. Reflects the production computation as of 2026-08-16. Build provenance — every version gate, the backtests behind each shipped change, and the data files behind each section — is set out across this tab.