Where the data comes from
The foundation is a deep well of minor league history: roughly 162,000 individual player-seasons stretching from 1978 through 2013, covering every level from complex-league rookie ball up through Triple-A. Once a new season gets underway, those same leagues’ stats are pulled straight from Major League Baseball’s own statistics feed and folded into that historical pool in the same format. That means a player who debuted last week is measured on the exact same scale as someone who played in 1985. The comparisons aren’t limited to recent history; they draw on decades of it.
Turning raw stats into rates
A box score full of hits, walks, and strikeouts doesn’t mean much on its own: a player with 400 plate appearances and one with 40 can’t be judged by raw totals. So the first step converts everything into rates: how often a hitter walks or strikes out per plate appearance, their on-base and slugging percentages, and a power measure that isolates extra-base pop from batting average. Speed gets its own treatment, blended from how often a player succeeds when stealing a base, how often they attempt one, how often they stretch a double into a triple, and how often their legs turn a single into a run. Pitchers get the same idea applied to their side of the game: strikeout rate, walk rate, and home-run rate per batter faced, along with a workload measure that distinguishes a starter’s outings from a reliever’s.
Adjusting for context
A .300 on-base percentage in a pitcher-friendly rookie league and a .300 on-base percentage in a hitter-friendly Triple-A league aren’t the same accomplishment. Every rate stat is measured against the average for that specific league, level, and season, so performance is judged relative to the competition a player actually faced. Two further adjustments sharpen that picture. The first accounts for sample size: a player with only a few dozen plate appearances has a rate stat that’s still mostly noise, so it gets pulled toward the league average until more playing time proves it’s real, while a player with hundreds of plate appearances is judged much more on their own numbers. The second accounts for age: performing well while young for your level says a lot more about future potential than the same performance from someone repeating a level at an older age, so younger players get a modest boost and older players a modest discount.
Putting everyone on the same scale
Once a stat has been adjusted for league, sample size, and age, it’s converted into a standardized score that says how many standard deviations above or below average a player is. That step is what makes it possible to compare a hitter’s power to their strikeout rate, or a hitter in 1991 to a hitter in 2011, on one common scale. For hitters, the traits carried through in this form are plate discipline (walks and strikeouts), power, and an overall speed score. For pitchers, it’s strikeout rate, walk rate, home-run rate, and how a pitcher is used.
Finding the closest matches
With every player reduced to the same handful of standardized traits, comparing two players becomes a matter of measuring how far apart they are on each one, then combining those gaps into a single similarity score. Traits that tend to matter more for future success (like power and strikeout avoidance for hitters, or missing bats for pitchers) are weighted more heavily than traits that matter less. That combined score is what determines a player’s list of closest comps. One more rule keeps the comparisons honest: a prospect is only measured against how other players performed up through that same point in their career, not against what those players went on to do later. A player currently in High-A is compared to other players’ High-A-and-below track records, never to a level they haven’t reached yet.
Building a career profile
Rather than comparing single seasons in isolation, the site builds a running career-to-date profile for every player, updated as they climb each level. A player’s more recent seasons carry more weight in that profile than seasons from further back, so the picture reflects who a player is right now rather than treating every year of their career as equally telling.
Excluding stats after a player used up rookie eligibility
A career profile is meant to capture a prospect’s development, not the occasional rehab assignment or tune-up outing a player takes once he’s already an established big leaguer. So once a player has exceeded MLB’s rookie limits (130 career at-bats for hitters, 50 career innings for pitchers), any minor league seasons from that point forward are left out of his career profile and out of the comp pool entirely, even though they still show up in his season-by-season stat log on the site. Seasons up through the one in which he crossed that limit are still included.
For the same reason, a player only serves as a comp target once he’s 28 or older as of last season — old enough that his major league career (or lack of one) is treated as settled rather than still in progress. The site’s “Rookie-eligible” filter uses this same 28-year cutoff, alongside the rookie-limit check above, to narrow the matches list down to players who still read as active prospects.
What the comps suggest for the future
Once a prospect’s closest comps are identified, the next question is simple: what did those comps actually go on to do? The site looks at each comp’s career major league performance — leaving aside any comp who never reached the majors at all, since a bust has no big-league performance to measure, just an absence of one — and, weighting more-similar comps more heavily, builds a realistic range of outcomes: a below-average case, a typical case, and an above-average case. Alongside that range, the site separately shows what share of comparable players made the majors at all, since simply getting there is its own part of the story a stat-line range can’t tell on its own. It’s not a single prediction so much as an honest picture of the range of paths players like this one have actually taken.
How players are ranked
By default, the Matches list isn’t sorted by median comp outcome alone. That median is only calculated from the comps who actually reached the majors, so on its own it says nothing about how many of a player’s comps got there in the first place — a player whose few successful comps did very well, but whose pool is mostly players who never made the majors at all, can look identical by median outcome to a player whose comps almost all reached, even though those are very different levels of realistic risk.
So the default ranking blends the median outcome together with a few other signals: how often a player’s comps actually reached the majors at all, the best-case (90th-percentile) outcome among those comps rather than just the typical one, and — for pitchers specifically — whether a player profiles as a starter or a reliever, since a shutdown reliever and a solid full-workload starter can look similar by a rate stat alone despite very different real value.
Each of those signals is measured by comparing a player against other current prospects — specifically, the same “rookie-eligible” population the site’s own eligibility filter uses elsewhere — rather than against the full 45-plus years of players in the database. That keeps the ranking about how a player stacks up against today’s prospect pool, not diluted by decades of established major leaguers’ old minor league track records.
This blended ranking is what Comp Strength Value reflects, and it’s the list’s default sort. The individual numbers behind it — median outcome, best-case outcome, reach probability — are still shown on every player and can still be sorted on directly if you’d rather look at one in isolation.
Measuring big-league success: wRAA and fRAA
To rank comps and build a range of outcomes, a comp’s entire major league career needs to be boiled down to one number, and that number needs to mean the same thing whether the comp was a hitter or a pitcher. For hitters, that number is wRAA (weighted Runs Above Average): every plate appearance is valued using the same run-value weights Fangraphs uses for wOBA (a walk is worth less than a double, a double less than a home run, and so on), compared against that season’s league-average hitter, and summed across a career. A positive wRAA means a hitter created more runs than a league-average hitter would have gotten from the same number of plate appearances; a negative wRAA means fewer.
Pitchers get the mirror-image treatment with fRAA (FIP Runs Above Average): each season’s FIP (Fielding Independent Pitching, a run estimate built only from strikeouts, walks, hit-by-pitches, and home runs, the outcomes a pitcher controls most directly regardless of the defense behind them) is compared against that season’s actual league-average FIP, converted into a runs-per-nine-innings gap, and accumulated over every inning a pitcher threw. Because both stats land on the same real-runs-above-average scale, a Median, 90th-percentile, or Ceiling comp value reads the same way no matter which side of the ball a player is on. Comps who never reached the majors are left out of that range rather than counted as a zero-value outcome — a bust has no big-league production to average in, they simply never got the chance. That risk is instead reflected separately, in the share of comps who reached the majors at all (see “How players are ranked” above), so a strong comp pool with a real bust risk still reads as strong on outcome while the risk itself stays visible rather than silently dragging the outcome numbers down.
What you see on the site
The public site is a trimmed-down window into the full database: any player who’s had roughly ten games’ worth of a look is included, each shown alongside their twenty closest comps. That works out to different raw numbers for hitters and pitchers because they accumulate their relevant stat at different rates per game. Hitters get about 4.3 plate appearances per game, while pitchers face about 12.0 batters per appearance once you blend starters and relievers together. Ten games’ worth of playing time is roughly 40 career plate appearances for a hitter, and roughly 120 batters faced for a pitcher. Those two numbers look mismatched at a glance, but they’re aiming at the same underlying bar: a real, if brief, look at how a player performs.