Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Tuesday, April 21, 2009

Putting the "Erm..." in Sabermetrics

Particularly absurd post on Mets Geek charting each team's bullpen's average fastball speed against strikeouts. Don't even try to wrap your mind around what it might even mean to average the pitch speeds of an entire bullpen together, the results are random noise, not "a slight trend."

Tuesday, September 25, 2007

Mets vs. Nationals

Oh man, what a brutal game. Once again, solid-looking offense that still fails to deliver while our pitching falls apart. Pelfrey was looking so good at the beginning. It just seems weird that the Mets could outhit Washington and yet end up nine runs down. I mean, we had a bunch of walks, but nine runs' worth? I guess we were making hits when they didn't matter, while the Nats made theirs count. And how is Ronnie Belliard so good against the Mets defense?

One of the viewer questions answered by the broadcast team was about the OPS offensive stat, which I hadn't heard of before. At first it sounded pretty reasonable from a math standpoint—a sum of two percentages, which isn't extremely rigorous as far as quantifying events, but comes pretty close for disjoint events (e.g., the percentage of days in which I play tennis plus the percentage of days in which I play squash is very close to the percentage of days in which I play racquet sports, since I don't play other racquet sports and very rarely play both in the same day) and is very easy to calculate.

(I don't have a problem with the loose application of terms like "average" and "percentage" in baseball stats. Bad assumptions about what these kinds of terms say about the underlying data is already a huge problem in the general population's understanding of things like economic data ("the median income is rising!"), so the more we can encourage people to think about what dynamic is actually driving a given number, whether it's called "average" or "percentage" or something else, the better.)

But if you actually start to think about what the two stats that are added together to get OPS, on-base percentage and slugging percentage, actually stand for, the meaningfulness of it starts looking a little questionable.

Slugging percentage is pretty straightforward: it's a hitter's total number of bases gained on hits, divided by their at-bats. As a predictive metric, that's easy to understand: we should expect a player slugging .500 to get about five bases in ten at-bats. The stat appears to say something very simple about a player's hitting ability, and only their hitting ability. In statistical terms, it's the expected value of bases a runner will hit for in an official at-bat.

On-base percentage is a little more subtle. It's a ratio between two multiple-term expressions, but basically what it boils down to is, absent any mistakes or decisions made by a manager or a defensive player other than the pitcher, how likely is it that a given batter currently at the plate will end up on base? This means that on-base percentage incorporates some offensively valuable metrics—the ability to draw a walk, for example—that are not accounted for by slugging percentage. But note also that, statistically, on-base percentage is a probability, not an expected value.

So what do these things say about a player when they're added together? By the admission of the people who came up with OPS, nothing real. But the OPS does represent something like "offensive value" that neither on-base nor slugging percentage captures on its own: a given player may get more extra base hits than another, but their greater ability to draw walks means that in the end they bring a similar offensive value to the team.

Is there a more meaningful way of calculating that same information? I'm not sure. Wikipedia makes it sound as though OPS was conceived as a more easily calculated alternative to SLOB (the product of the two stats instead of the sum) and "runs created" (the product of SLOB and at-bats). Simple addition "works" as an alternative because the two statistics are of the same order of magnitude. But neither SLOB nor runs created deals with the problem of trying to combine an expected value and a probability.

I might prefer instead to replace the "hits" term in the numerator of the on-base percentage with the "total bases" value used to measure slugging. The resulting figure would be a valid expected value, but taking into account plate appearances, such as walks, that are excluded from official at-bats. At the very least, we should no longer allow difficulty of calculation to limit the statistics used for comparing aspects of play. There may be any number of reasons for eschewing SLOB, but the fact that it requires a sliderule to determine should not be among them.

Anyway, all of that aside, OPS as a measure of offensive value reminds me of IQ as a measure of human intelligence. Indeed, there is even a statistic called OPS+, which measures on-base and slugging percentages against the park-adjusted totals for the league, and then scales them so the league average is at 100. A player with an OPS+ over 150 could be said to have a genius-level offense, while one with an OPS+ of 50 might be considered an offensive imbecile.

In fact, the original purpose of OPS+ was to identify challenged players who required extra instruction during batting practice. Then later it was misapplied in an attempt to demonstrate that Negro League batters were innately inferior to their Major League counterparts. (Okay, this paragraph is made up.)

OPS and OPS+ are useful metrics for comparing the offensive abilities of different players. OPS+ probably correlates pretty well with RBI and does a fair job of predicting career runs, not to mention player salary. A pocket calculator with access to the statistics could be programmed to organize starting batting orders strictly by OPS+ and probably come pretty close to making the same decisions as a Major League manager. If you had to guess which non-pitchers (nullus) on an A-class rookie squad would go on to make the big leagues, you could probably do worse than to pick the group of players whose OPS+ figures are over, say, 125.

And yet, OPS is not something unknown that we've managed to measure with some sort of test and then stuck with because it's just so damn useful. There's no mystery behind what OPS "is": it's just something we've defined arbitrarily, this formula of different well-understood concrete statistics cobbled together, and which we know for a fact has no intrinsic meaning except as an abstraction. It's a convenience, a tool for making comparisons without having to get into the fine details of the fundamental attributes (both inborn and learned) like upper body strength, mental reaction time, experience at the game of baseball, and sprinting ability, many of which might prove difficult to measure, if not impossible. We would be mistaken to take the utility of OPS as evidence of an underlying phenomenon of the baseball player body, some organ the fitness of which determines offensive ability.

Implicit in our OPS-centered discussion of "offensive value" is the knowledge that no single figure can hope to sum up all of what makes a given player valuable in an offense, and that real offensive value is only truly meaningful in the context of an actual in-game situation. We have no reason to consider the existence of "general offensive ability" anything other than a statistical abstraction; and certainly not as the product of some "general offensive ability factor" behind all the various quantifiable and non-quantifiable aspects that figure into offensive value.

With one exception, everything that I've stated about OPS and OPS+ applies analogously to IQ. Both can be useful metrics for making comparisons, can be used to estimate or predict more concrete information to which we may not have access, and can lead to poor decision-making when misapplied or taken out of context. The exception is that whereas we can say exactly how we determine OPS, how the actual figure, devoid of intrinsic meaning though it may be, is calculated from raw data. In contrast, we have no real idea what an IQ test measures. We call it "intelligence," but there's no breakdown in terms of rates of neural activity, or brain volume, or anything measurable beyond the fact that it correlates with a bunch of other things to an extent that we find it useful in making comparisons and predictions. We give it extrinsic meaning, but there is still nothing at all to believe that IQ "is" anything, that it comes from anywhere, that it is anything other than a statistical abstraction, like OPS+, useful to us only because our limited minds cannot deal with the whole array of information, some of it perhaps not even quantifiable, that lurks behind that single figure.

Friday, August 31, 2007

The Disappearance of .400 Hitting

I've been reading this Stephen Jay Gould book, Full House, which is all about how various things in the world that we usually describe as "trends" are in actuality better understood as side effects of changes in statistical variation. One of the two main examples he uses throughout the book is the disappearance of batting averages over .400 over the last century. This "trend" does not represent a decline in hitting ability over time, but rather a narrowing of variation in hitting ability.

As the excellence of play in professional baseball has increased over the years, the difference in ability between the best and worst batters and the best and worst fielders has become smaller. In addition, MLB management has occasionally tweaked the rules to keep the mean batting average essentially constant from season to season. Also, the fact that .000 represents a hard lower limit (you can't have a negative batting average) means that the distribution of averages will skew right. As a consequence of these three factors, the right-hand tail of the batting average distribution curve has been brought in closer: whereas a century or so ago, the handful of batters out on the end of that curve were hitting above .400, that narrower curve now puts them significantly below .400. Here's some graphs showing how the curve has changed shape over the years:



So I've pretty much been thinking of all sorts of trends that may or may not be understood better with this sort of model. One that came to mind this morning was the stock market. It is sort of a general rule that market indexes and averages march ever upward. There are corrections and recessions and stuff, but the overall trend seems to be one of continuous increase. Does this represent a miracle of capitalism, that all firms will tend to gain in value over time, or is it merely a side effect of a widening variation among firms?

I didn't look for historical data—I know that the mean value of publicly traded firms has increased, so all I'm really interested in is the shape of the distribution of firms' values at a single point in time—but I was able to find the current values of all the US corporations listed on the NASDAQ. Here's the histogram showing how many firms are valued in each range (Y-axis is number of companies, X-axis is value in millions of dollars as of market close on 8/31):



It's the same kind of right-skewed curve as with batting averages, which makes sense: $0 represents another hard lower limit on firm value, because a company can't have negative worth without disappearing from the market. Gould talked about the idea of a "drunkard's walk": if a drunk leaves a bar, and is staggering back and forth on a sidewalk between the bar on one side a ditch on the other, then even if he staggers completely randomly (even chance of staggering towards the wall or towards the ditch) then he will still always end up in the ditch.

In the same way, it could be completely random whether a given company increases or decreases in value each day: some firms will move to the right in the distribution curve, some will move left, but the left limit will never drop below zero, and some very few (Microsoft) will inch the right limit further out. There needs be no general tendency for company values to increase, just variation increasing in the only direction it can over time.

Of course, that doesn't mean that the value of a company is random. Just that the shape of the curve could be explained by randomness. To actually demonstrate that it's all random you'd need to track individual firms over time. And in fact I would expect there to actually be tendency to increase in value as a result of the steadily increasing labor pool. But it is interesting that the growing mean value of corporations is not indicative of anything greater than an increase in variability.